Multi-modal data generation method, device, system and equipment and storage medium

By generating object-centered multimodal data, the applicability problem of multimodal data acquisition equipment when the object position changes is solved, and the accuracy and flexibility of robot operation are improved.

CN120645211APending Publication Date: 2025-09-16PASSINI ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510762006.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing multimodal data acquisition equipment cannot be applied when the position of an object changes, resulting in inaccurate robot operation.

Method used

By acquiring object image data, first joint image data and second joint posture data, the posture data of the fingertip in the object coordinate system is determined, and multimodal data centered on the object is generated.

Benefits of technology

It enables the use of multimodal data when the position of an object changes, is suitable for scenarios where the positions of different objects change, and improves the accuracy and flexibility of robot operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120645211A_ABST
    Figure CN120645211A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal data generation method, device and system, equipment and a storage medium. The method comprises the following steps: acquiring object image data, first joint image data, second joint pose data and force / touch sensing data; according to the object image data at any same time point, determining object pose data of the operated object under an image sensor coordinate system at the same time point; determining fingertip pose data of a fingertip in an object coordinate system according to the object pose data, the first joint image data and the second joint pose data at the same time point; and generating multi-modal data including object pose data, fingertip pose data and force / touch perception data at each same time point. According to the technical scheme of the embodiment of the invention, the multi-modal data taking the object as the center is generated, and the method can be suitable for different object position change scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a method, apparatus, system, device, and storage medium for generating multimodal data. Background Art

[0002] A robot multimodal data acquisition device is a device used to collect multiple types of data required for robot operation. It can simultaneously collect different forms of information, such as visual information, sound information, and position information, so that the robot can perceive the environment more comprehensively and make intelligent decisions.

[0003] However, while the data collected by existing multimodal data acquisition equipment enables robots to replicate the user's hand movements and perform a series of operations according to the user's hand movements, during actual operations, robots often need to focus more on the manipulation of the object being manipulated rather than the movement trajectory or state of their hands. If the position of the manipulated object changes during execution, the current multimodal data recorded by the robot is no longer applicable. Summary of the Invention

[0004] The present invention provides a multimodal data generation method, apparatus, system, device and storage medium to generate object-centered multimodal data, thereby being applicable to different object position change scenarios.

[0005] According to one aspect of the present invention, a multimodal data generation method is provided, which is applied to a multimodal data generation system. The method is executed by a controller in the multimodal data generation system, comprising:

[0006] Acquiring object image data, first joint image data, second joint posture data, and force / tactile perception data;

[0007] Determine, based on the object image data at any same time point, the object pose data of the operated object in the image sensor coordinate system at the same time point;

[0008] Determining fingertip posture data of the fingertip in the object coordinate system according to the object posture data, the first joint image data, and the second joint posture data at the same time point;

[0009] Generate multimodal data including object posture data, fingertip posture data and force / tactile perception data at the same time point.

[0010] According to another aspect of the present invention, a multimodal data generation device is applied to a multimodal data generation system, wherein the device is configured in a controller in the multimodal data generation system, and includes:

[0011] A data acquisition module is used to acquire object image data, first joint image data, second joint posture data and force / tactile perception data;

[0012] An object posture determination module is used to determine the object posture data of the operated object in the image sensor coordinate system at any same time point based on the object image data at the same time point;

[0013] a fingertip posture determination module, configured to determine fingertip posture data of the fingertip in the object coordinate system based on the object posture data, the first joint image data, and the second joint posture data at the same time point;

[0014] The multimodal data generation module is used to generate multimodal data including object posture data, fingertip posture data and force / tactile perception data at the same time point.

[0015] According to another aspect of the present invention, a multimodal data generation system is provided, comprising a controller, an image acquisition device, a motion capture glove reference component, a motion capture glove rigidly connected to the motion capture glove reference component, and a tactile module disposed at the end of the motion capture glove and connected to the fingertips; the motion capture glove is disposed with a posture sensor;

[0016] The image acquisition device is used to photograph the operated object to generate object image data during the execution of the object operation task, and photograph the reference component of the motion capture glove to obtain first joint image data, and send the object image data and the first joint image data to the controller;

[0017] The motion capture glove is used to collect second joint posture data through the posture sensor during the execution of the object operation task, and send the second joint posture data to the controller;

[0018] The tactile module is used to generate force / tactile perception data during the execution of the object operation task, and send the force / tactile perception data to the controller;

[0019] The controller is used to execute the above-mentioned multimodal data generation method.

[0020] According to another aspect of the present invention, an electronic device is provided, comprising:

[0021] at least one processor; and

[0022] a memory communicatively connected to the at least one processor; wherein,

[0023] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can perform the multimodal data generation method described in any embodiment of the present invention.

[0024] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the multimodal data generation method described in any embodiment of the present invention when executed.

[0025] The technical solution of the embodiment of the present invention determines the object posture data of the operated object in the image sensor coordinate system through object image data, and determines the fingertip posture data of the fingertips in the object coordinate system based on the object posture data, the first joint image data and the second joint posture data, thereby realizing the generation of object-centered multimodal data, so that when the position of the operated object changes during actual operation, the multimodal data can still be used, which is suitable for scenarios where the positions of different objects change.

[0026] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0028] Figure 1A is a flowchart of a multimodal data generation method provided according to embodiment 1 of the present invention;

[0029] Figure 1B 1 is a schematic structural diagram of a multimodal data generation system provided according to the first embodiment of the present invention;

[0030] Figure 2 is a flowchart of a multimodal data generation method provided according to the second embodiment of the present invention;

[0031] Figure 3 1 is a schematic structural diagram of a multimodal data generating device provided according to a third embodiment of the present invention;

[0032] Figure 4 3 is a schematic diagram of the structure of an electronic device for implementing the multimodal data generation method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0033] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0034] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0035] Example 1

[0036] Figure 1A This is a flowchart of a multimodal data generation method provided in Example 1 of the present invention. This embodiment is applicable to the case of generating object-centered multimodal data. The method can be performed by a multimodal data generation device, which can be implemented in the form of hardware and / or software. The multimodal data generation device can be configured in an electronic device, which can be a controller in a multimodal data generation system.

[0037] This method can be applied to multimodal data generation systems, such as Figure 1B The structure diagram of a multimodal data generation system shown in FIG. 1 is a diagram showing the structure of a multimodal data generation system. The multimodal data generation system 10 includes a controller 11, an image acquisition device 12, a motion capture glove reference component 13, a motion capture glove 14 rigidly connected to the motion capture glove reference component 13, and a tactile module 15 deployed at the end of the motion capture glove 14 and connected to the fingertips; a posture sensor is deployed on the motion capture glove 14.

[0038] The image acquisition device 12 is used to capture an image of the object being manipulated during the object manipulation task, as well as an image of the first joint of the motion capture glove reference component 13. For example, the image acquisition device 12 may be a visual camera. The visual camera may be a wearable RGBD (Red, Green, Blue, and Depth) camera capable of simultaneously capturing color images and depth information.

[0039] In actual use, the image acquisition device 12 can be worn on the user's chest and secured with a strap. Upper body movement during filming can prevent the manipulated object and the motion capture glove's reference components from being obscured. The image acquisition device 12 uploads the captured image data to the controller 11 for real-time or offline processing.

[0040] The motion capture glove reference component 13 can serve as a reference or benchmark component for the motion capture glove 14. The motion capture glove reference component 13 is rigidly connected to the motion capture glove 14 to serve as a fixed reference point for the motion capture glove, providing a benchmark for the subsequent pose calculation of the second joint. For example, the motion capture glove reference component 13 can be a pair of hexahedral visual bracelets with a hollowed-out center for the arm to pass through. Each of the six sides can be affixed with a different QR code to represent a different side, facilitating the precise identification and positioning of the motion capture glove reference component during the execution of a task operation.

[0041] The motion capture glove reference component 13 is rigidly connected to the motion capture glove 14, and each side of the motion capture glove reference component 13 has a clear size and position relationship with the motion capture glove 14. During actual operation, the six-sided design of the motion capture glove reference component 13 ensures that at least one side of the QR code is visible during hand operation, thus avoiding obstruction.

[0042] During actual operation, the motion capture glove 14 is worn over the operator's hands. A posture sensor can be deployed on the motion capture glove 14 to collect posture data of the second joint in the second joint coordinate system. The posture sensor can, for example, be composed of multiple EMF (Electromagnetic Field Sensor) sensors or IMU (Inertial Measurement Unit) sensors with a certain degree of metal interference resistance.

[0043] It should be noted that the first joint and the second joint are any two joint positions from the arm to the fingertip; wherein the second joint is arranged close to the fingertip, and the first joint is arranged away from the fingertip. In this embodiment, the first joint described above can be the wrist joint of the motion capture glove reference component 13; the second joint can be the finger or palm joint of the motion capture glove 14. Taking the motion capture glove reference component 13 as an example of a hexahedron visual bracelet, during the execution of the object operation task, the image acquisition device 12 performs image acquisition on the hexahedron visual bracelet, and can obtain two-dimensional code images of different faces, and the acquired two-dimensional code image is used as the first joint image. During the execution of the object operation task, the posture sensor deployed on the motion capture glove 14 collects the posture of the finger or palm joint, and the collected posture of the finger or palm joint is used as the second joint posture.

[0044] The tactile module 15 is mounted at the end of the motion capture glove 14 and connected to the fingertips. It measures force and tactile information during contact between the fingertips and the object being measured. This tactile module can utilize a multi-tactile area array sensor, capable of measuring nearly fifteen types of force and tactile information, including distributed force, resultant force, resultant torque, temperature, hardness, viscosity, slip, static friction, and dynamic friction.

[0045] The image acquisition device 12, motion capture glove 14, and haptic module 15 are each connected to a controller 11 in communication, for example, by cables, to ensure data transmission speed and reliability. The controller 11 is configured to transmit sensor data collected by the controller 11 for real-time or offline processing. The controller 11 can be worn by the wearer in the form of a backpack and primarily includes components such as a processor, a power supply, and external interfaces.

[0046] It should be noted that when the operator wears the above-mentioned devices, the device calibration program needs to be started, specifically the calibration of the motion capture gloves and the tactile module. Specifically, the posture sensor of the motion capture gloves obtains the posture information of each finger joint, specifically the posture information of each finger joint relative to the zero position. Therefore, before using the device, the operator needs to put on the motion capture gloves and place the palm flat on the table for zero-position calibration. At this time, the posture sensor will record the zero-position information of each finger joint. The contact force measurement of the tactile module is relative information. During the zero-position calibration, the fingertip contact force needs to be zero. At this time, the tactile module records the zero-position information of the contact force and calibrates the device based on the zero-position information.

[0047] The method of this embodiment is applied to the multimodal data generation system 10 and is executed by the controller 11 in the multimodal data generation system 10, such as Figure 1A As shown, the method specifically includes:

[0048] S110 , acquiring object image data, first joint image data, second joint posture data, and force / tactile perception data.

[0049] S120 : Determine object pose data of the operated object in the image sensor coordinate system at any same time point based on the object image data at the same time point.

[0050] S130 , determining the fingertip posture data of the fingertip in the object coordinate system according to the object posture data, the first joint image data, and the second joint posture data at the same time point.

[0051] S140: Generate multimodal data including object posture data, fingertip posture data, and force / tactile perception data at the same time points.

[0052] The object image data may be obtained by an image acquisition device capturing images of an operated object during an operator's operation task.

[0053] The first joint image data can be obtained by an image acquisition device capturing an image of the motion capture glove reference component. The first joint image data can reflect the spatial position and posture of the joints of the motion capture glove reference component during the execution of an object manipulation task. Taking the motion capture glove reference component as an example, the image acquisition device captures an image of the visual bracelet during the execution of the manipulation task to obtain bracelet image data, and the bracelet image data is used as the first joint image data. If the visual bracelet is a hexahedral bracelet with different QR code images attached to different surfaces, the bracelet image data captured is the QR code image data, and the QR code image data is used as the first joint image data.

[0054] Among them, the second joint posture data can be obtained by collecting the posture data of the operator's finger or palm joints during the execution of the object operation task by the posture sensor deployed on the motion capture glove; the second joint posture data can reflect the posture changes of the operator's finger or palm joints during the execution of the object operation task.

[0055] The first joint may be the joint position of the motion capture glove reference component. If the motion capture glove reference component is worn on the operator's wrist, the first joint may represent the wrist joint. The second joint may be the joint position of the finger or palm of the hand wearing the motion capture glove.

[0056] The force / tactile perception data can be obtained by collecting the tactile and force data generated by the fingers relative to the object during the operation of the object by the tactile module.

[0057] It should be noted that due to the different acquisition frequencies of the image acquisition equipment, motion capture gloves and tactile module, the collected object image data, first joint image data, second joint posture data and force / tactile perception data do not belong to the same time dimension. Therefore, before processing the object image data, first joint image data, second joint posture data and force / tactile perception data, the data can be aligned in the time dimension to obtain the object image data, first joint image data, second joint posture data and force / tactile perception data at the same time point.

[0058] Optionally, after obtaining the object image data, first joint image data, second joint posture data and force / tactile perception data, the device with the lowest acquisition frequency can be selected from the image acquisition device, motion capture glove and tactile module, and the acquisition frequency of the device can be used as the reference frequency, and other data can be aligned with the reference data.

[0059] Specifically, if the acquisition frequency of the image acquisition device is 30Hz, the acquisition frequency of the motion capture glove is 60Hz, and the acquisition frequency of the tactile module is 100Hz, then the acquisition frequency corresponding to the object image data or the first joint image data acquired by the image acquisition device is used as the reference frequency, and the second joint posture data acquired by the motion capture glove and the force / tactile perception data acquired by the tactile module are aligned with the reference frequency. Exemplarily, for example, the first acquisition time points corresponding to the object image data or the first joint image data acquired by the image acquisition device are T1, T2...Tn, respectively, the second acquisition time points corresponding to the second joint posture data acquired by the motion capture glove are t1, t2...tn, respectively, and the third acquisition time points corresponding to the force / tactile perception data acquired by the tactile module are m1, m2...mn, respectively. Taking any time point Tn among the first acquisition time points as an example, the time point ts closest to the time point Tn is selected from the second acquisition time point, and the second joint posture data at the time point ts is used as the second joint posture data at the time point Tn, thereby achieving alignment of the second acquisition time point with the first acquisition time point. Similarly, the time point ms closest to time point Tn is selected from the third acquisition time point, and the force / tactile perception data at time point ms is used as the force / tactile perception data at time point Tn, thereby aligning the third acquisition time point with the first acquisition time point. Ultimately, the object image data, first joint image data, second joint posture data, and force / tactile perception data corresponding to each acquisition time point T1, T2, ..., Tn are obtained.

[0060] In order to further achieve accurate alignment of the object image data, the first joint image data, the second joint posture data and the force / tactile perception data in the time dimension, interpolation processing can also be used to perform data alignment.

[0061] In an optional embodiment, the object image data, the first joint image data, the second joint posture data, and the force / tactile perception data are aligned in the time dimension to obtain the object image data, the first joint image data, the second joint posture data, and the force / tactile perception data at the same time point. This can also be achieved in the following manner, and the specific steps include:

[0062] Step a1: Select target reference data from the object image data, the first joint image data, the second joint posture data and the force / tactile perception data, and use the other data except the target reference data as the data to be aligned.

[0063] Step a2: Obtain the benchmark collection time corresponding to the target benchmark data.

[0064] Step a3: for any data to be aligned, determine a first adjacent acquisition time and a second adjacent acquisition time adjacent to the reference acquisition time among all data acquisition times of the data to be aligned.

[0065] Step a4: Determine the alignment data of the data to be aligned at the reference acquisition time based on the first data to be aligned acquired at the first adjacent acquisition time and the second data to be aligned acquired at the second adjacent acquisition time, so as to obtain object image data, first joint image data, second joint posture data and force / tactile perception data at each same time point.

[0066] For example, any one of the object image data, the first joint image data, the second joint posture data, and the force / tactile perception data can be used as the target reference data, or the data acquired by the device with the lowest acquisition frequency can be used as the target reference data. For example, if the image acquisition device has the lowest acquisition frequency, the object image data or the first joint image data can be used as the target reference data, and the second joint posture data and the force / tactile perception data can be used as the data to be aligned.

[0067] The data acquisition time corresponding to the target reference data is used as the reference acquisition time. For any data to be aligned, determine the first adjacent acquisition time and the second adjacent acquisition time adjacent to the reference acquisition time in all the data acquisition times of the data to be aligned. Specifically, if the reference acquisition time is T1, T2, ..., Tm, ..., Tn, and the total data acquisition time of the data to be aligned is t1, t2, ..., tn, for any time point Tm of the reference acquisition time, select two time points adjacent to the acquisition time point Tm from all the data acquisition time points of the data to be aligned as the first adjacent acquisition time and the second adjacent acquisition time. For example, if time point Tm = 1000ms, and all the data acquisition time points of the data to be aligned are: t1, t2, ..., 990ms, 1100ms, ..., tn, it can be seen that the first adjacent acquisition time adjacent to Tm can be 990ms, and the second adjacent acquisition time can be 1100ms.

[0068] The data mean of the first data to be aligned acquired at the first adjacent acquisition time and the second data to be aligned acquired at the second adjacent acquisition time is used as the alignment data of the data to be aligned at the reference acquisition time. Continuing with the above example, the data mean between the data a at the first adjacent acquisition time point 990ms and the data b at the first adjacent acquisition time point 1100ms in the data to be aligned is used as the alignment data at the acquisition time point Tm=1000ms. Thus, the alignment data corresponding to the data to be aligned at the reference acquisition time points T1, T2, ..., Tn can be obtained, thereby realizing the alignment between the object image data, the first joint image data, the second joint posture data and the force / tactile perception data, and obtaining the object image data, the first joint image data, the second joint posture data and the force / tactile perception data at the same time points (T1, T2, ..., Tn).

[0069] Optionally, the method of performing data alignment on the data to be aligned at the reference acquisition time can also be: assuming that the reference acquisition time is 9.2s, the first adjacent acquisition time of the first data to be aligned is 9s, and the second adjacent acquisition time of the second data to be aligned is 10s. Assuming that the posture of the first data to be aligned is T1 and the posture of the second data to be aligned is T2, then T1 to T2 are divided into several intermediate postures, taking 10 as an example, T1.1, T1.2,…, T2 are obtained, and the intermediate postures change linearly, then the alignment data Tn with a time of 9.2s can be obtained by difference, Tn∈[T1.1, T2].

[0070] Compared with searching for the nearest data starting from the lowest frequency data, the above data alignment method has higher alignment accuracy, further reduces data errors caused by acquisition delays, and improves data alignment accuracy when frequencies do not match.

[0071] After aligning the object image data, first joint image data, second joint posture data, and force / tactile perception data, the object posture data of the manipulated object in the image sensor coordinate system at any given time point is determined. The image sensor coordinate system is a three-dimensional rectangular coordinate system established with the focal center of the image acquisition device as its origin and the optical axis of the image acquisition device as its Z-axis. Assuming the image acquisition device is a camera, the image sensor coordinate system is the camera coordinate system. The object posture data is used to describe the position and posture of the manipulated object in three-dimensional space within the image sensor coordinate system.

[0072] For the object image data at any same time point, determining the object pose data of the operated object in the image sensor coordinate system at the same time point. In an optional embodiment, determining the object pose data of the operated object in the image sensor coordinate system at the same time point based on the object image data at any same time point includes:

[0073] Step b1: According to the object image data at any same time point, based on the mask image prediction model, obtain the mask image data corresponding to the object image data at the same time point.

[0074] Step b2: input the object image data and its corresponding mask image data into a pre-trained pose estimation model to obtain the object pose data of the operated object in the image sensor coordinate system at the same time point output by the model.

[0075] The mask image prediction model is used to locate the manipulated object in the object image data and generate a segmentation mask corresponding to the object image. For example, the mask image prediction model can be a YOLO (You Only Look Once) model, a deep learning model used for target detection. The object image data can specifically be RGB image data containing the manipulated object, and the mask image data can specifically be binary image data containing the manipulated object. By inputting the object image data into the mask image prediction model, mask image data corresponding to the object image data can be obtained.

[0076] The pose estimation model is used to estimate the pose of the operated object in the image sensor coordinate system, and can be pre-trained and pre-deployed in the controller. This embodiment also provides a model training method for the pose estimation model.

[0077] In an optional embodiment, a plurality of object sample images are obtained; each object sample image contains the same type or a different type of operated object; a mask sample image corresponding to each object sample image is generated, and each object sample image and its corresponding mask sample image are used as a target sample image; each target sample image is labeled to obtain a true pose value corresponding to each target sample image; the true pose value corresponding to each target sample image is input into a pre-built pose estimation model to obtain a predicted pose output by the model; the pose estimation model is trained based on the true pose value and the predicted pose until a preset model training end condition is met, thereby obtaining a trained pose estimation model.

[0078] The mask sample images of the object sample images can be obtained by mask image prediction using a mask image prediction model. Each object sample image includes images of the area containing the manipulated object, whether of the same or different types. The true pose value is the actual position and attitude of the manipulated object in the image sensor coordinate system; the position is the object's position in three-dimensional space, and the attitude is the rotation state of the manipulated object, typically expressed as quaternions or Euler angles.

[0079] Specifically, each target sample image is labeled to obtain the true pose value of the corresponding target sample image. The labeling process can be implemented by manual labeling, automatic labeling, or semi-automatic labeling. The true pose value corresponding to each target sample image is input into a pre-built pose estimation model to obtain the predicted pose output by the model; the current loss value is determined based on the true pose value and the predicted pose; based on the current loss value, the pose estimation model is trained until the loss value reaches a set threshold, or the loss value tends to be stable, or the number of iterations reaches a set threshold, etc., to obtain a trained pose estimation model. Among them, the pre-built pose estimation model can be, for example, a PoseNet (pose estimation network) model.

[0080] The object image data and its corresponding mask image data are input into a pre-trained pose estimation model, and the model outputs the object pose data of the operated object in the image sensor coordinate system.

[0081] The above technical solution obtains a posture estimation model through pre-training and deployment, and performs image analysis on the object image data of the operated object based on the posture estimation model, thereby determining the object posture data of the operated object in the image sensor coordinate system, and realizing accurate prediction of the object posture data of the operated object in the image sensor coordinate system, thereby improving the accuracy of subsequent fingertip posture data estimation based on the object posture data, and further improving the accuracy of multimodal data generation.

[0082] The fingertip pose data of the fingertip in the object coordinate system is determined based on the object pose data, the first joint image data, and the second joint pose data at the same time point. For example, the object pose data, the first joint image data, and the second joint pose data can be input into a pre-trained fingertip pose estimation model to obtain the pose data output by the model.

[0083] The fingertip pose estimation model can be pre-trained by relevant technical personnel and deployed on the controller. For example, fingertip sample data including object pose samples, first joint image samples, and second joint pose samples can be generated, and the fingertip sample data can be labeled to obtain the fingertip pose truth value corresponding to the fingertip sample data; the fingertip sample data and its corresponding fingertip pose truth value are input into a pre-built network model to obtain the fingertip predicted pose output by the model; based on the fingertip pose truth value and the fingertip predicted pose, the network model is trained until the model training end condition is met, thereby obtaining a fingertip pose estimation model that has completed training.

[0084] Force / tactile perception data can include nearly 15 types of tactile perception information collected or measured by the tactile module, including fingertip distributed force, resultant force, resultant torque, temperature, hardness, viscosity, slip sensation, static friction, and dynamic friction. Based on the dimensional relationship between the body coordinate system and the fingertip, as well as the position of the fingertip in the object coordinate system, the force / tactile perception data can be converted to the object coordinate system and combined with the fingertip position trajectory to form a tactile perception sequence corresponding to that trajectory.

[0085] Generate multimodal data including object posture data, fingertip posture data and force / tactile perception data at the same time point.

[0086] The multimodal posture data of the present invention can be applied to human-computer interaction scenarios and machine learning scenarios. Taking the human-computer interaction scenario as an example, based on the object-centered multimodal data, a robot can imitate the movements of a human operator to complete a specific task. After the user wears the various devices of the multimodal data generation system and completes a certain operation task, the object-centered multimodal data obtained contains the three-dimensional posture of the operated object and the three-dimensional posture information of the operator's fingertips. It can be mapped to the robot's dexterous hand to replicate the user's operation task. For example, if a user grabs a bottle of dishwashing liquid and moves it from point A to point B, the various data acquisition devices in the multimodal data generation system will record multimodal information such as the grip position and grip force of the finger on the dishwashing liquid bottle, and record the path the finger needs to traverse to move the bottle from point A to point B. All the information required to complete the task of grabbing and moving the dishwashing liquid is stored in the multimodal data. This data can be used to drive the robot to complete the task. For example, when a bottle of dishwashing liquid appears in front of the robot, based on the finger posture trajectory centered on the object, the robot can imitate the user's operation, calculate the grasping position and strength of the dexterous hand on the dishwashing liquid bottle, control itself to grasp the dishwashing liquid, and then drive the robot to reproduce the user's action based on the multimodal data of the dishwashing liquid moving from A to B to complete the movement task.

[0087] Taking machine learning scenarios as an example, multimodal data can be used for imitation learning, reinforcement learning, and self-supervised learning. In imitation learning, by combining data such as visual observation, force feedback, and finger movements, robots can more comprehensively imitate human demonstrations and master complex skills such as manipulating objects or fine assembly. In reinforcement learning, multiple sensory data are used as state representations and reward signals. For example, in grasping tasks, the tactile sensor collects moderate contact force as reward information, making reinforcement learning more stable and effective, and accelerating robot strategy learning. In self-supervised learning, the natural correspondence between different modal data can be used to use the event of moving the manipulated object from point A to point B in space as a sign of the success of the robot's manipulation task, creating a self-supervised learning signal and reducing reliance on manual labeling.

[0088] The technical solution of the embodiment of the present invention determines the object posture data of the operated object in the image sensor coordinate system through object image data, and determines the fingertip posture data of the fingertips in the object coordinate system based on the object posture data, the first joint image data and the second joint posture data, thereby realizing the generation of object-centered multimodal data, so that when the position of the operated object changes during actual operation, the multimodal data can still be used, which is suitable for scenarios where the positions of different objects change.

[0089] Furthermore, the multimodal data generation method of this embodiment can be configured within the multimodal data generation system described in the embodiment. The system's components are all wearable, making them easy to carry, convenient to use, and applicable to a wider range of scenarios. During actual object manipulation operations, the motion capture glove reference component can be moved to avoid image occlusion within the field of view, ensuring the image acquisition accuracy of the image acquisition device, while also being easy to operate and cost-effective.

[0090] Example 2

[0091] Figure 2 This is a flowchart of a multimodal data generation method provided in Example 2 of the present invention. This embodiment is optimized and improved based on the above technical solutions.

[0092] Furthermore, the step "determine the fingertip posture data of the fingertip in the object coordinate system based on the object posture data, the first joint image data and the second joint posture data at the same time point" is refined into "determine the reference posture data of the fingertip in the image sensor coordinate system based on the positional relationship between the first joint and the second joint based on the first joint image data and the second joint posture data; determine the fingertip posture data of the fingertip in the object coordinate system based on the reference posture data and the object posture data." to improve the method of determining the fingertip posture data of the fingertip in the object coordinate system.

[0093] It should be noted that for the parts not described in detail in the embodiments of the present invention, reference can be made to the descriptions of other embodiments. Figure 2 As shown, the method includes the following specific steps:

[0094] S210 , acquiring object image data, first joint image data, second joint posture data, and force / tactile perception data.

[0095] S220 : Determine object pose data of the operated object in the image sensor coordinate system at any same time point based on the object image data at the same time point.

[0096] S230 : Determine reference posture data of the fingertip in the image sensor coordinate system according to the first joint image data and the second joint posture data and based on the positional relationship between the first joint and the second joint.

[0097] S240 : Determine fingertip posture data of the fingertip in the object coordinate system according to the reference posture data and the object posture data.

[0098] S250: Generate multimodal data including object posture data, fingertip posture data, and force / tactile perception data at the same time points.

[0099] The positional relationship between the first joint and the second joint is the relative positional relationship between the motion capture glove reference component worn at the first joint and the motion capture glove worn at the second joint. The relative positional relationship between the motion capture glove reference component and the motion capture glove can be predetermined and stored in the controller for subsequent parameter calculation.

[0100] For example, a visual wristband with a hexahedron as the motion capture glove base component is associated with a different QR code image. Due to the rigid connection between the motion capture glove base component and the glove, the positional and dimensional relationships between the wristband's faces and the glove are fixed. Therefore, these positional and dimensional relationships can be pre-stored in the controller for easy recall during subsequent calculations.

[0101] In an optional embodiment, determining reference pose data of the fingertip in the image sensor coordinate system based on the first joint image data and the second joint pose data and the positional relationship between the first joint and the second joint includes:

[0102] Step c1: Determine the initial posture data of the fingertip in the second joint coordinate system according to the second joint posture data.

[0103] Optionally, the second joint posture data may include a first posture of the metacarpophalangeal joint in the second joint coordinate system, a second posture of the proximal interphalangeal joint relative to the metacarpophalangeal joint, a third posture of the distal interphalangeal joint relative to the proximal interphalangeal joint, and a fourth posture of the fingertip relative to the distal interphalangeal joint; based on the first posture, second posture, third posture and fourth posture, the initial posture data of the fingertip in the second joint coordinate system is determined.

[0104] For example, if the second joint is the palm joint, the second joint coordinate system can be a spatial coordinate system with the center of the palm as the origin, the Z axis pointing toward the fingertips, the Y axis perpendicular to the back of the hand, and the X axis parallel to the palm plane. Assuming that the motion capture glove is worn at the second joint during actual operation, the second joint coordinate system is the motion capture glove coordinate system.

[0105] For example, the initial pose data of the fingertip in the second joint coordinate system is determined as follows:

[0106]

[0107] in, Represents the first pose of the metacarpophalangeal joint in the second joint coordinate system; Represents the second position of the proximal interphalangeal joint relative to the metacarpophalangeal joint; The third position of the distal interphalangeal joint relative to the proximal interphalangeal joint; In other embodiments of the present invention, the second joint coordinate system may be based on a reference component as its origin, and the reference component is provided on the forearm.

[0108] Step c2: Determine first joint posture data of the first joint in the image sensor coordinate system based on the first joint image data.

[0109] For example, a joint pose prediction model can be pre-trained to predict the first joint pose data of the motion capture glove reference component in the image sensor coordinate system. The joint pose prediction model can be trained in the same way as the pose estimation model.

[0110] Specifically, the YOLO model is used to predict the mask image data corresponding to the first joint image data, and the first joint image data and its corresponding mask image data are input into the pre-trained joint pose prediction model to obtain the first joint pose data of the motion capture glove reference component in the image sensor coordinate system output by the model. The first joint pose data is also 6D pose data.

[0111] Step c3: Determine the relative posture data between the first joint and the second joint based on the positional relationship between the first joint and the second joint.

[0112] It should be noted that the positional relationship between the first joint and the second joint is the relative positional relationship between the motion capture glove reference component worn at the first joint position and the motion capture glove worn at the second joint position, which can be measured in advance by relevant technical personnel and stored in the controller.

[0113] Taking the visual bracelet with a hexahedron as the motion capture glove reference component as an example, the first joint image data is the QR code image data of different faces; the current bracelet face can be determined based on the first joint image data, and the relative position relationship between the current bracelet face and the motion capture glove can be determined based on the pre-stored relative position relationship between each bracelet face and the motion capture glove, that is, the relative position relationship between the first joint and the second joint.

[0114] Step c4: Determine reference pose data of the fingertip in the image sensor coordinate system based on the initial pose data, the first joint pose data, and the relative pose data.

[0115] For example, the reference pose data of the fingertip in the image sensor coordinate system is determined as follows:

[0116]

[0117] in, Represents the initial pose data of the fingertip in the second joint coordinate system; Represents the first joint pose data of the first joint in the image sensor coordinate system; Represents the relative pose data between the first joint and the second joint.

[0118] Furthermore, according to the reference pose data and object pose data Determine the fingertip pose data in the object coordinate system

[0119]

[0120] The technical solution of this embodiment determines the fingertip posture data of the fingertip in the object coordinate system based on the initial posture data, the first joint posture data, the relative posture data and the object posture data, thereby achieving accurate determination of the posture data of the fingertip in the object coordinate system, thereby realizing the generation of object-centered multimodal data.

[0121] It is understood that the present invention can employ both offline and real-time processing during the multimodal data generation process. When real-time processing is employed, data alignment is performed using the lowest frequency acquisition device as a reference.

[0122] Example 3

[0123] Figure 3 This is a schematic diagram of the structure of a multimodal data generation device provided in the third embodiment of the present invention. The multimodal data generation device provided in the embodiment of the present invention is applicable to the case of generating multimodal data centered on an object. The multimodal data generation device can be implemented in the form of hardware and / or software, such as Figure 3 As shown, the device specifically includes: a data acquisition module 301, an object posture determination module 302, a fingertip posture determination module 303 and a multimodal data generation module 304.

[0124] in,

[0125] The data acquisition module 301 is used to acquire object image data, first joint image data, second joint posture data and force / tactile perception data;

[0126] The object pose determination module 302 is configured to determine the object pose data of the operated object in the image sensor coordinate system at any same time point based on the object image data at the same time point;

[0127] a fingertip posture determination module 303 for determining fingertip posture data of the fingertip in the object coordinate system based on the object posture data, the first joint image data, and the second joint posture data at the same time point;

[0128] The multimodal data generation module 304 is configured to generate multimodal data including object posture data, fingertip posture data, and force / tactile perception data at the same time points.

[0129] The technical solution of the embodiment of the present invention determines the object posture data of the operated object in the image sensor coordinate system through object image data, and determines the fingertip posture data of the fingertips in the object coordinate system based on the object posture data, the first joint image data and the second joint posture data, thereby realizing the generation of object-centered multimodal data, so that when the position of the operated object changes during actual operation, the multimodal data can still be used, which is suitable for scenarios where the positions of different objects change.

[0130] Optionally, the object pose determination module 302 is specifically configured to:

[0131] According to the object image data at any same time point, based on the mask image prediction model, obtain the mask image data corresponding to the object image data at the same time point;

[0132] The object image data and its corresponding mask image data are input into a pre-trained pose estimation model to obtain the object pose data of the operated object in the image sensor coordinate system at the same time point output by the model.

[0133] Optionally, the pose estimation model is trained in the following manner:

[0134] Acquire a plurality of object sample images; each of the object sample images contains the same type or different types of operated objects;

[0135] Generate a mask sample image corresponding to each of the object sample images, and use each of the object sample images and its corresponding mask sample image as a target sample image;

[0136] Labeling each of the target sample images to obtain the true pose value corresponding to each of the target sample images;

[0137] Inputting the true pose value corresponding to each of the target sample images into a pre-built pose estimation model to obtain a predicted pose output by the model;

[0138] The pose estimation model is trained according to the pose true value and the predicted pose until a preset model training end condition is met, thereby obtaining a pose estimation model that has completed training.

[0139] Optionally, the fingertip posture determination module 303 includes:

[0140] a reference posture determining unit, configured to determine reference posture data of the fingertip in the image sensor coordinate system based on the first joint image data and the second joint posture data and on a positional relationship between the first joint and the second joint;

[0141] The fingertip posture determination unit is used to determine the fingertip posture data of the fingertip in the object coordinate system according to the reference posture data and the object posture data.

[0142] Optionally, a reference pose determination unit includes:

[0143] an initial posture determining subunit, configured to determine initial posture data of the fingertip in the second joint coordinate system according to the second joint posture data;

[0144] a wristband posture determination subunit, configured to determine first joint posture data of the first joint in an image sensor coordinate system based on the first joint image data;

[0145] a relative posture determination subunit, configured to determine relative posture data between the first joint and the second joint according to a positional relationship between the first joint and the second joint;

[0146] The reference posture determination subunit is used to determine the reference posture data of the fingertip in the image sensor coordinate system according to the initial posture data, the first joint posture data and the relative posture data.

[0147] Optionally, the device further includes:

[0148] a module for determining data to be aligned, configured to, after acquiring the object image data, the first joint image data, the second joint posture data, and the force / tactile perception data, select target reference data from the object image data, the first joint image data, the second joint posture data, and the force / tactile perception data, and use other data except the target reference data as data to be aligned;

[0149] A reference time determination module, configured to obtain a reference acquisition time corresponding to the target reference data;

[0150] An adjacent time determination module is configured to determine, for any data to be aligned, a first adjacent acquisition time and a second adjacent acquisition time adjacent to the reference acquisition time among all data acquisition times of the data to be aligned;

[0151] A time alignment module is used to determine the alignment data of the data to be aligned at the reference acquisition time based on the first data to be aligned acquired at the first adjacent acquisition time and the second data to be aligned acquired at the second adjacent acquisition time, so as to obtain object image data, first joint image data, second joint posture data and force / tactile perception data at the same time points.

[0152] The multimodal data generation device provided in the embodiment of the present invention can execute the multimodal data generation method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0153] Example 4

[0154] Figure 4 A schematic diagram of the structure of an electronic device 40 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0155] like Figure 4 As shown, the electronic device 40 includes at least one processor 41 and a memory, such as a read-only memory (ROM) 42, a random access memory (RAM) 43, etc., which is communicatively connected to the at least one processor 41. The memory stores a computer program that can be executed by the at least one processor, and the processor 41 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 42 or the computer program loaded from the storage unit 48 into the random access memory (RAM) 43. Various programs and data required for the operation of the electronic device 40 can also be stored in the RAM 43. The processor 41, ROM 42, and RAM 43 are connected to each other via a bus 44. An input / output (I / O) interface 45 is also connected to the bus 44.

[0156] Multiple components in the electronic device 40 are connected to the I / O interface 45, including an input unit 46, such as a keyboard, a mouse, etc.; an output unit 47, such as various types of displays, speakers, etc.; a storage unit 48, such as a magnetic disk, an optical disk, etc.; and a communication unit 49, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 49 allows the electronic device 40 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0157] The processor 41 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 41 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 41 executes the various methods and processes described above, such as the multimodal data generation method.

[0158] In some embodiments, the multimodal data generation method can be implemented as a computer program that is tangibly contained in a computer-readable storage medium, such as storage unit 48. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 40 via ROM 42 and / or communication unit 49. When the computer program is loaded into RAM 43 and executed by processor 41, one or more steps of the multimodal data generation method described above can be performed. Alternatively, in other embodiments, processor 41 can be configured to perform the multimodal data generation method in any other suitable manner (e.g., by means of firmware).

[0159] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0160] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0161] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0162] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0163] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0164] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.

[0165] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.

[0166] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A method for generating multimodal data, characterized in that: Applied to a multimodal data generation system, the method is executed by a controller in the multimodal data generation system, comprising: Acquiring object image data, first joint image data, second joint posture data, and force / tactile perception data; Determine, based on the object image data at any same time point, the object pose data of the operated object in the image sensor coordinate system at the same time point; Determining fingertip posture data of the fingertip in the object coordinate system according to the object posture data, the first joint image data, and the second joint posture data at the same time point; Generate multimodal data including object posture data, fingertip posture data and force / tactile perception data at the same time point.

2. The method according to claim 1, characterized in that The determining, based on the object image data at any same time point, the object pose data of the operated object in the image sensor coordinate system at the same time point includes: According to the object image data at any same time point, based on the mask image prediction model, obtain the mask image data corresponding to the object image data at the same time point; The object image data and its corresponding mask image data are input into a pre-trained pose estimation model to obtain the object pose data of the operated object in the image sensor coordinate system at the same time point output by the model.

3. The method according to claim 2, characterized in that The model training method of the pose estimation model is as follows: Acquire a plurality of object sample images; each of the object sample images contains the same type or different types of operated objects; Generate a mask sample image corresponding to each of the object sample images, and use each of the object sample images and its corresponding mask sample image as a target sample image; Labeling each of the target sample images to obtain the true pose value corresponding to each of the target sample images; Inputting the true pose value corresponding to each of the target sample images into a pre-built pose estimation model to obtain a predicted pose output by the model; The pose estimation model is trained according to the pose true value and the predicted pose until a preset model training end condition is met, thereby obtaining a pose estimation model that has completed training.

4. The method according to claim 1, wherein Determining the fingertip posture data of the fingertip in the object coordinate system according to the object posture data, the first joint image data, and the second joint posture data at the same time point includes: determining reference pose data of the fingertip in the image sensor coordinate system according to the first joint image data and the second joint pose data and based on a positional relationship between the first joint and the second joint; The fingertip posture data of the fingertip in the object coordinate system is determined according to the reference posture data and the object posture data.

5. The method according to claim 4, characterized in that The determining, according to the first joint image data and the second joint posture data, reference posture data of the fingertip in the image sensor coordinate system based on a positional relationship between the first joint and the second joint, includes: Determining initial position data of the fingertip in the second joint coordinate system according to the second joint position data; determining first joint pose data of the first joint in an image sensor coordinate system according to the first joint image data; Determining relative posture data between the first joint and the second joint according to a positional relationship between the first joint and the second joint; Determine reference pose data of the fingertip in the image sensor coordinate system according to the initial pose data, the first joint pose data and the relative pose data.

6. The method according to claim 1, wherein After acquiring the object image data, the first joint image data, the second joint posture data, and the force / tactile perception data, the method further includes: Selecting target reference data from the object image data, the first joint image data, the second joint posture data, and the force / tactile perception data, and using other data except the target reference data as data to be aligned; Obtaining a reference acquisition time corresponding to the target reference data; For any data to be aligned, determining a first adjacent acquisition time and a second adjacent acquisition time adjacent to the reference acquisition time among all data acquisition times of the data to be aligned; Based on the first data to be aligned acquired at the first adjacent acquisition time and the second data to be aligned acquired at the second adjacent acquisition time, the alignment data of the data to be aligned at the reference acquisition time is determined to obtain object image data, first joint image data, second joint posture data and force / tactile perception data at each same time point.

7. A multimodal data generating device, characterized in that: Applied to a multimodal data generation system, the device is configured in a controller in the multimodal data generation system, comprising: A data acquisition module is used to acquire object image data, first joint image data, second joint posture data and force / tactile perception data; An object posture determination module is used to determine the object posture data of the operated object in the image sensor coordinate system at any same time point based on the object image data at the same time point; a fingertip posture determination module, configured to determine fingertip posture data of the fingertip in the object coordinate system based on the object posture data, the first joint image data, and the second joint posture data at the same time point; The multimodal data generation module is used to generate multimodal data including object posture data, fingertip posture data and force / tactile perception data at the same time point.

8. A multimodal data generation system, characterized in that: The system includes a controller, an image acquisition device, a motion capture glove reference component, a motion capture glove rigidly connected to the motion capture glove reference component, and a tactile module deployed at the end of the motion capture glove and connected to the fingertips; A posture sensor is deployed on the motion capture glove; The image acquisition device is used to photograph the operated object to generate object image data during the execution of the object operation task, and photograph the reference component of the motion capture glove to obtain first joint image data, and send the object image data and the first joint image data to the controller; The motion capture glove is used to collect second joint posture data through the posture sensor during the execution of the object operation task, and send the second joint posture data to the controller; The tactile module is used to generate force / tactile perception data during the execution of the object operation task, and send the force / tactile perception data to the controller; The controller is used to execute the multimodal data generation method according to any one of claims 1 to 6.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the multimodal data generation method according to any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the multimodal data generation method according to any one of claims 1 to 6 when executed.

Citation Information

Patent Citations

  • Hand shape classification method based on image processing

    CN110705465A

  • Reading method and system for cross-modal task instruction understanding

    CN116070173A

  • Mechanical arm sensing method based on multi-modal data fusion

    CN117103277A

  • Object pose labeling method, device and system of training sample

    CN120031923A

  • Action generator, action control method and storage medium stored with program for executing the same

    JP1998340354A