A multi-modal data acquisition method and system

CN122821072APending Publication Date: 2026-09-25ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611049642.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-14
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0006]本发明实施例提供一种多模态数据采集系统及方法,旨在解决现有手部采集装置因刚性器件堆叠干扰自然操作动作,以及多模态观测基础不稳定导致重建精度不足的技术问题

Benefits of technology

[0015]本说明书第五方面提供一种计算机程序产品,包括计算机程序/指令,该计算机程序/指令被处理器执行时实现第一方面所述方法的步骤。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821072A_ABST
    Figure CN122821072A_ABST
Patent Text Reader

Abstract

A multi-modal data acquisition method and system, the method is applied to a system comprising a head device worn on the head of an operator and a hand device worn on the hand of the operator, the head device comprising a first clock, a camera and a first inertial measurement unit, the hand device comprising a second clock and a plurality of second inertial measurement units, the second clock is time-synchronized with the first clock, the method comprising: the head device acquires a hand image through the camera, acquires head inertial data through the first inertial measurement unit, the hand image and the head inertial data comprise timestamps provided by the first clock; send the hand image and the head inertial data to a server; the hand device acquires hand inertial data of the hand through the plurality of second inertial measurement units, the hand inertial data comprises timestamps provided by the second clock; send the hand inertial data to the server.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments in this specification relate to the field of embodied intelligent data processing technology, specifically to a multimodal data acquisition method and system. Background Technology

[0002] With the rapid development of embodied intelligence technology, training high-quality Vision-Language-Action Models (VLAs) requires a large amount of real-world first-person perspective interaction data. This type of data typically requires the simultaneous recording of multimodal information, including the operator's environmental images, head movement trajectories, fine hand movements, and tactile feedback upon contact with objects. To acquire this data, the industry is increasingly adopting wearable head-mounted devices to replace traditional robotic teleoperation solutions. These devices, worn on different parts of the body, simultaneously capture multi-source sensory signals during natural operation.

[0003] Existing distributed wearable data acquisition solutions typically include a visual acquisition module for the chest or head and hand acquisition modules for both hands. The hand acquisition module, in order to simultaneously acquire information on finger bone posture, spatial position, and contact force, often integrates multiple rigid sensors on the back of the hand, knuckles, and palm. For example, some existing technologies use magnetic positioning principles, placing a magnetic field transmitting base station on the back of the hand and magnetic induction slave stations on each finger, inferring the spatial pose of the fingers relative to the back of the hand by calculating changes in the magnetic field. Other solutions employ exoskeletons, fixing inertial measurement units or angle sensors to a rigid frame to constrain finger movement trajectories. This concentrated stacking of rigid devices, connection structures, and wiring units on the hand surface inevitably increases the overall thickness and local rigidity of the hand device. When the operator performs fine motor actions such as grasping, pinching, and twisting, the increased mechanical resistance restricts the natural range of motion of the finger joints, causing the acquired motion data to deviate from actual human operating habits, thus affecting the accuracy and generalization ability of subsequent embodied intelligent model training data.

[0004] Furthermore, existing solutions still have room for improvement in terms of stability during multimodal observation. Solutions relying solely on inertial measurement units (IMUs) are susceptible to accumulated zero-drift errors, leading to decreased attitude estimation accuracy after prolonged operation. Solutions relying solely on visual markers are prone to target loss or positioning jumps when the hand is occluded or lighting conditions change drastically. While some existing technologies attempt to integrate multiple sensors, the lack of differentiated observation designs for the rigid and flexible joint characteristics of the hand makes it difficult to consistently provide stable spatial constraints between macroscopic hand back displacement and microscopic finger movements in complex dynamic scenarios. This incomplete observational foundation leads to data gaps or excessive noise during subsequent high-precision 3D hand reconstruction and hand-eye coordination analysis.

[0005] On the other hand, existing distributed acquisition systems rely heavily on real-time high-bandwidth wireless transmission, limiting their application in weak network environments. Traditional architectures typically require each acquisition device to aggregate raw image streams, high-frequency inertial data, and tactile signals to an edge gateway or directly upload them to the cloud in real time. In outdoor operations, industrial sites, or mobile acquisition scenarios, network bandwidth fluctuations or signal interruptions can easily lead to data transmission delays, packet loss, or even forced termination of acquisition tasks. If the sampling rate is reduced or image quality is compressed to adapt to weak network environments, the integrity and detail richness of the data will be further sacrificed, making it difficult to meet the high-fidelity training sample requirements of embodied intelligent models. Therefore, how to reduce the interference of devices with natural operations and improve the system robustness in weak network environments while ensuring the synchronization and integrity of multimodal data remains a pressing issue in the current technological field. Summary of the Invention

[0006] This invention provides a multimodal data acquisition system and method, aiming to solve the technical problems of insufficient reconstruction accuracy caused by the interference of rigid device stacking with natural operation in existing hand acquisition devices and the instability of multimodal observation foundation.

[0007] This specification provides a multimodal data acquisition method for a system including a head-mounted device worn on an operator's head and a hand-mounted device worn on the operator's hand. The head-mounted device includes a first clock, a camera, and a first inertial measurement unit (IMU). The hand-mounted device includes a second clock and a plurality of second IMUs, wherein the second clock is time-synchronized with the first clock. The method includes:

[0008] The head-mounted device acquires hand images via the camera and head inertial data via the first inertial measurement unit. The hand images and head inertial data include timestamps provided by the first clock. The hand images and head inertial data are then sent to the server.

[0009] The hand device collects hand inertial data through the plurality of second inertial measurement units, the hand inertial data including a timestamp provided by the second clock; and sends the hand inertial data to the server.

[0010] A second aspect of this specification provides a multimodal data acquisition system, comprising: a head-mounted device configured to be worn on an operator's head, integrating a first clock, a camera, and a first inertial measurement unit; and a hand-mounted device configured to be worn on the operator's hand, integrating a second clock and a second inertial measurement unit; wherein the second clock is synchronized with the first clock.

[0011] The head-mounted device is used to: acquire hand images via the camera, acquire head inertial data via the first inertial measurement unit, wherein the hand images and the head inertial data include timestamps provided by the first clock; and send the hand images and the head inertial data to the server.

[0012] The hand device is used to: collect hand inertial data of the hand through the plurality of second inertial measurement units, the hand inertial data including a timestamp provided by the second clock; and send the hand inertial data to the server.

[0013] A third aspect of this specification provides a computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to perform the method described in the first aspect.

[0014] A fourth aspect of this specification provides a computing device including a memory and a processor, wherein the memory stores executable code, and the processor executes the executable code to implement the method described in the first aspect.

[0015] This specification provides a computer program product in a fifth aspect, including a computer program / instructions that, when executed by a processor, implement the steps of the method described in the first aspect.

[0016] The technical solution provided in this invention significantly reduces the overall weight and local rigidity of the hand device by placing the IMU on the surface of the hand device to collect hand inertial data, acquiring hand images and head-head relationship data through the head device, and obtaining time-aligned multimodal data through time alignment of the hand and head devices. This structural design reduces the mechanical constraints on the operator's finger joint movements, allowing natural actions such as grasping and pinching to be performed without interference, thereby improving the accuracy of the collected data and operational comfort. Secondly, by acquiring time-aligned multimodal data as described above, a complementary multimodal observation foundation of inertial and visual modes is constructed. The visual mode provides high-precision absolute spatial position constraints, effectively correcting the drift error of the inertial measurement unit; the inertial mode maintains the continuity of the motion trajectory when visual occlusion or temporary loss of markers occurs. This complementary mechanism provides stable and complete data support for subsequent high-precision spatial reconstruction of the hand and hand-eye coordination analysis. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments in this specification, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a schematic diagram of a multimodal data acquisition system according to one embodiment of this specification;

[0019] Figure 2 This is a schematic diagram of the head device in one embodiment of this specification;

[0020] Figure 3 This is a schematic diagram of the hand device in one embodiment of this specification;

[0021] Figure 4 This is a schematic diagram of a tactile sensing layer in one embodiment of this specification;

[0022] Figure 5 This is a flowchart of a dynamic time synchronization method in one embodiment of this specification;

[0023] Figure 6 This is a flowchart of a method for calibrating user hand parameters in one embodiment of this specification;

[0024] Figure 7 This is a flowchart of a method for calibrating hand parameters in one embodiment of this specification;

[0025] Figure 8 This is a flowchart of a data acquisition method in one embodiment of this specification;

[0026] Figure 9 This is a flowchart of the method for verifying hand parameters in the embodiments of this specification;

[0027] Figure 10 This is a flowchart of a method for verifying hand parameters using a hand device in one embodiment of this specification. Detailed Implementation

[0028] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.

[0029] In one embodiment of this application, a multimodal data acquisition system is provided, which is mainly aimed at the first-person perspective data acquisition needs in embodied intelligence training scenarios. Figure 1As shown, the multimodal data acquisition system includes a head device 10, a left-hand device 20L, and a right-hand device 20R (which can be specifically implemented as a data glove). The system may also include an edge device 30 and a server 40. Figure 1 (Not shown in the image). The head-mounted device 10 is configured to be worn on the operator's head, serving as the main visual perception node and time synchronization master node for the entire system; the left-hand device 20L and right-hand device 20R are configured to be worn on the operator's left and right hands, respectively, serving as slave nodes for acquiring multimodal hand motion and tactile information. The edge device 30 is preferably a smartphone or portable tablet terminal, used to provide a human-machine interface, task management, and status management. The edge control device can run a data acquisition device control application to perform the following functions: during the acquisition phase, viewing the device's online status, synchronization status, battery status, and abnormal alarm information to assist users in on-site confirmation and task management, and executing necessary acquisition parameter configurations; during the data upload phase, controlling the device to upload task data to the cloud platform under suitable network conditions.

[0030] Server 40 is used to receive and process the uploaded data packets after the data acquisition task is completed. Each acquisition device (head device or hand device) has independent local data processing and storage capabilities, thereby reducing the dependence on real-time high-bandwidth wireless links and enabling it to adapt to long-term continuous data acquisition tasks in weak network or no network environments such as outdoor and industrial sites.

[0031] like Figure 1 As shown, the head-mounted device 10 preferably adopts a headband structure in terms of physical form. In one embodiment, a multi-camera 101 may be integrated inside the headband. Figure 1 As shown, the multi-camera system is arranged in an arc along the front of the headband to cover the main field of view where the operator's hands move naturally in front of them. In one specific implementation, the multi-camera system comprises six camera modules, four of which are visible light cameras and two are infrared cameras.

[0032] A pair of visible light cameras, C1 and C2, located in the middle of the front area of ​​the headband, form a stereo vision baseline, primarily used for environmental depth perception, hand target detection, and visual odometry calculation of head trajectory. Baseline distance. The distance can be set between 60mm and 80mm, balancing depth resolution for near-field hand operation (30cm-50cm) with the effective range for far-field environmental perception. The field of view (FOV) of the two cameras is approximately 90 degrees, with the overlapping area facing the operator's natural operating space in front of their chest, used to perform a binocular stereo matching algorithm to obtain a dense depth map.

[0033] Two infrared cameras, IR1 and IR2, located on either side, are specifically designed to track infrared marker units (such as infrared lights or IR LEDs) emitted from handheld devices, obtaining the three-dimensional spatial coordinates of the marker points through triangulation. The front of the infrared cameras integrates narrowband filters that allow only light with wavelengths of 850nm or 940nm to pass through, effectively suppressing interference from ambient visible light and ensuring stable extraction of marker point coordinates even in strong light or against complex textured backgrounds.

[0034] The two outermost visible light cameras, C3 and C4, have a wider field of view, which expands the overall environmental coverage, reduces blind spots, and captures environmental context information at the edge of the operator's field of vision, preventing target loss due to rapid head rotation. This multi-camera layout can ensure high accuracy in the central area while also meeting the needs of environmental perception over a large area.

[0035] refer to Figure 2 The schematic diagram of the head-mounted device shown illustrates that the head-mounted device 10 also includes an Inertial Measurement Unit (IMU). The IMU is installed as close as possible to the center of mass of the headband to reduce the interference of linear acceleration on attitude calculation. The acceleration, angular velocity, and other data it collects can assist visual algorithms in head attitude estimation and provide continuous pose inference in the event of rapid movement or loss of visual features.

[0036] The inertial measurement unit is preferably a nine-axis IMU, which includes: a 3-axis accelerometer for measuring the linear acceleration (including gravitational acceleration) of an object in the three orthogonal directions of X, Y, and Z; a 3-axis gyroscope for measuring the angular velocity of the object rotating around the three axes of X, Y, and Z; and a 3-axis magnetometer for measuring the magnetic field strength and direction of the surrounding environment (usually referring to the Earth's magnetic field).

[0037] The head-mounted device 10 also includes a system-on-chip (SoC, hereinafter referred to as the headband SoC), a battery, and a storage medium. In one embodiment, the head-mounted device 10 may further include an audio interaction unit, which may include a microphone and a speaker. Additionally, the head-mounted device 10 may also include buttons for inputting user commands.

[0038] As the main control chip for head-mounted devices, the headband SoC integrates the acquisition, processing, synchronization, and storage scheduling functions of multimodal sensing data onto a single silicon chip through a highly integrated system-on-a-chip design. In one embodiment, the headband SoC may include functional modules such as a computing unit, a sensor interface, a time synchronization module, a wireless communication controller, a storage management module, and a status management module.

[0039] The computing unit may include a multi-core central processing unit (CPU) for running the operating system, task scheduling, and device status management. Sensor interface connection. Figure 1 The multi-view camera shown integrates a high-precision Inter-Integrated Circuit (I2C) bus or Serial Peripheral Interface (SPI) bus for real-time reading of acceleration, angular velocity, and geomagnetic data from the inertial measurement unit. The time synchronization module may include a hardware timer and timestamp injection logic. This module can time based on the hardware timer and directly add a global timestamp with microsecond-level precision at the instant the sensor data is generated.

[0040] The wireless communication controller integrates multiple communication links to form a system communication architecture where different links work together, each undertaking a different function, rather than a single wireless link handling all data and control tasks. For example, these multiple communication links may include low-bandwidth links (such as Bluetooth Low Energy (BLE)), high-bandwidth links (Wireless Fidelity (Wi-Fi)), and low-power links (such as Sub-GHz wireless communication modules).

[0041] Among these features, low-power links can be dedicated to multi-device time synchronization links. Taking Sub-GHz modules as an example, Sub-GHz modules operate in frequency bands below 1 gigahertz (GHz) (e.g., 433 MHz, 868 MHz, or 915 MHz). Utilizing their physical characteristics of longer wavelengths, stronger diffraction capabilities, and lower penetration loss, they can be used to construct time synchronization links with handheld devices. Due to the relatively less channel interference and longer transmission distance in the Sub-GHz band, this mechanism significantly improves the determinism and coverage of time synchronization between head-mounted and handheld devices. It ensures that even under conditions of significant operator movement, facing away from the device, or non-line-of-sight (NLOS), the slave node can still stably receive the synchronization reference, thereby maintaining the microsecond-level time alignment accuracy of the multimodal data acquisition system in all scenarios. This further enhances the system's engineering practicality in weak network or high-interference environments such as industrial sites and outdoor wide-area environments.

[0042] Low-bandwidth links can be used to connect to edge devices to upload status data from the head unit, allowing users to verify and manage the data at the data acquisition site. Conversely, high-bandwidth links can be used during off-collection periods or when network conditions are favorable to establish high-bandwidth data transmission links, uploading massive amounts of raw data from local storage to the server. Furthermore, during hand parameter calibration and periodic verification, the head unit can act as a WiFi hotspot to connect with the glove device, receiving small amounts of collected data from the glove device for calibration or verification.

[0043] The storage management module integrates a high-speed Direct Memory Access (DMA) controller and interfaces with embedded Multi-Media Card (eMMC) or Universal Flash Storage (UFS). The storage management module is used for reading, compressing, and writing images to disk, as well as reading and writing audio data from the IMU to disk.

[0044] In addition, such as Figure 2 As shown, the status management module in the headband SoC is used to sense and manage the health of the system in real time. For sensor status management, the headband SoC periodically detects the heartbeat signals and register status of each sensor (camera, IMU, haptic sensor, etc.) to identify whether a sensor is offline, whether the data is abnormal, or whether it is in an overheat protection state. For disk storage status management, it monitors the write speed, remaining space capacity, and file system integrity of the storage medium in real time, triggering an alert or stopping recording when the remaining space is below a threshold. For power management, it reads the battery voltage, current, and temperature in real time through an integrated high-precision fuel gauge chip, accurately calculates the remaining power using coulomb counting, and dynamically adjusts the system power consumption strategy (such as reducing the sampling rate or shutting down unnecessary modules). For error status reporting, once any of the above anomalies (such as synchronization loss, storage failure, sensor malfunction, or low battery) are detected, the headband SoC immediately generates a status frame containing an error code and a timestamp, reporting it to the edge control device via a low-bandwidth wireless link (such as Sub-GHz or BLE) or recording it at the end of the local log. In terms of device status control, the headband SoC executes operations such as device startup, shutdown, reset, mode switching (such as switching from calibration mode to acquisition mode) and firmware upgrade according to received external instructions or internal preset strategies, thereby realizing automated management and closed-loop control of the entire life cycle of the hand device and the entire acquisition system.

[0045] The left-hand device 20L and the right-hand device 20R are basically symmetrical in hardware structure, both adopting the form of flexible data gloves. Taking the right-hand device 20R as an example, such as... Figure 1As shown, the right-hand device 20R may include a glove base 201, multiple inertial-infrared integrated modules, tactile sensors, a glove SoC, a battery, and storage media. The glove base 201 is made of a highly elastic, breathable, flexible fabric material to adapt to different users' hand shapes and reduce the feeling of restriction when wearing it.

[0046] Multiple inertial-infrared integrated modules are distributed across the back of the hand and key joints of each finger. For example, a reference module 202a is placed on the back of the hand, and joint modules 202b are placed at the base and tip of each finger. This module employs a miniaturized multi-layer PCB stacked design, with the inertial measurement unit (IMU) sub-board at the bottom and the optical emission sub-board at the top. A six-axis or nine-axis IMU chip is mounted on the IMU sub-board, and its sensitive axis coordinate system is strictly aligned with the module's mechanical coordinate system. The optical emission sub-board integrates a high-brightness infrared LED and a miniature collimating lens. The collimating lens shapes the divergent beam emitted by the LED into a conical beam with a half-angle of approximately 30 degrees, ensuring sufficient illumination intensity for long-distance identification while avoiding excessively large light spots that could cause confusion between adjacent markers. The two sub-boards are interconnected via blind vias and are encapsulated within a lightweight, high-temperature resistant engineering plastic housing. The housing surface has light-transmitting windows corresponding to the infrared LED positions, while the remaining parts are shielded to prevent internal stray light leakage.

[0047] Each integrated inertial-infrared module internally integrates an IMU and an infrared light-emitting diode (LED) on the same small circuit board, as described above, ensuring that the inertial data and the active optical marker are strictly aligned in space or have a fixed, known offset. This integrated design not only reduces the size and weight of the handheld device but also eliminates extrinsic calibration errors between multiple sensors, facilitating subsequent multimodal fusion. It is understood that the embodiments in this specification are not limited to integrating the IMU and infrared LED together; for example, the IMU and infrared LED can be configured with a fixed relative position, or they can be configured as separate units fixed by rigid connectors.

[0048] In one embodiment, a data box can be placed on the back of the hand area of ​​the data glove. Multiple infrared marker units with a preset spatial distribution are arranged on multiple surfaces of the data box, and an inertial measurement unit is built into it. The back-of-hand data box, as a rigid target of the back of the hand, can be stably observed by a headband infrared camera and used to acquire the attitude information of the back of the hand in the camera coordinate system, the headband coordinate system, and the world coordinate system. The infrared marker units distributed on the surface of the data box preferably form an asymmetric three-dimensional distribution to improve recognizability from different viewpoints and provide stable constraints for subsequent geometric pose solving based on 2D-3D correspondence.

[0049] Tactile sensors are located on the inner side of the fingertips and the palm area of ​​the glove to capture pressure distribution information during hand-object interaction. Preferably, the tactile sensors employ a flexible, layered structure to form the tactile sensing layer, balancing flexibility, coverage area, and pressure distribution detection capability. Specifically, Figure 4 This is a schematic diagram of the tactile sensing layer in the embodiments of this specification. Figure 4 The location of the tactile sensor is indicated by a blue curve in the lower right corner. (Example) Figure 4 As shown, the tactile sensing layer can be formed by stacking at least three layers of flexible materials. The middle layer is a pressure-sensitive variable resistance fiber fabric layer, whose local resistance changes when subjected to pressing, squeezing, or contact force. Horizontal and vertical conductors are respectively provided in the upper and lower fabric layers located on the upper and lower sides of the middle layer. These horizontal and vertical conductors are preferably formed of flexible conductive material and distributed at predetermined intervals in the palm area. When the three layers are stacked, the horizontal and vertical conductors in the upper and lower layers can form multiple cross-sensing units through the middle variable resistance fiber fabric layer. During actual data acquisition, the system can poll or matrix-scan the horizontal and vertical conductors to obtain resistance changes or corresponding ADC sampling value changes in different cross-regions. Since different regions cause local resistance changes in the middle variable resistance fiber fabric layer under stress, the system can obtain contact intensity distribution information in different areas of the palm. Furthermore, a mapping relationship between resistance change or ADC change and contact pressure can be established through pre-calibration, thereby converting the original electrical signal into force information or pressure distribution information. Through the aforementioned layered flexible structure, the palm tactile sensing unit can achieve multi-point pressure sensing of the palm contact area without significantly increasing the thickness and rigidity of the glove. This is beneficial for maintaining natural hand grip and operation, as well as for acquiring richer hand-object interaction tactile data.

[0050] It is understandable that, in addition to the variable resistance fiber fabric scheme based on the piezoresistive effect, capacitive, piezoelectric, or optical tactile sensing mechanisms can also be used to implement tactile sensors, and there is no limitation on this.

[0051] Similar to headband SoCs, glove SoCs may include functional modules such as computing units, sensor interfaces, time synchronization modules, wireless communication controllers, storage management modules, and state management modules. Figure 3As shown, the glove SoC can perform the following functions: reading or writing IMU / haptic data to a disk; controlling infrared LEDs; synchronizing time with the headband as a synchronization slave; collaborating across multiple communication links; and managing device status. These multiple communication links include, for example, WiFi, Sub-GHz communication links, and Bluetooth. The hand device can act as a slave, synchronizing time with the headband device via a Sub-GHz communication link; uploading status information to a mobile phone via Bluetooth; acting as a client connecting to the headband's WiFi hotspot, uploading a small amount of collected data to the headband as verification data; and uploading the full amount of collected data to the server via WiFi connection.

[0052] In the embodiments described in this specification, the edge device 30 does not directly participate in the real-time reception of high-bandwidth raw data, but rather serves as the entry point for task management. Operators can start or stop acquisition tasks, view the battery level and storage status of each node, perform device pairing, and trigger self-test processes through the application on the edge device 30. The edge device 30 also displays a visual guidance animation to prompt the user to perform specific hand gestures to complete the calibration of the user's hand parameters. After all acquisition tasks are completed, the server 40 receives data packets uploaded by each acquisition device (including head and hand devices) via a wired or high-speed wireless network, and performs offline high-precision 3D reconstruction, multimodal data alignment, and training sample generation. Since the raw data has already been recorded at a high quality on the edge, the cloud processing process does not need to worry about data packet loss caused by network fluctuations, thus ensuring the integrity of the training data.

[0053] In terms of power management, both the head-mounted device 10 and each hand-held device have built-in independent rechargeable lithium battery packs and are equipped with efficient power management units (PMUs). Considering the need for long-term data acquisition, the circuit design of each node incorporates multi-level power management modes. Under normal acquisition conditions, all sensors and the main control chip operate at full speed; in standby or paused states, high-frequency sensors enter sleep mode, with only the communication module listening for wake-up signals; when the battery level falls below a preset threshold, the system automatically triggers a protection mechanism, prioritizing the saving of currently acquired data and safely shutting down to prevent file system corruption or data loss due to sudden power outages. The battery compartment design fully considers wearing comfort; the head-mounted device's batteries are evenly distributed on the back of the headband to balance the center of gravity, while the hand-held devices' batteries are concentrated in the data box on the back of the hand, avoiding adding extra weight to the finger joints and affecting dexterity.

[0054] In one embodiment of this application, high-precision time synchronization among multiple devices is a crucial prerequisite for achieving multimodal data fusion. Since the head device 10, the left-hand device 20L, and the right-hand device 20R each have independent clock sources, their local clock frequencies exhibit slight differences due to temperature variations, aging, and manufacturing tolerances. Long-term operation leads to accumulated timestamp discrepancies. Directly using unsynchronized timestamps for data alignment will result in significant spatiotemporal misalignment between the visual image and the inertial attitude in high-speed motion scenarios, thereby compromising the accuracy of 3D reconstruction.

[0055] to this end, Figure 5 This diagram illustrates a dynamic time synchronization method using a master-slave architecture as described in an embodiment of this specification. The head device 10 serves as the time master node, and its internal first clock... It is considered the global standard time; the left and right hand devices act as time slave nodes, with their internal second clocks. It needs to periodically move closer to the first clock.

[0056] like Figure 5 As shown, in step S501, the head device 10 obtains the current time. .

[0057] The head device 10 can start from its first clock at preset time intervals (e.g., every 100ms or 200ms). Get the current time .

[0058] In step S503, the head device 10 sends a time-related message to the hand device 20. Synchronization signal packets.

[0059] The time synchronization process is performed via the dedicated low-latency wireless link (e.g., a Sub-GHz link). Header device 10 acquires the current time each time... Then, a synchronization signal packet is generated and broadcast. The data payload of this synchronization signal packet contains the transmission time. The precise value.

[0060] In step S505, the hand device 20 calculates the receiving time. and Time deviation .

[0061] The handheld device 20L or 20R immediately records the moment of reception upon receiving the synchronization signal packet. This moment is determined by the current second clock. Provided. Subsequently, the handheld device calculates the raw time deviation. .

[0062] In step S507, the hand device 20 is based on the time deviation Adjust the clock.

[0063] In one implementation, It is not a true clock offset because it includes the propagation delay of the wireless signal in the air. And the fixed processing delay introduced by the receiver hardware demodulation and interrupt response. Therefore, the actual clock offset It can be calculated using the following formula:

[0064]

[0065] in, Depending on the physical distance between devices, in close-range wearable scenarios (usually less than 1 meter), its value is approximately 3.3 ns, which can be ignored or set as a constant relative to the microsecond-level synchronization accuracy requirements; These are inherent system parameters, which can be precisely measured via a wired connection during the factory calibration phase and then embedded into the handheld device's firmware. Through the above calculations, the handheld device obtains an estimate of its clock offset relative to the master node at the current moment.

[0066] To overcome the random noise interference that may occur in a single measurement (such as instantaneous jitter caused by multipath effects in the wireless channel), this embodiment employs a sliding window filtering algorithm to filter the data obtained from multiple consecutive measurements. The sequence is processed. Let the first... The deviation obtained from the measurement is The smoothing deviation after filtering It can be represented as:

[0067]

[0068] In the formula, This is a smoothing coefficient, ranging from 0 to 1, and is typically set to 0.1 to 0.3. Smaller values... A higher value can effectively suppress high-frequency noise, but it will reduce the tracking speed for clock frequency drift; a larger value... The higher the value, the faster the response but the weaker the noise immunity. The system can dynamically adjust based on the current wireless link quality. Value: Increase appropriately when the Signal Strength Indicator (RSSI) is stable and the bit error rate is low. To converge quickly; otherwise, to decrease. In order to maintain stability.

[0069] In one implementation, a smoothed clock skew is obtained. Subsequently, the handheld device does not directly modify the local clock count, as a time-lapse could potentially cause confusion in upper-layer application logic. Instead, the system uses a software phase-locked loop (PLL) to dynamically adjust the second clock. The counting frequency. Specifically, if This indicates that the node clock is lagging, requiring a slight increase in the interrupt frequency of the local timer or the insertion of an additional counting pulse within a cycle; if This indicates that the node clock is ahead, requiring a slight reduction in the counting frequency or skipping a few counting pulses. Frequency adjustment amount. Proportional to the amount of deviation:

[0070]

[0071] in, and These are the proportional gain and integral gain coefficients, respectively. Through this closed-loop control, the second clock... The frequency gradually locks to the first clock. The frequency of these frequencies ensures that the relative error between the two remains within the microsecond or even nanosecond range. The timestamp, after synchronization correction, is directly embedded into the metadata header of each frame of sensor data (including IMU data, tactile data, and infrared status codes), ensuring that all locally recorded data carries a unified global time reference.

[0072] In addition to radio time synchronization, this embodiment also introduces a timing-encoded auxiliary synchronization mechanism based on active infrared marker units to further improve the alignment accuracy of visual and inertial data at the sub-frame level. Multiple infrared marker units on the handheld device not only serve as spatial positioning feature points but also function as time beacons. The handheld main control chip controls these infrared LEDs to flash at high frequency according to a preset timing-encoded pattern. This encoding pattern uses a combination of pulse width modulation (PWM) and time slot encoding: each synchronization cycle is divided into several time slots, and the combination of LED on / off states represents specific binary encoded information, including the current frame number and the time difference increment from the most recent radio synchronization time. .

[0073] Infrared cameras IR1 and IR2 in the head-mounted device 10 acquire images at a fixed high frame rate (e.g., 120fps or 240fps). During image preprocessing, the system not only extracts the two-dimensional coordinates of the light spot but also decodes the brightness change sequence of the light spot to reconstruct the embedded time information. Due to the extremely high speed of light, the transmission delay of the infrared light signal is negligible; therefore, the state of the light spot captured in the image directly reflects the local time of the hand-mounted device at that moment. By comparing the decoded time information with the midpoint of the head-mounted camera's exposure time, the system can calculate the subtle phase difference between visual acquisition and inertial acquisition. Even if the radio synchronization link experiences a brief packet loss due to sudden interference, the infrared visual channel can still provide high-frequency relative time verification, preventing the time drift from rapidly expanding in a short period. This dual synchronization mechanism (low-frequency radio calibration frequency + high-frequency optical signal phase verification) constitutes a robust time synchronization system.

[0074] Through the aforementioned time alignment process, each record ultimately stored in the local file system—whether a video frame, inertial vector, or tactile pressure map—has a precisely corresponding time label. This provides a solid temporal foundation for the multimodal fusion algorithm in the subsequent offline processing stage, enabling the algorithm to accurately correlate the visually observed hand position with the inertial-calculated posture, thereby effectively solving the failure problem of single modality under occlusion or rapid movement.

[0075] In practical engineering applications, each inertial measurement unit (IMU) inherently exhibits axial misalignment during manufacturing, and mechanical installation tolerances are unavoidably introduced during assembly onto the carrier. This results in an unknown static rotational deviation between the body coordinate system of each sensor and the preset theoretical carrier coordinate system. Without compensation for this deviation, systematic errors will occur during multi-sensor data fusion. To address this, a motion alignment process can be performed on the multiple IMUs within the handheld device.

[0076] During motion alignment, the user of the hand device maintains a rigid hand posture and performs slow, multi-directional rotation in space for three to five seconds to record sensor data. Using the IMU on the back of the hand as the base node and the IMUs at the fingers as child nodes, the relative attitude matrix of each child node relative to the base node is calculated based on the recorded attitude data. The expression for this relative attitude matrix changing over time is as follows:

[0077] R rel (t)=R base (t)⁻¹×R node (t)

[0078] =[R init_base ×R body_base (t)]⁻¹×[R init_node ×R body_node (t)]

[0079] =R body_base (t)⁻¹×R init_base ⁻¹×R init_node ×R body_node (t),

[0080] Where R init_base R represents the attitude matrix at the initial moment of the base node. body_base (t) represents the rotation matrix of the base node at time t relative to its initial orientation, R init_node R represents the pose matrix of the child node at the initial moment. body_node (t) represents the rotation matrix of the child node at time t relative to its initial orientation, and the intermediate term M=R in the formula is... init_base ⁻¹×R init_node This represents the difference in reference frames fixed under ideal conditions; the system aims to find an optimal alignment rotation matrix M=R. aligner Any solution used to solve this R aligner All methods described herein are protected within the scope of this patent.

[0081] In practice, the alignment parameters are solved by constructing a nonlinear optimization problem. The optimization variables include the alignment rotation matrix so3, which is parameterized using the angular axis. aligner ∈ℝ³ and the average relative rotation matrix so3 expressed in angular axis parameterization mean ∈ℝ³. For each sampling time t, first calculate the transformed relative attitude R. rel (t)=R body_base (t)⁻¹×R aligner ×R body_node (t), and then calculate the residual vector residual(t) = AngleAxis(R). mean_rel ⁻¹×R rel (t))∈ℝ³, where the AngleAxis function represents the operation of converting a rotation matrix into an angular axis vector, R mean_rel For SO3 mean The reconstructed average relative rotation matrix is ​​optimized by minimizing the sum of squared L2 norms of the residual vectors at all sampling times, i.e., minΣ. t The optimal R for each child node is obtained by solving ||residual(t)||² using an iterative optimization algorithm. aligner This completes the motion alignment.

[0082] In addition, a confidence scoring mechanism can be established to assess the reliability of the alignment results. This score is based on the product of the root mean square residual cost and the motion amplitude, where the motion amplitude is defined as the cumulative rotation angle during alignment. The wider the motion coverage during alignment and the smaller the final root mean square residual cost, the higher the reliability of the result. The root mean square residual angle value (RMS) can be calculated first based on the optimized total cost. residual_deg =√(2×mean_cost_per_node / N_samples)×180 / π, where mean_cost_per_node represents the average optimization cost per node, and N_samples represents the total number of samples involved in the calculation. The ranking score is then calculated using the exponential decay function.

[0083] ranking=clamp(100.0×exp(-RMS_residual_deg×0.3),0,100),

[0084] The clamp function is used to limit the value to the range of zero to one hundred. This score directly reflects the alignment quality. The higher the score, the more accurate the relative attitude calibration between the sensors and the more sufficient the motion excitation. The system automatically filters valid alignment data or prompts the user to re-execute the alignment action based on the score to ensure the accuracy of subsequent attitude calculation.

[0085] Figure 6 This is a flowchart illustrating the method for calibrating user hand parameters in an embodiment of this specification. In one embodiment, the calibration mode is a crucial initialization stage before the system formally begins data acquisition. Its core objective is to establish an individualized hand skeleton model adapted to the current user's hand shape characteristics, wearing position, and flexible deformation state, and to establish a precise mapping relationship between multi-sensor nodes and anatomical key points. Unlike traditional methods that rely on manual measurement input or dedicated calibration fixtures, this embodiment employs an automated modeling strategy based on multimodal data fusion. It utilizes the semantic understanding capabilities of visible light vision, the geometric constraints of infrared active marking, the relative attitude deduction capabilities of the inertial measurement unit, and the contact event detection capabilities of the tactile unit to complete high-precision calibration without the need for external auxiliary equipment.

[0086] In step S601, the edge device 30 sends instructions to the head device 10 and the hand device 20 (including the left hand device and / or the right hand device) to indicate the start of calibration.

[0087] Once the user has put on the head device 10 and hand device 20, the calibration process is triggered by the application on the edge device 30. The edge device 30 then sends a start command to each acquisition device through the wireless communication link, so that the system switches from the standby state to the calibration guidance state, laying the control foundation for subsequent multi-sensor synchronous acquisition and model initialization.

[0088] In step S603, the edge device 30 guides the wearer to perform preset actions.

[0089] The display screen of the edge device 30 presents a visually guided animation, instructing the user to move both hands to a preset field of view approximately 40cm to 60cm in front of the head device 10, maintaining the initial posture of open palms and naturally extended fingers. The visible light cameras C1-C4 in the head device 10 monitor in real time whether the hands are completely within the overlapping field of view. If fingers are detected to be outside the field of view or obstructed, the interface immediately prompts the user to adjust the position or angle of the hands until the observation conditions required for calibration are met.

[0090] Subsequently, the guidance interface further prompts the user to perform a series of finger-touching actions, such as the tip of the thumb touching the tips of the index, middle, ring, and little fingers in sequence, and the fingertips of both fingers lightly touching each other. These preset actions provide the necessary physical constraints for subsequent tactile verification and skeletal parameter optimization.

[0091] In step S605, the hand device 20 controls the infrared marking unit to turn on.

[0092] Upon receiving the instruction to begin calibration, the hand device 20 immediately controls the activation of multiple infrared marker units deployed on the back of the hand and at each finger joint, in order to cooperate with the head device in acquiring infrared marker images.

[0093] In one embodiment, the hand device can activate each infrared marker unit according to a preset code and a preset timing coding pattern, thereby controlling multiple infrared marker units to flash periodically according to the preset timing coding pattern. Simultaneously, the hand device records multiple coded states corresponding to the multiple infrared marker units as marker state data. These multiple coded states are timestamped by a second clock. The timestamped states are precisely marked by the second clock of the hand device 20 and share the same time reference with subsequently acquired inertial data, thus ensuring strict alignment between optical and inertial observations on the time axis. In this way, the infrared marker units not only provide spatial positioning features, but their flashing sequence also carries identification coding information, enabling the head device 10 to accurately distinguish the affiliation of each infrared marker unit in the case of densely distributed multiple marker points, providing geometric constraint input for subsequent hand parameter calibration.

[0094] In step S607, the head device 10 acquires hand image data through a camera.

[0095] Once the hand positions are confirmed and the observation conditions are met, the visible light camera and infrared camera in the head-mounted device 10 simultaneously begin acquisition. The visible light camera captures RGB images of the hands, while the infrared camera simultaneously captures infrared marker images of the active infrared light spots emitted by each inertial-infrared integrated module on the hand-mounted device. In one embodiment, both the RGB image and the infrared marker image include time-synchronized timestamps.

[0096] In the embodiments described in this specification, because infrared LEDs have unique timing coding characteristics, the system can easily distinguish the marking points at different fingers and different joints, avoiding the problem of easy confusion in traditional passive marking.

[0097] In step S609, the hand device 20 collects hand inertial data through each IMU.

[0098] Each inertial-infrared integrated module's IMU outputs real-time triaxial acceleration, triaxial angular velocity, and triaxial magnetic force data as hand inertial data. In one embodiment, the hand inertial data includes a time-synchronized timestamp.

[0099] In step S611, the hand device 20 collects tactile data through a tactile sensor.

[0100] As the user performs a preset finger-tapping action according to the instructions of the edge device, the tactile sensor collects pressure distribution signals in the palm and fingertip areas at high frequency as tactile data. In one embodiment, the tactile data includes a time-synchronized timestamp.

[0101] In step S613, the hand device 20 sends hand inertial data and tactile data to the head device 10.

[0102] The hand device 20 can transmit hand inertial data and tactile data with a unified timestamp to the head device 10 via a wireless communication link, ensuring strict alignment of multimodal data in the time dimension. In one embodiment, the head device 10 can provide a Wi-Fi hotspot, upon which the hand device 20 establishes a high-bandwidth communication connection with the head device 10. Specifically, the system-on-chip (SoC) inside the head device 10 integrates a Wireless Fidelity (Wi-Fi) communication controller. When the system enters calibration mode or receives a verification command from the edge device 30, the headband SoC automatically activates the Soft Access Point (SoftAP) function, broadcasting a beacon frame containing a preset Service Set Identifier (SSID) and a security key. This SSID typically contains a device serial number or task identifier, enabling the hand device 20 to quickly identify and distinguish the target network from other environmental noise during scanning. The glove SoC of the hand device 20 has a built-in Wi-Fi client module. After the calibration process starts, it actively scans for surrounding wireless signals and automatically initiates an association request with the head device 10's hotspot based on pre-stored pairing information or connection credentials forwarded by the edge device 30. After both parties complete a four-way handshake authentication, an encrypted data transmission channel is established. At this point, the hand device 20 obtains an Internet Protocol (IP) address dynamically assigned by the head device 10, forming a point-to-point local area network topology. Because the calibration process has high requirements for data integrity and real-time performance, this Wi-Fi connection uses User Datagram Protocol (UDP) or Transmission Control Protocol (TCP) to encapsulate data packets, and adds a sequence number verification and retransmission mechanism at the application layer to ensure the reliability of data transmission in close-range wear scenarios.

[0103] In one embodiment, the hand device 20 also sends the marking status data corresponding to the infrared marking unit to the head device 10.

[0104] In step S615, the head device 10 calibrates the hand parameters based on hand image data, inertial data, and tactile data.

[0105] Figure 7 This is a flowchart of a method for calibrating hand parameters in one embodiment of this specification.

[0106] like Figure 7As shown, in step S701, the positions of key points of the hand are determined based on the RGB image.

[0107] In one implementation, the coordinates of 21 standard key points, including the wrist, metacarpophalangeal joints, proximal interphalangeal joints, distal interphalangeal joints, and fingertips, can be extracted using a two-dimensional hand key point detection network.

[0108] Because of the use of a multi-camera array, the system utilizes epipolar constraints to match identical feature points across different camera views and calculates the 3D positions of these key points in the camera coordinate system using a triangulation algorithm. ,in This represents the keypoint index. Epipolar constraints are the core geometric principle in multi-view stereo vision systems used to establish matching relationships between corresponding points in different camera views. Essentially, it describes the projection characteristics of the plane (i.e., the epipolar plane) formed by the optical centers of two cameras and the same 3D point in space onto their respective imaging planes. Specifically, when multiple visible light cameras in the head-mounted device 10 simultaneously observe the hand, any keypoint of the hand in space (such as a fingertip or joint) appears as a pixel in the left camera image. The corresponding matching point of this point in the right camera image must lie on a specific straight line, which is the epipolar line. The position and direction of the epipolar line are entirely determined by the relative pose (rotation matrix and translation vector) between the two cameras and the camera's intrinsic parameter matrix. Triangulation algorithms are the core computational method in multi-view stereo vision systems for recovering 3D spatial coordinates from 2D image observations based on geometric principles. The basic idea is to utilize the geometric property that the projection rays of multiple cameras onto the same spatial point should ideally intersect at a single point, and to estimate the 3D position of that point by solving for the optimal intersection point.

[0109] This process generates an initial shape of the hand skeleton based on visual semantics, providing a global scale reference for the hand structure, but its accuracy is limited by image resolution and occlusion.

[0110] In step S703, the position of the infrared marker unit is determined based on the infrared marker image.

[0111] Specifically, by calculating the pixel coordinates of the light spot in the left and right infrared images and combining this with known camera extrinsic parameters, the absolute three-dimensional position of the center point of each inertial module in the camera coordinate system can be recovered. ,in Corresponding module number. Because the inertial module is rigidly (or quasi-rigidly) connected to key parts of the finger bones. It provides high signal-to-noise ratio spatial observations of bone nodes, which can serve as strong constraints for correcting visual skeleton drift.

[0112] In one embodiment, the head device 10 may first determine the affiliation of each infrared marker unit based on the aforementioned marker status data, and then determine the position of each infrared marker unit based on the infrared marker image and the affiliation of each infrared marker unit.

[0113] In step S705, the relative position of the IMU is determined based on the inertial data.

[0114] Specifically, the pitch and roll angles of each IMU can be calculated using the triaxial acceleration and triaxial angular velocity data collected by each IMU, and the heading angle can be estimated by combining the magnetometer data (if configured), thereby obtaining the rotation matrix of each IMU relative to the world coordinate system. By multiplying the inverse of the rotation matrix of the current IMU with the rotation matrix of the adjacent IMU, the relative rotation relationship between adjacent IMUs can be obtained, thereby revealing the spatial orientation change between adjacent phalanges.

[0115] Subsequently, an initial kinematic chain model can be constructed based on a preset default human hand bone length parameter library (containing average statistical values ​​for adult men, women, and children). The preset default finger bone length parameter plays a crucial role in initial scale anchoring.

[0116] By relative rotation relationship Combined with the default bone length, the spatial extension direction and endpoint position of each phalanx are calculated through forward kinematics recursion, thereby estimating the relative positional distribution of key finger points. The advantage of this path is that it is unaffected by lighting and occlusion, and can provide continuous attitude changes.

[0117] In step S707, the key point positions, infrared marker unit positions, and IMU relative positions are fused to generate hand parameters.

[0118] In one embodiment, the relative positional distribution of the aforementioned key finger points Furthermore, the relative position information between IMUs observed through the visual path can be introduced as an external geometric constraint. Since visual triangulation or infrared marker localization can provide the absolute three-dimensional coordinates of the center point of the inertial-infrared integrated module in the camera coordinate system, it implicitly contains real physical distance information. The system compares the relative pose calculated by the inertial path with the relative position observed by the visual path to construct an optimization objective function that includes the bone length parameter. When there is a residual between the inertial calculated fingertip spacing and the visually observed module spacing, the optimization algorithm adjusts the bone length parameter to minimize the residual, thereby gradually correcting the default prior length to an individualized parameter that adapts to the current user's actual hand shape.

[0119] In another implementation, a unified multimodal joint optimization framework can be constructed, which integrates key point positions, infrared marker unit positions, and IMU relative poses as inputs for solving hand skeleton parameters. Its core lies in establishing a mathematical mapping relationship between each observation data and the parameters of the model to be optimized and minimizing the comprehensive residual.

[0120] In step S709, the contact detection result is determined based on the tactile data.

[0121] For tactile data, a dynamic threshold can be set. When the pressure signals of two fingertip areas (such as the tip of the thumb and the tip of the index finger) simultaneously exceed the threshold and the duration is greater than a preset window (such as 50ms), a valid contact event is determined to have occurred. .

[0122] In step S711, the hand parameters are optimized based on the contact detection results.

[0123] The aforementioned contact events can be compared with the hand skeleton results generated based on visible light, infrared, and inertial information. Specifically, if the skeleton results obtained by fusing visual, infrared, and inertial information indicate that the corresponding fingers should be in contact or close contact, but the tactile channel does not detect the corresponding contact signal, or if the tactile signal indicates that contact has occurred but the skeleton results do not reflect the corresponding finger-to-finger relationship, it can be determined that there is a deviation in the current tactile channel mapping relationship, threshold setting, or local skeleton estimation, and further correction or calibration parameter adjustment can be triggered.

[0124] Based on this, key hand point information obtained from RGB images, absolute position observation information of inertial nodes based on infrared markers, finger skeleton estimation results based on inertial and default bone lengths, and contact events obtained from tactile data can be jointly fused. During this fusion process, geometric consistency, node correspondence, structural constraints, and tactile contact constraints among multiple types of information can be utilized to optimize the current user's finger bone length, joint connection relationships, the position of inertial nodes in the skeleton model, and the correspondence between tactile channels and key skeleton regions. Through this process, a finger skeleton model adapted to the current user's wearing state can be established, and further, a spatial mapping relationship between active infrared markers, inertial nodes, tactile regions, and the hand skeleton model can be established.

[0125] Specifically, a nonlinear optimization problem can be constructed to minimize the residuals among multiple observation data. The cost function is defined. as follows:

[0126]

[0127] In the formula, This represents the length vector of each phalanx bone to be optimized. Represents the joint angle state. This represents the installation offset of the inertial module relative to the skeleton; and These are the visual keypoint prediction function and the module position prediction function derived from the kinematic model, respectively. The fingertip spacing calculated based on the current model. The ideal spacing in the contact state (theoretically close to 0); This is an indicator function that only takes effect when a tactile touch event is detected; These are the weighting coefficients for each mode. Wherein, the above can be considered as... The corresponding joint angle is used as the initial estimate of θ, and the default bone length parameter is used as the initial value of L, together forming the initial values ​​of the optimization algorithm.

[0128] In this context, tactile data serves as a "truth check." If the skeletal model obtained through the fusion of vision and infrared sensors shows a distance of 15mm between the tips of the thumb and index finger (no contact), but the tactile unit clearly reports a strong contact signal, it indicates that the current bone length estimate is too long or there is a deviation in the joint zero point. In this case, the optimization algorithm will significantly increase the penalty weight for this item, forcing the model parameters... and The system iterates in the direction of shortening the fingertip distance until the calculated spacing matches the contact state perceived by touch. Conversely, if the model shows contact but there is no tactile signal, it may mean that the tactile sensitivity threshold is set too high or the model is over-compressed, and the system will adjust the parameters accordingly. This mechanism, which uses tactile events as hard constraints, effectively solves the problem of the accumulation of tiny-scale errors that are difficult to detect in purely visual or purely inertial solutions.

[0129] The optimization solution can be obtained iteratively using the Levenberg-Marquardt algorithm. Initially, the system loads default skeletal parameters; in each iteration, the algorithm calculates the reprojection error, geometric position error, and haptic constraint error under the current parameters, and updates the parameter vector. Typically, after 10 to 20 iterations, the cost function... It converges to a stable value, at which point we obtain , and These are individualized model parameters adapted to the user. These parameters not only include precise knuckle length but also implicitly include sensor offset compensation for the tightness of the gloves.

[0130] After calibration, the head-mounted device can score the calibrated hand parameters using a preset algorithm, and determine whether recalibration is necessary based on the score. If the tactile and visual / inertial data of a finger-tapping action conflict significantly and cannot be eliminated through optimization (e.g., the user did not perform the action correctly or the sensor malfunctions), the system will mark that channel as abnormal and guide the user to re-perform the specific action or check the device's fit. Through this closed-loop automated calibration process, this embodiment can complete high-precision personalized modeling within tens of seconds, significantly reducing the barrier to entry for embodied intelligent data acquisition and ensuring the anatomical accuracy of subsequent data collection.

[0131] After the head device confirms that the calibrated hand parameters are valid, it can send the calibrated hand parameters to the server 40 so that the server can estimate the hand posture based on the collected data and the calibrated hand parameters.

[0132] Figure 8 This is a flowchart illustrating a data acquisition method in one embodiment of this specification. In one embodiment of the invention, [the method is described in detail]. Figure 8 The data acquisition and processing process shown may specifically include the following steps.

[0133] Step S801: The head device 10 acquires a hand image, including a timestamp, via a camera.

[0134] In this step, the head device 10, acting as a first-person visual perception node, continuously captures visual information about the operator's environment and hands using the visible light camera in its integrated multi-view camera system. To ensure accurate alignment of the visual data with other modal data in subsequent processing, the camera sensor directly injects a timestamp provided by a first clock via a hardware timer at the moment each frame exposure ends or readout is completed. This hardware-level stamping mechanism avoids time stamping errors caused by operating system task scheduling jitter, ensuring that the timestamps of the hand images accurately reflect the physical moment when photons arrive at the sensor plane. In a different embodiment, an event camera can be used instead of a traditional frame camera to acquire hand images. The event camera responds asynchronously to pixel-level brightness changes and outputs an event stream with precise timestamps. This approach provides higher temporal resolution and dynamic range in high-speed motion scenarios, thereby providing richer temporal details for subsequent pose estimation.

[0135] by Figure 1Taking the headband shown as an example, the four visible light cameras C1, C2, C3, and C4 located in the middle and outer sides of the front of the headband are mainly responsible for acquiring broadband spectral texture images of the environment in front of the operator and their hands. Among them, C1 and C2 form a stereo vision baseline to provide dense depth perception, while C3 and C4 expand the field of view coverage to capture edge context. At the end of each frame exposure, these visible light cameras are directly injected with a microsecond-level timestamp provided by the first clock by the hardware timer in the on-chip system, generating a first-view visible light image sequence containing rich semantic information and natural texture features, providing basic data support for subsequent hand detection, key point recognition, and environmental understanding.

[0136] In one embodiment, the head device 10 also utilizes two infrared cameras IR1 and IR2 located at specific positions on the front side of the headband to acquire narrowband infrared light signals emitted by the active infrared marking unit on the hand device 20. The narrowband filter integrated at the front end of the infrared camera effectively suppresses interference from ambient visible light, making the acquired hand image mainly appear as a high-contrast discrete light spot distribution rather than continuous texture; the infrared camera is also stamped with a precise timestamp originating from the same first clock at the moment the exposure ends, thereby ensuring strict alignment of infrared light spot observation and visible light texture observation on the time axis. This dual-channel parallel acquisition mechanism allows the concept of hand image in this embodiment to encompass two complementary types of data streams: visible light images provide natural scene representations unaffected by artificial marking, suitable for initial pose estimation and semantic segmentation based on learning methods; while infrared images provide a geometrically constrained point set resistant to illumination changes and background obfuscation, suitable for high-precision rigid body pose calculation and drift correction.

[0137] Step S803: The hand device 20 acquires hand inertial data, including timestamps, via the IMU.

[0138] Multiple IMUs integrated within the hand device 20 continuously output triaxial acceleration, triaxial angular velocity, and optional triaxial magnetometer data at high-frequency sampling rates (e.g., 200Hz or higher). The second clock in the hand device 20 is highly synchronized with the first clock in the head device 10, so each IMU data sample carries a timestamp based on a unified time reference. This hand inertial data reflects the instantaneous motion state of each joint in three-dimensional space and is the core input source for continuous deduction of finger posture.

[0139] Because IMU data is generated at a much higher frequency than image frame rates and is sensitive to transmission delays, the hand device 20 can temporarily store this data locally using a circular buffer and retain the original microsecond-level timestamps during packaging to prevent loss of time information due to communication congestion. In one different embodiment, a miniature magnetic positioning system can be used to replace or partially assist the IMU in acquiring hand motion data. By placing a magnetic field emitter in the palm and arranging magnetic induction coils at the fingertips, induced voltage signals can be acquired. These induced voltage signals can be used to calculate the six degrees of freedom pose of the fingers relative to the palm. Although this method is susceptible to interference from environmental ferromagnetic materials, it can provide drift-free absolute position observations in unobstructed and electromagnetically pure environments, thus complementing the IMU data.

[0140] In one embodiment, the hand device 20 can also collect tactile data through a tactile sensor to assist in determining the hand posture.

[0141] Step S805: Head device 10 acquires head inertial data, including timestamps, via IMU.

[0142] In addition to handling visual acquisition, the head-mounted device 10 also senses the six degrees of freedom of head motion in real time through its integrated first inertial measurement unit (IMU). The head inertial data is also timestamped by a first clock, ensuring a strict correspondence between it and the hand images acquired by the same device on the time axis. The key to this design is that the head inertial data not only serves as an auxiliary input for head pose estimation but also acts as a bridge to transform local hand observations into the world coordinate system. When visual features are missing or the camera experiences rapid motion blur, the head inertial data maintains the continuity of the head trajectory, thus ensuring that the hand image's positioning in the global coordinate system does not jump. In practical engineering deployments, considering the differences in dynamic characteristics between head and hand movements, the sampling rate and filtering parameters of the head IMU can be configured independently of the hand IMU to adapt to relatively stable head movements but with drastic changes in the field of view.

[0143] While data is being acquired, a status management thread runs in the background. This thread polls the health status of each key module at a low frequency (e.g., 1Hz), including remaining battery power (SOC), remaining storage percentage, sensor temperature, wireless link signal strength (RSSI), and time synchronization deviation. This status information is packaged into short status frames and periodically sent to the edge device 30 via a low-bandwidth status management link. Users at the acquisition site can view the device's operation in real time on their mobile phone screen. If a device's battery is found to be too low or the synchronization deviation exceeds the allowable range (e.g., more than 1ms), the system will immediately issue an audible and visual alarm, prompting the user to pause the task and intervene. This real-time closed-loop feedback mechanism greatly reduces the risk of losing the entire valuable acquisition data due to device failure.

[0144] After collecting data, each acquisition device (including the head device and the left and right hand devices) first stores the collected data in local storage media. During local storage, the collected data can be organized and packaged using a unified task number, unified timestamp, and device identification information to facilitate subsequent multi-node data aggregation and post-processing reconstruction.

[0145] When the user issues a "stop data collection" command through edge device 30, or when the preset data collection duration is reached, the system enters the task completion phase. At this time, each data collection device stops sensor sampling, closes the data file currently being written, and generates a manifest file containing task metadata. This manifest records the task's unique identifier (UUID), start and end times, calibration model version, device serial number, and checksum of each data file.

[0146] Step S807: The head device 10 sends the hand image and head inertial data to the server 40.

[0147] After each data acquisition device finishes collecting data, the system enters upload standby mode. Considering the possibility of weak network conditions at the acquisition site, the system does not force the immediate upload of massive amounts of raw data. Instead, each acquisition device first attempts to establish a high-speed connection with server 40 or a local relay server. If network conditions meet a preset threshold (e.g., Wi-Fi signal strength greater than -65dBm and sufficient bandwidth), the system automatically starts a background transmission process, uploading data packets from local storage in fragments. This "local recording, post-upload" strategy completely eliminates the reliance on real-time wireless links during the acquisition process, ensuring zero data loss and high fidelity even in outdoor environments with weak or no network access.

[0148] In one different embodiment, an edge computing node can be used instead of the remote server 40 as the data receiving and preliminary processing terminal. That is, the data uploaded by the head device 10 can be directly received at the collection site through a portable high-performance computing device, and data quality verification and preprocessing can be performed in real time. This can retain the robustness of offline collection and shorten the cycle from collection to feedback, which is suitable for scientific research or debugging scenarios that require rapid iterative verification.

[0149] The head-mounted device 10 can send RGB images of the hand and head inertial data to the server 40. In one embodiment, the head-mounted device 10 can also send infrared marker image data it has acquired to the server 40.

[0150] Step S809: The hand device 20 sends the hand inertial data to the server 40.

[0151] In parallel with the head device, the hand device 20 also sends its locally stored hand inertial data to the server 40. Since the volume of hand inertial data is much smaller than image data, its bandwidth requirements are lower, but the integrity of its timestamps is extremely important. After receiving the two data streams from the head device 10 and the hand device 20, the server 40 performs multimodal alignment based on the unified timestamps carried by each stream, reconstructing the operator's complete hand-eye-head coordinated motion trajectory during the acquisition period. If a network interruption occurs during transmission, the system supports a breakpoint resumption mechanism, retransmitting only unacknowledged data fragments, thereby improving transmission efficiency while ensuring data temporal continuity.

[0152] In one embodiment, the hand device 20 can also send the collected induced voltage data, tactile data, and other data to the server 40.

[0153] In this process, the data transmission between the hand device 20 and the head device 10 can be asynchronous. As long as both data have a valid timestamp based on the same time base, the server 40 can complete the accurate spatiotemporal registration in the post-processing stage without having to force the real-time synchronous transmission of the two data streams at the acquisition end.

[0154] Step S811: The server estimates the hand pose based on multimodal data.

[0155] During the offline processing phase, the system can fuse multimodal data to estimate hand pose. This fusion process assigns different weights to each modality's results based on confidence, temporal continuity, observability, and consistency with the back of the hand pose. Specifically, the inertial skeleton result serves as the primary estimation result, providing continuity of finger pose and overall structure; the visual result serves as an auxiliary estimation result, used to constrain, correct, or verify consistency of the inertial result in local timeframes. This fusion method yields more accurate and stable finger keypoint pose estimation results under conventional acquisition modes.

[0156] In one implementation, during calibration mode, a finger skeletal model of the user's current wearing state, along with the correspondence between each IMU in the hand and key points of the hand skeleton, has been established. During the routine data acquisition phase, multiple IMUs on each hand continuously output pose information for corresponding parts. For example, fourteen IMUs are set up for each hand, distributed on the back of the hand and key parts of each finger. Adjacent IMUs can obtain rotational state information of corresponding joints, thereby forming three-degree-of-freedom pose changes for each joint.

[0157] The rotational information output by the IMU can be fed into the calibrated finger skeleton model to reconstruct the posture relationships of each finger joint at the current moment and further calculate the positions of key points in each finger. To reduce the impact of rapid movement, local vibration, or integration errors on the results, a dynamic compensation term can be constructed by combining the accelerometer output to correct the joint posture and key point positions, thereby forming a finger skeleton reconstruction result based on inertial information. This inertial-based skeleton reconstruction path is the main source for obtaining the finger key point pose in this invention.

[0158] Simultaneously, based on the hand images acquired by the visible light camera of the headband, 3D hand pose estimation can be performed on the image region within the hand target box to obtain vision-based finger joint pose results. For the left hand image, mirror normalization processing can be performed first to unify the reference representation of the left and right hands. The visual estimation results can output parameterized hand model parameters, or be further decoded to obtain the 3D hand joint positions in the camera coordinate system. Compared to the inertial path, the visual path is mainly used to provide auxiliary constraints, anomaly verification, and offline fusion reference, and its weight in the final finger keypoint results can be set lower than that of the inertial skeleton reconstruction path.

[0159] By using the above method, this invention combines the calibrated individualized skeletal model, the continuous posture information of inertial nodes, and the visual-assisted estimation results, so that the finger keypoint pose estimation has both high temporal continuity and local motion sensitivity, and can suppress the drift, error accumulation and individual model deviation problems existing in the pure inertial scheme through visual results, thereby improving the accuracy and stability of finger pose estimation in complex operation scenarios.

[0160] In one different embodiment, a deep learning-based end-to-end fusion network can be used to replace the traditional geometric optimization method to perform hand pose estimation. This network takes the original image sequence and IMU time-series signal as input directly and learns the confidence weights of different modalities in different scenarios through an attention mechanism. This enables robust pose regression without explicit modeling of physical constraints, and is particularly suitable for severe occlusion or non-rigid deformation scenarios that are difficult to handle by traditional geometric methods.

[0161] Based on the method flow consisting of steps S801 to S811, the combination of a distributed acquisition architecture and a unified time base achieves precise synchronization and decoupled transmission of multimodal data at the source end. This ensures that natural hand movements are not interfered with by cables or real-time communication, and provides a high-quality data foundation with spatiotemporal alignment for the backend. Furthermore, through multimodal fusion estimation on the server side, the inherent defects of single vision being susceptible to occlusion and single inertia being prone to cumulative drift are effectively overcome, so that the final output hand posture has both global accuracy and local continuity, meeting the stringent requirements of embodied intelligence training for data authenticity and completeness.

[0162] To further improve positioning reliability during long-term data acquisition, this invention introduces a periodic multi-view accuracy verification mechanism in the conventional acquisition mode. During conventional acquisition, the head device can trigger a verification at preset time intervals, preset frame intervals, or when the active infrared marker meets high-confidence observation conditions. This verification uses the rigid body posture of the back of the hand as a reference to periodically check the finger bone results obtained based on inertial node reconstruction and the node position results obtained based on active infrared marker observation, thereby determining whether the current hand reconstruction result meets expectations. When abnormal deviations are detected, corresponding correction or recovery operations are triggered to suppress cumulative errors and local drift during long-term acquisition.

[0163] Figure 9 This is a flowchart of a method for verifying hand parameters in an embodiment of this specification. The method includes the following steps S901-S911.

[0164] like Figure 9 As shown, in step S901: the head device 10 sends a verification instruction to the hand device 20.

[0165] In this step, the head device 10, acting as the system's master control node and time synchronization reference source, generates a verification command when preset verification trigger conditions are met and sends it to the hand device 20 via a low-latency wireless link. The trigger conditions can be based on a fixed time interval (e.g., every 5 minutes), a fixed frame interval (e.g., every 3000 frames), or real-time monitored abnormal data indicators (such as a sudden increase in inertial integral residual or a decrease in visual tracking confidence). The verification command may include information such as the task identifier for this verification, the expected duration, and infrared light control parameters, enabling the hand device 20 to enter a dedicated verification working mode instead of using the conventional acquisition configuration. This verification mechanism, initiated proactively by the head device, ensures precise timing coordination between the verification action and head visual acquisition, avoiding data window misalignment issues caused by asynchronous triggering at both ends.

[0166] In one different embodiment, an edge device 30 can be used instead of a head device 10 as the initiator of the verification command. The edge device 30 sends synchronous verification commands to both the head device 10 and the hand device 20 simultaneously via Bluetooth or Wi-Fi link. This approach is suitable for scenarios that require manual confirmation by the user or interactive verification in conjunction with an external guidance interface, and can more closely bind the verification process with the user's operational intent.

[0167] Step S903: The hand device 20 controls the infrared light to turn on.

[0168] Upon receiving the verification command, the hand device 20 immediately controls multiple infrared marker units deployed on the back of the hand and each finger joint to illuminate according to a preset timing coding pattern. Unlike the infrared lights in conventional acquisition mode, which can employ a low-power intermittent flashing strategy, in verification mode, the infrared lights operate in full-brightness continuous or high-frequency pulse mode to ensure that the infrared camera of the head device 10 can capture a sufficient number of high signal-to-noise ratio marker point images within a short verification window. The activation time of the infrared lights is precisely marked by the second clock of the hand device 20 and shares the same time reference with the subsequently acquired inertial data, thereby ensuring strict alignment of optical and inertial observations on the time axis.

[0169] The key to this step is that the infrared markers not only provide spatial positioning features, but their on / off timing also carries identification coding information, enabling the head device 10 to accurately distinguish the affiliation of each node in the case of densely distributed multiple marker points, providing unambiguous geometric constraint input for subsequent parameter verification.

[0170] Step S907: The hand device 20 collects hand inertial data via IMU.

[0171] While the infrared lights are on, multiple inertial measurement units (IMUs) in the hand device 20 continuously output triaxial acceleration, triaxial angular velocity, and optional triaxial magnetometer data. These hand inertial data reflect the instantaneous motion state of each joint of the hand during the calibration period and are the core basis for evaluating the consistency between the inertial calculation results under the current calibration parameters and external observations.

[0172] In one implementation, during the verification process, the user can be prompted by an edge device to perform specific small movements (such as slowly flexing or extending fingers or slightly rotating the wrist) within a short period of time. In another implementation, the verification process can be performed in a manner that is imperceptible to the user.

[0173] The IMU sampling rate can be maintained at the same high-frequency setting as conventional acquisition to fully capture subtle changes in the dynamic process. Each IMU data sample carries a microsecond-level timestamp provided by the second clock to ensure that it corresponds precisely in time with the emission time of the infrared lamp in step S903 and the subsequent transmitted image data.

[0174] Step S909: The hand device 20 sends hand inertial data to the head device 10.

[0175] The hand device 20 transmits the hand inertial data collected during the verification period to the head device 10 in real-time or near real-time via a wireless communication link. Compared to the strategy of prioritizing local storage and then uploading in batches afterward in the conventional acquisition mode, the data transmission in the verification mode emphasizes low latency and immediacy, so that the head device 10 can complete parameter verification as quickly as possible and decide whether to continue the current acquisition task or trigger recalibration. The transmitted data packet contains the complete raw IMU readings, the corresponding timestamp sequence, and the verification task identifier. After receiving it, the head device 10 can perform spatiotemporal registration with the locally cached infrared marker image. If a brief packet loss occurs during transmission, the system can fill in the missing segments based on timestamp interpolation or a retransmission request mechanism to ensure the integrity of the verification algorithm input. This end-side direct connection transmission architecture avoids the additional latency caused by server relay, enabling the verification loop to be completed within milliseconds, minimizing the impact of interruptions on the normal acquisition process.

[0176] In step S911, the head device verifies the hand parameters based on the infrared marker image and hand inertial data.

[0177] After acquiring time-aligned infrared marker images and hand inertial data, the head device 10 executes a multimodal joint verification algorithm to evaluate the validity of the current hand parameters.

[0178] Figure 10 This specification illustrates a flowchart of a method for verifying hand parameters using a hand device according to an embodiment of the present invention, which may specifically include the following steps.

[0179] Step S1001: Determine whether the positional relationship of the fingers relative to the back of the hand obtained based on hand inertial data meets expectations.

[0180] In this step, the system uses the rigid body pose of the back of the hand as a reference to verify the consistency of finger skeletal key points reconstructed based on multiple IMUs and individualized skeleton models. Because IMU data is prone to cumulative errors after long-term integration, and flexible gloves may slip during repeated wear or vigorous movement, the inertial-calculated finger positions gradually deviate from the actual anatomical structure. Therefore, the system compares the position vectors of each finger key point relative to the back of the hand coordinate system at the current moment with preset skeletal structure constraints to check whether they meet prior conditions such as constant phalanx length, reasonable joint range of motion, and motion continuity. If the relative position of a finger key point exceeds the allowable tolerance zone, or the displacement abrupt change between adjacent frames exceeds the physical threshold, the inertial observation of that channel is deemed abnormal. This checking process essentially uses the known rigid body pose of the back of the hand as an anchor point to reverse-verify whether the unknown finger pose calculated by the inertial link is still within the confidence domain, thus transforming the abstract drift problem into a quantifiable geometric deviation judgment.

[0181] In one different embodiment, residual testing based on a dynamic model can be used instead of pure geometric constraints to determine whether the inertial data meets expectations. That is, the dynamic equation of finger movement is constructed using the acceleration and angular velocity output by the IMU, and the residual between the theoretically predicted position and the actual integrated position is calculated. When the residual energy spectrum increases significantly in a specific frequency band, it is determined to be non-rigid interference caused by sensor loosening or poor contact. This method can identify potential fault modes that have not yet manifested as obvious positional deviations earlier.

[0182] Step S1003: Determine whether the positional relationship between the fingers and the back of the hand meets expectations based on the position of the infrared marker unit.

[0183] In parallel with the inertial path, the system also utilizes an infrared camera on the head unit to observe the active infrared markers from multiple perspectives and reconstruct the three-dimensional positions of each marker point in the camera coordinate system. Since the infrared markers and inertial nodes have a fixed spatial correspondence on the hand unit, and reference markers are usually also arranged on the rigid body of the back of the hand, the infrared marker positions in the finger area can be transformed to the back of the hand coordinate system, thereby determining whether the positional relationship between the fingers and the back of the hand meets the aforementioned preset constraints. This step provides an external observation channel independent of the inertial data to cross-verify the judgment results in step S901. Through the above checks, the system can identify abnormal deviations caused by accumulated inertial errors, local occlusion, false detection of active markers, rapid movement, or local model mismatch. For example, if the relative position of the marker point observed by infrared is consistent with the nominal value, but the inertial calculation result deviates, it can be confirmed that the deviation originates from IMU drift; conversely, if both deviate but in different directions, it may indicate that the glove has undergone overall deformation or wear displacement.

[0184] Step S1005: If the expected results are met, maintain the currently calibrated hand parameters.

[0185] When the verification results of steps S1001 and S1003 both indicate that the positional relationship between the finger key points and the infrared marker units relative to the back of the hand is within the normal range, the system determines that the current hand reconstruction result is reliable and there is no need to adjust the calibrated hand parameters. At this time, the system continues to use the existing individualized skeletal model, sensor installation offset, and tactile channel mapping relationship for subsequent pose estimation and data output.

[0186] Step S1007: In the event of minor anomalies, correct the results of local finger key points in the hand parameters.

[0187] If the verification results indicate that a deviation exists but is minor—for example, the position of a key finger point deviates from the expected value within an acceptable threshold, or the infrared marker unit shows only slight displacement—the system classifies it as a minor local anomaly. In this case, global recalibration is not triggered. Instead, based on the current high-confidence infrared observation results or skeletal structure constraints, the affected key finger points are corrected online.

[0188] Specifically, the absolute position provided by the infrared markers can be used as a strong constraint. Inverse kinematics or weighted fusion algorithms can be used to adjust the posture estimates of the corresponding phalanges, bringing them back to their reasonable anatomical positions. This local correction strategy can compensate for transient errors caused by short-term occlusion, local vibration, or minor slippage in real time without interrupting the acquisition process. This maintains data continuity and suppresses the propagation of errors to subsequent frames. Compared to directly discarding abnormal frames, this soft correction method better preserves the operator's complete motion trajectory during complex interactions, which is of great value for embodied intelligence models learning adaptive human behavior under non-ideal conditions.

[0189] Step S1009: In the event of a significant anomaly, trigger recalibration.

[0190] When the calibration results show a deviation exceeding a preset safety threshold—for example, multiple finger key points deviating significantly simultaneously, irreversible deformation of the infrared marker layout, or irreconcilable contradictions between inertial and infrared observation results—the system identifies it as a significant anomaly. Such anomalies typically indicate a fundamental change in the glove's wearing state (e.g., significant misalignment, fabric tearing) or a sensor hardware malfunction. In this case, local corrections cannot restore data quality, and continuing to use the current parameters will result in training samples containing systematic errors. Therefore, the system automatically triggers a recalibration process, suspends the current data acquisition task, and guides the user to re-perform the calibration action via an edge device to update the hand model parameters. This mechanism constitutes the system's safety fallback strategy, ensuring that valid data is only produced when the model matches physical reality. In actual deployment, the timestamp and contextual state information of the anomaly are recorded simultaneously with the recalibration trigger, facilitating subsequent offline analysis of the cause of the failure or optimization of the wearable design. Through this tiered response mechanism, the system minimizes unnecessary interruptions while ensuring data quality, achieving a balance between robustness and availability.

[0191] In this way, the system can periodically use the back of the hand reference frame to verify the consistency between finger bone results and active label observation results, thereby improving the positioning stability and data reliability during long-term continuous acquisition.

[0192] The multimodal data acquisition method and system provided in the embodiments of this specification, through innovative hardware integration, robust spatiotemporal synchronization mechanisms, and intelligent multi-source fusion algorithms, construct a highly adaptable, high-precision, comfortable-to-wear, and easily deployable embodied intelligent data acquisition platform. This platform not only meets the urgent need for high-quality first-person perspective data in current vision-language-action (VLA) model training, but also provides reliable technical support for the future development of robotic teleoperation, virtual reality interaction, and human-computer collaboration. Any modifications, equivalent substitutions, or improvements made by those skilled in the art to the above embodiments without departing from the spirit and essence of this invention should be included within the scope of protection of this invention.

[0193] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program a digital system themselves to "integrate" it onto a PLD, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Moreover, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used when writing program development code. The original code before compilation must also be written in a specific programming language, which is called a Hardware Description Language (HDL). There is not just one HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. Currently, the most commonly used are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also understand that by simply performing some logic programming on the method flow using the aforementioned hardware description languages ​​and programming it into an integrated circuit, the hardware circuit that implements the logic method flow can be easily obtained.

[0194] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, ASICs, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.

[0195] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or physical entities, or by products with certain functions. A typical implementation device is a server system. Of course, this application does not exclude the possibility that, with the future development of computer technology, the computer implementing the functions of the above embodiments can be, for example, a personal computer, a laptop computer, an in-vehicle human-machine interaction device, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0196] While one or more embodiments of this specification provide the operational steps of the methods described in the embodiments or flowcharts, more or fewer operational steps may be included based on conventional or non-inventive means. The order of steps listed in the embodiments is merely one possible order of execution among many steps and does not represent the only possible order. In actual device or end product execution, the methods shown in the embodiments or drawings may be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment, or even a distributed data processing environment). The terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, product, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, product, or apparatus. Without further limitations, the presence of other identical or equivalent elements in the process, method, product, or apparatus that includes said elements is not excluded. For example, the use of terms such as "first," "second," etc., is to denote names and does not indicate any particular order.

[0197] For ease of description, the above devices are described in terms of function, divided into various modules. Of course, when implementing one or more of these specifications, the functions of each module can be implemented in one or more software and / or hardware components, or a module that performs the same function can be implemented by a combination of multiple sub-modules or sub-units. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between devices or units, and may be electrical, mechanical, or other forms.

[0198] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0199] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0200] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0201] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0202] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0203] Computer-readable media include both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage, graphene storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0204] Those skilled in the art will understand that one or more embodiments of this specification can be provided as a method, system, or computer program product. Therefore, one or more embodiments of this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, one or more embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0205] One or more embodiments of this specification can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a particular task or implement a particular abstract data type. One or more embodiments of this specification can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0206] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, system embodiments are basically similar to method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. In the description of this specification, the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this specification. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described can be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0207] The above description is merely an embodiment of one or more embodiments of this specification and is not intended to limit the scope of these embodiments. Various modifications and variations can be made to these embodiments by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of the claims.

Claims

1. A multimodal data acquisition method, applied to a system including a head-mounted device worn on an operator's head and a hand-mounted device worn on the operator's hand, the head-mounted device including a first clock, a camera, and a first inertial measurement unit, the hand-mounted device including a second clock and a plurality of second inertial measurement units, the second clock being time-synchronized with the first clock, the method comprising: The head-mounted device acquires hand images via the camera and head inertial data via the first inertial measurement unit. The hand images and head inertial data include timestamps provided by the first clock. The hand images and head inertial data are then sent to the server. The hand device collects hand inertial data through the plurality of second inertial measurement units, the hand inertial data including a timestamp provided by the second clock; and sends the hand inertial data to the server.

2. The multimodal data acquisition method as described in claim 1 further includes: The head device establishes a first communication link with the hand device; The head device periodically broadcasts a synchronization signal packet containing the current first time value of the first clock through the first communication link; The hand device receives the synchronization signal packet, calculates the time deviation between the receiving time and the first time value, and adjusts the second clock based on the time deviation.

3. The multimodal data acquisition method as described in claim 1, wherein the handheld device includes multiple infrared marker units, the camera includes a multi-view camera, the multi-view camera includes a visible light camera and an infrared camera, and the method further includes: In response to hand movements of an operator wearing the system guided by an edge device, the head device acquires a calibrated hand image via the camera, the calibrated hand image including a timestamp provided by a first clock, and the calibrated hand image including an RGB image captured by the visible light camera and an infrared marker image captured by the infrared camera; The hand device acquires calibrated hand inertial data through the plurality of second inertial measurement units, and the calibrated hand inertial data includes a timestamp provided by a second clock; The hand device collects calibrated tactile data of the fingertips and / or palm areas through the tactile sensor, and the calibrated tactile data includes a timestamp provided by a second clock; The hand device provides the calibrated hand inertial data and the calibrated tactile data to the head device; The head device calibrates hand parameters based on the calibrated hand image data, the calibrated hand inertial data, and the calibrated tactile data.

4. The method according to claim 3, wherein the head device calibrates hand parameters based on the calibrated hand image data, the calibrated hand inertial data, and the calibrated tactile data, comprising: The head device determines the positions of key hand points based on RGB images, determines the positions of infrared marker units based on infrared marker images, and determines the relative positions of the plurality of second inertial measurement units based on the calibration inertial data. Hand parameters are generated based on the key point positions, the infrared marker unit positions, and the relative positions of the plurality of second inertial measurement units.

5. The method according to claim 3, further comprising: The hand device controls the multiple infrared marker units to flash periodically according to a preset timing coding pattern, and records multiple coding states as marker state data, wherein the multiple coding states are timestamps provided by a second clock; The hand device provides the marked state data to the head device for calibrating hand parameters.

6. The method according to claim 3, wherein the infrared marking unit and the second inertial measurement unit are integrated into a single unit.

7. The multimodal data acquisition method as described in claim 4, wherein the hand device includes a tactile sensor, and the method further includes: The hand device collects pressure distribution signals in the palm and finger areas through the tactile sensor, and generates tactile data with a timestamp provided by a second clock; The hand device provides the tactile data to the head device; The head device determines the contact detection result based on the tactile data, and optimizes the hand parameters based on the contact detection result.

8. The method according to claim 7, further comprising: During the data acquisition process, the head device sends a verification instruction to the hand device. Based on the verification command, the hand device controls the activation of multiple infrared marker units and collects verification inertial data through multiple inertial measurement units; the verification inertial data is then sent to the head device. The head device acquires verification infrared marker images of the multiple infrared marker units through the infrared camera, and performs hand parameter verification based on the verification infrared marker images and verification inertial data.

9. The method according to claim 8, wherein the verification of hand parameters based on the verification infrared marker image and verification inertial data comprises: Based on the hand parameters and the verification inertial data, determine the first positional relationship of the fingers relative to the back of the hand, and determine whether the first positional relationship meets the preset constraint conditions; The position of each infrared marker unit is determined based on the verified infrared marker image. The second positional relationship of the finger relative to the back of the hand is determined based on the position of each infrared marker unit. It is then determined whether the second positional relationship meets the preset constraint conditions.

10. The multimodal data acquisition method as described in claim 1, wherein the head device sends the hand image and the head inertial data to a server, and the hand device sends the hand inertial data to the server, comprising: During the data acquisition process, the head device and the hand device respectively write the raw data with timestamp information into their respective local storage media; After the data acquisition task is completed, in response to the user's upload command or when the network conditions meet the preset threshold, the original data in the local storage medium is read. The raw data read is encapsulated into a data packet containing a unified task identifier and uploaded to the server via a data transmission link for timestamp-based multimodal data alignment processing.

11. The multimodal data acquisition method as described in claim 1, wherein during the data acquisition process, the head device and the hand device also send device status information to the edge device through a second communication link; The device status information includes at least one of the following: battery status, remaining storage space, sensor operating status, and synchronization link signal strength.

12. A multimodal data acquisition system, comprising: The head-mounted device, configured to be worn on the operator's head, integrates a first clock, a camera, and a first inertial measurement unit; the hand-mounted device, configured to be worn on the operator's hand, integrates a second clock and a second inertial measurement unit; the second clock is synchronized with the first clock. The head-mounted device is used to: acquire hand images via the camera, acquire head inertial data via the first inertial measurement unit, wherein the hand images and the head inertial data include timestamps provided by the first clock; and send the hand images and the head inertial data to the server. The hand device is used to: collect hand inertial data of the hand through the plurality of second inertial measurement units, the hand inertial data including a timestamp provided by the second clock; and send the hand inertial data to the server.

13. The multimodal data acquisition system of claim 12, wherein the handheld device comprises multiple infrared marker units; the camera comprises a multi-view camera, wherein the multi-view camera comprises a visible light camera and an infrared camera. The hand device is also used to: in response to guidance of hand movements of an operator wearing the system by an edge device, acquire a calibrated hand image via the camera, the calibrated hand image including a timestamp provided by a first clock, the calibrated hand image including an RGB image captured by the visible light camera and an infrared marker image captured by the infrared camera; The hand device acquires calibrated hand inertial data through the plurality of second inertial measurement units, and the calibrated hand inertial data includes a timestamp provided by a second clock; The hand device is also used to: acquire calibrated tactile data of the fingertips and / or palm areas via the tactile sensor, the calibrated tactile data including a timestamp provided by a second clock; and provide the calibrated hand inertial data and the calibrated tactile data to the head device; The head device is also used to calibrate hand parameters based on the calibrated hand image data, the calibrated hand inertial data, and the calibrated tactile data.

14. The multimodal data acquisition system of claim 13, wherein the hand device includes a tactile sensor. The hand device is also used to: acquire pressure distribution signals in the palm and finger areas via the tactile sensor, generate tactile data with a timestamp provided by a second clock; and provide the tactile data to the head device; The head device is also used to: determine the contact detection result based on the tactile data, and optimize the hand parameters based on the contact detection result.

15. The multimodal data acquisition system as described in claim 13, wherein the head device includes a headband, and the visible light camera and infrared camera in the multi-view camera unit are arranged in an arc shape in the front area of ​​the headband.