A multi-modal smart glasses system and method for embodied intelligence data collection
Patent Information
- Application Number
- CN202610976468.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-02
- Publication Date
- 2026-09-25
AI Technical Summary
[0003]目前主流的数据采集方案主要分为三类:其一为VR/AR头盔类头戴设备,此类设备集成度较高但体积重量大,佩戴易产生疲劳,且外观突兀,难以在日常公开场景使用,社会接受度低;其二为消费级智能眼镜,此类设备形态轻量化,但仅具备基础拍照与视觉记录功能,缺乏深度感知能力与精细手部动作捕捉能力,定位精度不足,无法满足具身数据的多维度采集需求;其三为分离式采集方案,通过外部动捕系统、数据手套与头戴相机组合实现多模态采集,但此类方案需预先布置场地设备,活动范围受限,且多设备间时间同步难度大,数据对齐精度难以保障
本发明采用眼镜形态的一体化集成设计,将手部动作捕捉、立体视觉感知与空间定位模块融合于常规眼镜结构中,在保持日常佩戴外观的同时实现多维度感知功能,有效提升设备的场景适应性与社会接受度;通过分布式多相机布局与统一的时空同步标定机制,实现手部动作、环境深度与设备位姿数据的精准对齐,避免多设备分离采集带来的配准误差,保障多模态数据的时空一致性;采用被动式视觉感知方案结合红外标记捕捉技术,无需额外投射可见光或依赖外部配套设备,即可在不同光照环境下稳定工作,适配室内外多种真实场景的连续数据采集,整体方案结构紧凑、功耗可控,能够为具身智能模型训练提供自然、高质量的多模态交互数据。
Smart Images

Figure CN122816461A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and in particular to a multimodal smart glasses system and method for embodied intelligence data acquisition. Background Technology
[0002] With the rapid development of embodied intelligence technology, the learning of skills by intelligent agents through interaction with the physical environment has become a core research direction. High-quality multimodal interaction data is a key foundation for supporting the training of embodied intelligence models. Existing embodied intelligence data acquisition needs to simultaneously meet multiple requirements such as first-person visual perception, fine hand movement recording, spatial pose localization, and environmental geometry reconstruction, and also needs to be adapted to long-term continuous acquisition in real-life scenarios.
[0003] Currently, mainstream data acquisition solutions can be divided into three categories: First, VR / AR headsets, which are highly integrated but bulky and heavy, causing fatigue when worn, and have an obtrusive appearance, making them difficult to use in everyday public settings and resulting in low social acceptance; Second, consumer-grade smart glasses, which are lightweight but only have basic photo and visual recording functions, lacking depth perception and fine hand motion capture capabilities, and have insufficient positioning accuracy, failing to meet the multi-dimensional data acquisition needs; Third, separate acquisition solutions, which combine external motion capture systems, data gloves, and head-mounted cameras to achieve multimodal acquisition, but these solutions require pre-deployment of equipment in the field, limiting the range of activity, and making time synchronization between multiple devices difficult, resulting in unreliable data alignment accuracy.
[0004] In summary, existing technologies cannot simultaneously achieve high-precision hand motion capture, environmental depth perception, and spatial positioning in a lightweight form that is close to that of ordinary glasses. Furthermore, the spatiotemporal fusion accuracy of multimodal data is insufficient, making it difficult to meet the needs of continuous and natural embodied intelligent data collection in real-world scenarios. Summary of the Invention
[0005] To address the aforementioned technical problems, this application provides a multimodal smart glasses system and method for embodied intelligent data acquisition.
[0006] In a first aspect, this application provides a multimodal smart glasses system for embodied intelligent data acquisition, including a frame structure and a computing and communication module. The frame structure is worn on a user's head and carries optical components, and the computing and communication module is used for multimodal data processing and external transmission. The system is characterized by further comprising: The infrared hand capture module includes two miniature infrared cameras and a wearable infrared LED marker assembly. The infrared cameras are arranged on both sides of the nose pad at the lower edge of the frame, with the optical axis pointing downward and inward to cover the chest operating space. They are used to acquire hand marker images and reconstruct the three-dimensional skeletal posture of the hand. The binocular RGB vision module includes two RGB cameras symmetrically arranged on the front of the lens frame, used to acquire first-view stereo images and calculate and generate dense depth maps of the environment. The SLAM environment perception module includes two wide-angle RGB cameras respectively arranged on the outer sides of the left and right temples, which are used to execute visual SLAM algorithms to obtain the device's six-degree-of-freedom pose and environment map; The time synchronization and calibration unit is used to perform hardware-triggered synchronization of all cameras and complete the joint calibration and coordinate system unification of multiple cameras, so as to realize the spatiotemporal alignment and fusion of hand skeleton, depth data and device pose.
[0007] Optionally, the baseline distance of the infrared camera is 3-5cm, and the field of view covers an operating range of 20-50cm in front of the chest; the infrared LED marker assembly includes at least 10 actively emitting markers, corresponding to the fingertips and metacarpals of the hand respectively.
[0008] Optionally, the baseline distance between the two RGB cameras of the binocular RGB vision module is 6-10cm, matching the human interpupillary distance range, with a horizontal field of view of not less than 90°, and a dense depth map is generated through a stereo matching algorithm.
[0009] Optionally, the wide-angle RGB camera of the SLAM environment perception module has a field of view of not less than 120° and is arranged facing outward; the computing and communication module integrates a six-axis IMU, and the IMU data is fused with the visual SLAM data to improve the robustness of pose estimation.
[0010] Optionally, the time synchronization and calibration unit uses a unified hardware trigger signal to achieve microsecond-level exposure synchronization of each camera. Through joint calibration, it obtains the intrinsic parameters of each camera, the extrinsic parameters between cameras, and the hand-eye transformation matrix, and unifies the multimodal data to the same world coordinate system.
[0011] Secondly, this application provides a method for acquiring embodied intelligent data using multimodal smart glasses, comprising the following steps: S1. The infrared LED marker images of the hand are acquired through the infrared hand capture module. The three-dimensional coordinates of the markers are obtained through binocular matching and triangulation. The hand skeleton posture data is then fitted and generated. S2. Synchronously acquire first-view stereo image pairs through binocular RGB vision modules, obtain disparity maps through epipolar correction and stereo matching, and calculate and generate dense environmental depth maps. S3. Collect environmental images through the SLAM environment perception module, and after feature extraction, pose calculation and optimization, output the device's six-degree-of-freedom pose and local environment map; S4. By synchronizing time and transforming coordinates, the hand skeleton posture data, environmental depth map and device pose data are aligned to a unified spatiotemporal coordinate system, and the fused multimodal acquisition data is output.
[0012] Optionally, in step S1, the hand skeleton fitting adopts inverse kinematics solution with physical constraints. Based on the 21-joint hand skeleton model, the hand pose is calculated by combining the position of the marker point and the joint angle constraints. When some marker points are occluded, the pose data is completed by kinematic constraints and inter-frame prediction.
[0013] Optionally, in step S2, the stereo matching employs a semi-global matching algorithm or a lightweight deep stereo matching network, and the depth calculation satisfies the formula: ; Where Z is the target point depth value, f is the equivalent focal length of the camera, B is the baseline distance of the binocular camera, and d is the pixel parallax value.
[0014] Optionally, in step S3, visual SLAM adopts a tight coupling scheme based on feature points, and combines loop closure detection and global bundle adjustment to eliminate accumulated drift; in step S4, hardware triggering is used to achieve time synchronization error of each sensor less than 1ms, and multimodal data spatial alignment is achieved through extrinsic parameter transformation.
[0015] In summary, this application includes at least one of the following beneficial technical effects: This invention employs an integrated design in the form of eyeglasses, merging hand motion capture, stereoscopic vision perception, and spatial positioning modules into a conventional eyeglass structure. While maintaining a comfortable everyday appearance, it achieves multi-dimensional perception capabilities, effectively enhancing the device's scene adaptability and social acceptance. Through a distributed multi-camera layout and a unified spatiotemporal synchronization calibration mechanism, it achieves precise alignment of hand motion, environmental depth, and device pose data, avoiding registration errors caused by separate data acquisition from multiple devices and ensuring the spatiotemporal consistency of multimodal data. Utilizing a passive visual perception scheme combined with infrared marker capture technology, it can operate stably under different lighting conditions without the need for additional visible light projection or external supporting equipment. It is adaptable to continuous data acquisition in various real-world indoor and outdoor scenarios. The overall solution is compact, with controllable power consumption, providing natural, high-quality multimodal interactive data for embodied intelligent model training. Attached Figure Description
[0016] Figure 1 This is a flowchart outlining the process of multimodal smart glasses. Detailed Implementation
[0017] The embodiments of this application are described in detail below, and examples of the embodiments are shown in the accompanying drawings.
[0018] In the description of this specification, the references to "certain embodiments," "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples" refer to specific features, structures, materials, or characteristics described in connection with the described embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0019] The multimodal smart glasses system described in this invention adopts a modular design, with the entire system housed on a frame structure resembling ordinary eyeglasses. It is divided into two main parts: the front of the frame and the temples. The front of the frame houses the optical components of a binocular RGB vision module and an infrared hand-tracking module. The left temple integrates the main computing unit, battery, and left-side SLAM wide-angle camera; the right temple integrates a wireless communication module, a six-axis IMU, and a right-side SLAM wide-angle camera. Signal transmission between the frame and temples is achieved through a flexible circuit board at the hinge. The battery and computing components are distributed on both sides to balance the wearing weight.
[0020] The infrared hand capture module includes two miniature global shutter infrared cameras, a near-infrared illumination unit, and a wearable infrared LED marker assembly. The two infrared cameras are symmetrically arranged on either side of the nose pad along the lower edge of the frame, with a baseline distance of 3-5 cm. Their optical axes are positioned downwards and slightly inwards, forming an angle of 30°-45° with the vertical direction, resulting in a 60°×45° field of view. This convergent field of view covers the 20-50 cm operating space in front of the chest, avoiding nasal obstruction. The near-infrared illumination unit is integrated inside the nose pad, with a wavelength of 850 nm, providing uniform surface illumination and ensuring high contrast in the marker image.
[0021] The infrared LED marker assembly comprises 10 actively emitting markers, each consisting of a miniature infrared LED and a diffuser. It is attached to the hand via a flexible ring or adhesive backing, corresponding to the fingertips of the five fingers and five metacarpal bone locations on the back of the hand, and is powered by a miniature button battery. The markers cover the back of the hand and fingertips; when some markers are obscured, posture correction can be achieved through skeletal kinematic constraints.
[0022] The module's workflow is as follows: Infrared cameras synchronously acquire hand images at 120fps. Median filtering and adaptive thresholding are used to extract highlighted marker regions. Ellipse fitting is used to filter candidate markers that meet size constraints. Based on epipolar constraints, candidate points from the left and right cameras are stereo matched. Triangulation is performed using calibrated camera intrinsic and extrinsic parameters to obtain the 3D coordinates of the markers. Finally, the 3D markers are input into a 21-joint hand skeleton model. A physically constrained inverse kinematics solver is used to calculate joint angles, outputting hand skeleton pose data. During the solution process, joint angle limits and constant bone length constraints are incorporated. When the number of visible markers is less than 6, the pose from the previous frame is combined with Kalman filter prediction to complete the occlusion node positions, ensuring the continuity of the pose output.
[0023] The binocular RGB vision module features two symmetrically arranged RGB cameras on the left and right sides of the front of the frame, simulating the position of the human eye with a baseline distance of 6-10cm, matching the human interpupillary distance range. The cameras have a resolution of 1920×1080, a frame rate of 30fps, a horizontal field of view ≥90°, and support HDR mode to adapt to different lighting environments. This module implements two functions: first-person perspective video acquisition and binocular depth calculation. The depth calculation process is as follows: First, hardware triggering ensures synchronized exposure of the left and right cameras with a timestamp error of less than 1ms, and the original image is de-mosaiced and color-corrected. Second, using the calibrated camera intra-camera relative pose calculation correction transformation matrix, epipolar correction is performed on the binocular image to ensure that corresponding pixels are located on the same horizontal scan line. Subsequently, stereo matching calculation is performed to obtain a disparity map, and finally, a dense depth map is calculated based on the geometric relationship between disparity and depth. Stereo matching can employ a semi-global matching algorithm, calculating the matching cost through Census transform, aggregating the cost along eight directions, removing erroneous matches after left-right consistency checks, and then refining to sub-pixel quality. Alternatively, a lightweight deep learning approach can be used, employing the MobileStereoNet network architecture. This network includes a feature extraction module, a cost volume construction module, and a disparity regression module. The input is epipolar-corrected left and right eye RGB images, and the output is a pixel-level disparity map with the same size as the input image. The network can be deployed on edge computing units through quantization compression to meet the computational power requirements for real-time operation. The geometric formula for deep computing is: ; Where Z is the depth value of the target point relative to the camera plane, in meters; f is the equivalent focal length of the camera, in pixels; B is the baseline distance of the stereo camera, in meters; and d is the pixel disparity value of the target point in the left and right images, in pixels. The calculated depth map can be further filled with holes using the fast traversal method to generate complete dense depth data.
[0024] The SLAM environment perception module includes two global shutter wide-angle RGB cameras, respectively positioned on the outer sides of the left and right temples, facing forward and outward, with a field of view ≥120°, forming a distributed wide field of view perception layout. The module adopts the ORB-SLAM3 algorithm framework, tightly coupling and fusing the images from the two cameras through extrinsic parameter constraints, expanding the effective perception field of view, reducing feature loss during rapid head rotation, and improving the robustness of pose tracking.
[0025] The complete algorithm flow includes: First, extracting ORB feature points for each frame of image and establishing feature associations between adjacent frames through optical flow and descriptor matching; Second, completing system initialization and constructing an initial map through essential matrix factorization and triangulation; Then, performing feature matching between the current frame and the local map, solving the six-degree-of-freedom pose of the current frame through the PnP algorithm, and providing an initial guess by combining the IMU pre-integration results; When the keyframe insertion condition is met, triangulating new map points and performing local bundle adjustment to optimize the local pose and map point position; During operation, loop closure detection is performed through the bag-of-words model, and after identifying visited scenes, pose map optimization and global bundle adjustment are performed to eliminate the cumulative drift from long-term operation and output a globally consistent device pose trajectory and sparse environment map.
[0026] The six-axis IMU integrated with the computing and communication module has a sampling frequency of 300Hz. Through pre-integration and tight coupling fusion with visual SLAM, the tracking stability in fast motion and weak texture scenes is improved.
[0027] The time synchronization and calibration unit is responsible for achieving time alignment and spatial coordinate system unification for all sensors. Time synchronization uses an FPGA to generate a unified hardware trigger signal, connecting to the exposure pins of all cameras to achieve microsecond-level exposure synchronization. IMU data is marked with a global timestamp and aligned with camera frames through an interpolation algorithm, ensuring that the timestamp error of all data is less than 1ms. Joint calibration consists of three steps: First, each camera is individually calibrated using a checkerboard pattern to obtain the camera's intrinsic distortion coefficients; second, the relative extrinsic parameters of the infrared camera pair and the binocular RGB camera pair are calibrated sequentially using a stereo calibration board; then, by jointly observing the same calibration board, the rigid transformation matrix between the temple SLAM camera and the main binocular camera is solved; finally, the transformation relationship between the hand-marked coordinate system and the device coordinate system is determined using a hand-eye calibration method. The system defines a three-level coordinate system: the world coordinate system takes the device position at the time of SLAM initialization as the origin, the device coordinate system is fixed at the center of the mirror frame, and the hand coordinate system is fixed at the center of the wrist joint. All modal data are unified to the world coordinate system through extrinsic parameter transformation. The coordinate transformation formula is: ; in, These are the homogeneous coordinates of a point in the world coordinate system. The transformation matrix from the device coordinate system to the world coordinate system is obtained from the device pose output by the SLAM system; This is the extrinsic transformation matrix from the sensor coordinate system to the device coordinate system; This represents the homogeneous coordinates of a spatial point in the corresponding sensor coordinate system. This transformation can unify the hand skeleton points, depth point cloud, and device pose to the same spatial reference system, achieving spatial alignment of multimodal data.
[0028] like Figure 1 As shown, the embodied intelligent data acquisition method of the present invention includes the following steps: S1. Hand skeleton posture reconstruction: The infrared hand capture module is activated to synchronously acquire infrared images of hand marker points. The three-dimensional coordinates of the marker points are obtained through marker point detection, binocular matching and triangulation. The hand skeleton model is fitted by inverse kinematics solution and the hand joint posture sequence is output. S2. Environmental depth information acquisition: Start the binocular RGB vision module, synchronously acquire the first-view binocular image pair, obtain the disparity map through epipolar correction and stereo matching, calculate the dense depth map through the depth geometry formula, and output the aligned RGB-D data. S3. Device pose and environment mapping: Start the SLAM environment perception module, collect images from the wide-angle camera on the temple, combine IMU data to execute the visual SLAM algorithm, calculate the device's six degrees of freedom pose in real time, build a local environment map and eliminate drift through loop closure optimization; S4. Multimodal spatiotemporal fusion output: Based on the hardware synchronous clock, the timestamps of all data are aligned. The calibrated extrinsic matrix transforms all data to a unified world coordinate system and encapsulates it into a timestamp-aligned multimodal dataset for storage or external transmission.
[0029] This embodiment features a basic configuration with a carbon fiber frame, a 140mm front frame width, and a total weight under 80g. The infrared camera uses an OV9281 global shutter sensor with a resolution of 1280×800 and a frame rate of 120fps; the binocular RGB camera uses an OV13B10 sensor with a native resolution of 4160×3120, downsampled to 1080p; and the SLAM wide-angle camera uses an OV7251 global shutter sensor with a resolution of 640×480, a frame rate of 120fps, and a field of view of 120°. The computing unit uses a Qualcomm QCS610 SoC with 128GBeMMC storage, and each temple has a 400mAh lithium polymer battery, supporting continuous data acquisition for over 2 hours.
[0030] The software's operating parameters are as follows: infrared hand capture output frame rate of 120fps with a latency of less than 20ms; binocular depth output frame rate of 30fps with a depth map resolution of 640×480; and SLAM positioning output frame rate of 30fps, which can meet the continuous data acquisition needs in daily scenarios.
[0031] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.
Claims
1. A multimodal smart glasses system for embodied intelligent data acquisition, characterized in that, The system includes a frame structure and a computing and communication module. The frame structure is worn on the user's head and carries optical components. The computing and communication module is used for multimodal data processing and external transmission. The system is characterized by further comprising: The infrared hand capture module includes two miniature infrared cameras and a wearable infrared LED marker assembly. The infrared cameras are arranged on both sides of the nose pad at the lower edge of the frame, with the optical axis pointing downward and inward, forming an angle of 30°-45° with the vertical direction, forming an intersecting field of view to cover the operating space of 20-50cm in front of the chest, for acquiring hand marker images and reconstructing the three-dimensional skeletal posture of the hand. The binocular RGB vision module includes two RGB cameras symmetrically arranged on the front of the frame, with a baseline distance of 6-10cm, matching the range of human interpupillary distance, and a horizontal field of view of not less than 90°. It is used to acquire first-person stereo images and calculate and generate dense depth maps of the environment. The SLAM environment perception module includes two wide-angle RGB cameras respectively arranged on the outer sides of the left and right temples, with a field of view of not less than 120°, facing outwards to form a distributed wide field of view perception layout, used to execute visual SLAM algorithms to obtain the device's six-degree-of-freedom pose and environment map; the computing and communication module integrates a six-axis IMU, and the IMU data is fused with the visual SLAM data to improve the robustness of pose estimation. The time synchronization and calibration unit is used to perform hardware-triggered synchronization of all cameras. It uses a unified hardware trigger signal to achieve microsecond-level exposure synchronization of each camera. Through joint calibration, it obtains the intrinsic parameters of each camera, the extrinsic parameters between cameras, and the hand-eye transformation matrix, unifying the multimodal data into the same world coordinate system, and realizing the spatiotemporal alignment and fusion of hand skeleton, depth data, and device pose.
2. The multimodal smart glasses system for embodied intelligent data acquisition according to claim 1, characterized in that, The baseline distance of the infrared camera is 3-5cm; the infrared LED marker assembly includes at least 10 actively emitting markers, corresponding to the fingertips and metacarpals of the hand, respectively.
3. The multimodal smart glasses system for embodied intelligent data acquisition according to claim 1, characterized in that, The binocular RGB vision module generates a dense depth map through a stereo matching algorithm. The stereo matching uses a semi-global matching algorithm or a lightweight depth stereo matching network, and the depth calculation satisfies the following formula: ; Where Z is the depth value of the target point relative to the camera plane, f is the equivalent focal length of the camera, B is the baseline distance of the binocular camera, and d is the pixel disparity value of the target point in the left and right images.
4. The multimodal smart glasses system for embodied intelligent data acquisition according to claim 1, characterized in that, The visual SLAM adopts a feature point-based tight coupling scheme, which combines loop closure detection and global bundle adjustment to eliminate cumulative drift; the time synchronization and calibration unit achieves a time synchronization error of less than 1ms for each sensor, and realizes multimodal data spatial alignment through extrinsic parameter transformation.
5. A method for acquiring embodied intelligent data using multimodal smart glasses, implemented based on the system described in any one of claims 1-4, characterized in that, Includes the following steps: S1. The infrared LED marker images of the hand are acquired through the infrared hand capture module. The three-dimensional coordinates of the markers are obtained through binocular matching and triangulation. The hand skeleton posture data is then fitted and generated. S2. Synchronously acquire first-view stereo image pairs through binocular RGB vision modules, obtain disparity maps through epipolar correction and stereo matching, and calculate and generate dense environmental depth maps. S3. Collect environmental images through the SLAM environment perception module, and after feature extraction, pose calculation and optimization, output the device's six-degree-of-freedom pose and local environment map; S4. By synchronizing time and transforming coordinates, the hand skeleton posture data, environmental depth map and device pose data are aligned to a unified spatiotemporal coordinate system, and the fused multimodal acquisition data is output.
6. The method for multimodal smart glasses for embodied intelligent data acquisition according to claim 5, characterized in that, In step S1, the hand skeleton fitting adopts inverse kinematics solution with physical constraints. Based on the 21-joint hand skeleton model, the hand posture is calculated by combining the position of the marker point and the joint angle constraint. When some marker points are occluded, the pose data is completed by kinematic constraints and inter-frame prediction.
7. The multimodal smart glasses system and method for embodied intelligent data acquisition according to claim 1, characterized in that, In step S2, stereo matching employs a semi-global matching algorithm or a lightweight deep stereo matching network, and the depth calculation satisfies the formula: ; Where Z is the target point depth value, f is the equivalent focal length of the camera, B is the baseline distance of the binocular camera, and d is the pixel parallax value.
8. The multimodal smart glasses system and method for embodied intelligent data acquisition according to claim 1, characterized in that, In step S3, visual SLAM adopts a tight coupling scheme based on feature points, and combines loop closure detection and global bundle adjustment to eliminate accumulated drift; in step S4, hardware triggering is used to achieve time synchronization error of less than 1ms for each sensor, and multimodal data spatial alignment is achieved through extrinsic parameter transformation.