Multimodal vision fusion positioning and interaction device and collaborative working method thereof

By using a multimodal visual fusion positioning and interaction device and its collaborative working method, the problems of time synchronization and spatial alignment of sensor data in heterogeneous camera collaborative work were solved, achieving efficient positioning and interaction effects and improving user experience.

CN120807641BActive Publication Date: 2026-05-19AI TUER
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
AI TUER
Filing Date
2025-07-14
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

In heterogeneous camera collaborative architecture, the time synchronization and spatial alignment issues of sensor data lead to inaccurate fused data, affecting positioning and interaction effects; hierarchical positioning fusion algorithms have difficulty accurately extracting and utilizing information in complex environments, resulting in low positioning accuracy; and low voice interaction recognition accuracy affects user experience.

Method used

A multimodal visual fusion positioning and interaction device is adopted, including a ceiling SLAM module, a horizontal binocular motion capture module, a heterogeneous camera collaboration module, an embedded AI processing module, and a hierarchical positioning fusion module. Data fusion is performed through timestamp synchronization, field-of-view complementarity strategy, and extended Kalman filter. Combined with inertial measurement unit and binocular visual odometry, accurate and efficient positioning and interaction are achieved.

Benefits of technology

It improves positioning accuracy and stability, enhances the system's anti-interference capability, and enables real-time human skeleton detection, gesture recognition, and object-to-object interaction, meeting users' real-time interaction needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120807641B_ABST
    Figure CN120807641B_ABST
Patent Text Reader

Abstract

The application provides a multimodal vision fusion positioning and interaction device and a cooperative working method thereof, which comprises: a ceiling SLAM module, which constructs a top-down coordinate system by taking a target fixed feature on a ceiling as an anchor point of SLAM; a horizontal binocular motion capture module; a heterogeneous camera cooperative module, which is used for cooperating with the ceiling SLAM module and the horizontal binocular motion capture module; an embedded AI processing module, which is used for a lightweight HRNet model for real-time human skeleton monitoring; a SLAM layer for optimizing and extracting the target fixed feature on the ceiling by using an improved ORB-SLAM3 framework; and a dynamic calibration layer for fusing inertial measurement data and binocular vision odometry by extending a Kalman filter. The application realizes a more accurate, efficient and widely applicable multimodal vision fusion positioning and interaction function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, and in particular to a multimodal visual fusion positioning and interaction device and its collaborative working method. Background Technology

[0002] In heterogeneous camera collaborative architectures, data from different types of sensors may encounter time synchronization and spatial alignment issues during the fusion process, leading to inaccurate fused data and affecting positioning and interaction performance. For example, the inertial measurement unit and stereo cameras update data at different frequencies; improper synchronization can result in positioning delays or errors.

[0003] When processing complex sensor data, hierarchical positioning fusion algorithms may fail to accurately extract and utilize the effective information from each sensor due to limitations in the algorithm model, resulting in low positioning accuracy. For example, in complex indoor environments, the algorithm may struggle to accurately distinguish and fuse information from different sensors.

[0004] In AR navigation and XR virtual shooting scenarios, the way users interact with virtual objects may not be natural or smooth enough. For example, the accuracy and response speed of gesture recognition may be insufficient, leading to delays or misrecognition when users operate virtual objects.

[0005] Voice interaction functionality may not be perfect, with issues such as low recognition accuracy and inaccurate understanding of voice commands, affecting the user's interactive experience. For example, in noisy environments, the voice recognition system may not be able to accurately recognize the user's commands.

[0006] Therefore, a multimodal visual fusion positioning and interaction device and its collaborative working method are proposed. Summary of the Invention

[0007] This specification provides a multimodal visual fusion positioning and interaction device and its collaborative working method, which realizes more accurate, efficient and widely applicable multimodal visual fusion positioning and interaction functions.

[0008] This specification provides a multimodal visual fusion positioning and interaction device, including:

[0009] Ceiling SLAM module, horizontal binocular motion capture module, heterogeneous camera collaboration module, embedded AI processing module, and hierarchical localization fusion module;

[0010] The ceiling SLAM module includes a monocular camera and an inertial measurement unit, and constructs a top-view coordinate system using fixed features of the target on the ceiling as anchor points for SLAM.

[0011] The horizontal binocular motion capture module includes an adjustable baseline binocular RGB camera and an active ring light;

[0012] The heterogeneous camera collaboration module is used to coordinate with the ceiling SLAM module and the horizontal binocular motion capture module;

[0013] The embedded AI processing module includes a neural network processing unit and a lightweight HRNet model for real-time monitoring of the human skeleton.

[0014] The hierarchical localization fusion module includes a SLAM layer that uses an improved ORB-SLAM3 framework to optimize the extraction of fixed features of the target on the ceiling, and a dynamic calibration layer that fuses inertial measurement data with binocular visual odometry through an extended Kalman filter.

[0015] Optionally, the heterogeneous camera collaboration module includes:

[0016] A timestamp synchronization unit is used to align the data acquisition timestamps of the monocular camera, the adjustable baseline binocular RGB camera, and the inertial measurement unit.

[0017] A field-of-view complementarity strategy unit is used to activate or deactivate independently controlled functional sub-modules in the horizontal binocular motion capture module based on the visibility of the fixed features of the target on the ceiling by the monocular camera.

[0018] The data fusion preprocessing unit is used to convert the top-view coordinate system pose coordinates of the ceiling SLAM module and the horizontal coordinate position coordinates of the horizontal binocular motion capture module into global coordinate system pose coordinates.

[0019] Optionally, the hierarchical positioning fusion module includes:

[0020] The state variables of the extended Kalman filter include the degree of freedom position, degree of freedom attitude, degree of freedom velocity, inertial measurement unit accelerometer bias, and inertial measurement unit gyroscope bias.

[0021] A coupling fusion unit is used to construct an observation model based on the feature point reprojection error output by the horizontal binocular motion capture module, the pre-integration constraint of the inertial measurement unit, and the target fixed feature point projection constraint on the ceiling of the SLAM layer; wherein, the target fixed feature point projection constraint is a projection constraint on fixed features on the ceiling;

[0022] A covariance adjuster is used to adjust the covariance matrix of the observation model based on the variance of the angular velocity and linear acceleration measured by the inertial measurement unit.

[0023] This specification provides a collaborative working method for a multimodal visual fusion positioning and interaction device, including:

[0024] Ceiling image data is acquired and preliminary positioning is performed using a ceiling SLAM module;

[0025] When the ceiling markers are not visible, the horizontal binocular motion capture module is activated for supplementary positioning.

[0026] In outdoor scenes, horizontal binocular motion capture modules are used directly for positioning.

[0027] Real-time detection of the human skeleton using a lightweight HRNet model;

[0028] The system integrates inertial measurement unit data, binocular visual odometry, and human skeleton data to output positioning results.

[0029] Optionally, this includes increasing the weight of circular / linear structures when the improved ORB-SLAM3 framework extracts ceiling features.

[0030] Optional features include: automatically detecting ceiling QR codes and correcting installation angle deviations when the device starts up.

[0031] Optionally, it also includes: when applied to XR virtual shooting scenarios, a real-time preview function can be achieved through a horizontal binocular motion capture module.

[0032] Optionally, when the ceiling marker is not visible, enabling the horizontal binocular motion capture module for supplementary positioning includes:

[0033] When the monocular camera fails to detect a valid fixed feature on the ceiling for a preset number of consecutive frames, the visual inertial odometry mode of the horizontal binocular motion capture module is activated to predict motion using data from the inertial measurement unit.

[0034] When the monocular camera detects a valid target fixed feature on the ceiling, the position output by the visual inertial odometry mode is combined with the pose calculated by the SLAM layer for beam adjustment to complete the supplementary positioning.

[0035] In this invention, a multimodal visual fusion positioning and interaction device and method can significantly improve positioning accuracy in scenarios such as AR navigation, industrial inspection, and XR virtual photography. Compared with pure visual SLAM, using fixed features on the ceiling as anchor points, combined with inertial measurement unit data and binocular visual odometry, effectively reduces environmental interference and error accumulation, making positioning more accurate and stable. It exhibits strong anti-interference capabilities in complex indoor and outdoor environments. Through a heterogeneous camera collaborative architecture and hierarchical positioning fusion algorithm, even if some sensor data is interfered with or missing in certain situations, the system can still rely on other sensor information for accurate positioning and tracking. Through hardware acceleration and multi-task learning strategies in the embedded AI processing pipeline, real-time human skeleton detection, gesture recognition, and object interaction detection functions are realized. In practical applications, the system has a fast response speed, which can meet the user's real-time interaction needs. Attached Figure Description

[0036] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0037] Figure 1 This is a flowchart illustrating a collaborative working method for a multimodal visual fusion positioning and interaction device provided by the present invention. Detailed Implementation

[0038] The following description is intended to disclose the present invention and enable those skilled in the art to implement it. The preferred embodiments described below are merely examples, and other obvious variations will occur to those skilled in the art. The basic principles of the invention defined in the following description can be applied to other embodiments, modifications, improvements, equivalents, and other technical solutions that do not depart from the spirit and scope of the invention.

[0039] The following is in conjunction with the appendix Figure 1 Exemplary embodiments of the invention will be described more fully here. However, exemplary embodiments can be implemented in many forms and should not be construed as limiting the invention to the embodiments set forth herein. Rather, these exemplary embodiments are provided to make the invention more comprehensive and complete, and to facilitate a full communication of the inventive concept to those skilled in the art. The same reference numerals in the figures denote the same or similar elements, components, or parts, and therefore repeated descriptions of them are omitted.

[0040] Subject to the technical concept of this invention, the features, structures, characteristics or other details described in a particular embodiment may be combined in one or more other embodiments in a suitable manner.

[0041] In the description of specific embodiments, the features, structures, characteristics, or other details described in this invention are intended to enable those skilled in the art to fully understand the embodiments. However, it is not excluded that those skilled in the art can practice the technical solutions of this invention without one or more of the specific features, structures, characteristics, or other details.

[0042] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.

[0043] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0044] The terms “and / or” or “and / or” include all combinations of any one or more of the listed items.

[0045] In scenarios such as AR navigation and industrial inspection, when in environments with complex textures or uneven distribution of feature points, such as densely packed exhibit areas in museums or complex equipment layouts in factory workshops, pure visual SLAM may suffer from decreased positioning accuracy due to the difficulty in accurately extracting and matching feature points. For example, in a museum, the material and texture of some exhibits may make it difficult for the camera to accurately capture features, affecting the accuracy of positioning.

[0046] In outdoor environments with strong sunlight or uneven lighting, such as direct sunlight or areas with much shade, the performance of the visual sensor may be affected, leading to decreased image quality and consequently impacting positioning stability. For example, in bright light conditions such as construction sites, the camera may fail to capture clear images, causing positioning errors.

[0047] In AR guided tour scenarios, dynamic factors such as visitor movement and exhibit movement (e.g., adjustments to temporary exhibitions in museums) can render pre-built maps or models ineffective, leading to inaccurate positioning. For example, in a museum, visitor movement may obscure or alter feature points in the environment, affecting the localization performance of SLAM.

[0048] In industrial inspection scenarios, dynamic factors such as the operation of production equipment and the activities of personnel can interfere with the positioning and navigation of inspection robots. For example, when an automated production line in a factory is running, the constantly changing environment around the robot may lead to positioning errors.

[0049] Existing SLAM cameras may not work effectively in certain scenarios, such as low light, highly reflective surfaces, or transparent surfaces, where feature point extraction is difficult and affects positioning accuracy. For example, in nighttime or dimly lit indoor environments, the camera may not be able to acquire enough image information for positioning.

[0050] Visual sensors are susceptible to dust, dirt, and other contaminants, which can degrade image quality and consequently affect the accuracy and stability of positioning. For example, in dusty environments such as factory workshops, camera lenses are easily contaminated, impacting positioning performance.

[0051] In XR virtual shooting scenarios, traditional SLAM localization methods may not be effective when encountering situations where markers are not visible on the ceiling, such as inside a car or on top of an LED screen. The metallic materials and complex electronic devices inside a car may interfere with signal transmission, making localization difficult; while the high brightness and reflective properties of the top of an LED screen may prevent visual sensors from accurately capturing feature points.

[0052] In virtual shooting scenarios, frequent scene changes and rapid camera movements can lead to lost positioning or unstable tracking. For example, when shooting action films, rapid shot transitions and actors' movements may prevent the system from updating position information in a timely manner, affecting the shooting results.

[0053] In XR virtual shooting scenarios, real-time preview and interaction require processing large amounts of image and sensor data, placing high demands on the system's computing power and transmission speed. Insufficient system performance may lead to stuttering, latency, and other issues, affecting shooting efficiency and quality.

[0054] Compatibility issues may exist between different virtual shooting software and hardware devices, leading to poor data transmission and interaction. For example, some virtual shooting devices may not be compatible with specific software platforms, limiting creative flexibility.

[0055] This specification provides an embodiment of a multimodal visual fusion positioning and interaction device, comprising:

[0056] Ceiling SLAM module, horizontal binocular motion capture module, heterogeneous camera collaboration module, embedded AI processing module, and hierarchical localization fusion module;

[0057] The ceiling SLAM module includes a monocular camera and an inertial measurement unit, and constructs a top-view coordinate system using fixed features of the target on the ceiling as anchor points for SLAM.

[0058] The horizontal binocular motion capture module includes an adjustable baseline binocular RGB camera and an active ring light;

[0059] The heterogeneous camera collaboration module is used to coordinate with the ceiling SLAM module and the horizontal binocular motion capture module;

[0060] The embedded AI processing module includes a neural network processing unit and a lightweight HRNet model for real-time monitoring of the human skeleton.

[0061] The hierarchical localization fusion module includes a SLAM layer that uses an improved ORB-SLAM3 framework to optimize the extraction of fixed features of the target on the ceiling, and a dynamic calibration layer that fuses inertial measurement data with binocular visual odometry through an extended Kalman filter.

[0062] In the specific implementation of this specification, the ceiling SLAM module combines a monocular camera with an inertial measurement unit (IMU). A camera with a large field of view (FOV) is selected to capture rich feature information on the ceiling. Fixed features on the ceiling (such as light fixtures, vents, and manually marked points) serve as anchor points for SLAM, constructing a top-down coordinate system. The horizontal binocular motion capture module is equipped with an adjustable baseline binocular RGB camera and an active ring light, simultaneously outputting depth maps and RGB streams through binocular stereo vision. A heterogeneous camera collaboration module enables the two to work together. The embedded AI processing module integrates a neural network processing unit (NPU) and deploys a lightweight HRNet model to achieve real-time human skeleton detection. The hierarchical localization fusion module includes a SLAM layer using an improved ORB-SLAM3 framework and a dynamic calibration layer that fuses the IMU and binocular visual odometry through an extended Kalman filter (EKF).

[0063] Optionally, the baseline length of the adjustable baseline binocular RGB camera can be adjusted from 10cm to 50cm.

[0064] In the specific implementation of this specification, the baseline length of the adjustable baseline binocular RGB camera can be flexibly adjusted within the range of 10-50cm, providing data support for motion capture and interaction through the principle of stereo vision.

[0065] Optionally, the size of the lightweight HRNet model is compressed from the original 245MB to 85MB.

[0066] In the specific implementation of this specification, the lightweight HRNet model, after being pruned, quantized, and compressed, is deployed on the NPU to achieve real-time detection of the human skeleton at 22 key points.

[0067] Optionally, the ceiling SLAM module and the horizontal binocular motion capture module are kept at a distance of not less than 20cm by a split bracket.

[0068] In the specific implementation of this specification, a magnesium alloy split gimbal is used to physically isolate the ceiling SLAM module and the horizontal binocular motion capture module, with a distance of not less than 20cm between them to avoid electromagnetic interference. The magnesium alloy material ensures structural stability and reduces weight.

[0069] Dual-target calibration employs the Zhang method, a classic camera calibration method known for its high accuracy and reliability. For the extrinsic parameter calibration between the SLAM module and the motion capture module, a checkerboard joint optimization method is used. Precise calibration before shipment ensures accurate coordinate relationships between the various modules, providing a foundation for subsequent positioning and motion capture operations.

[0070] Optionally, the improved ORB-SLAM3 framework includes a weight adjustment mechanism for circular / linear structural features.

[0071] In the specific implementation of this specification, the improved ORB-SLAM3 framework optimizes feature point extraction for ceiling features. During the extraction process, the weight of circular / linear structures is increased. Since features such as ceiling lights and vents have specific shape characteristics, this optimization improves positioning accuracy and stability.

[0072] In museums and other venues, the equipment is suspended from the ceiling of corridors. The SLAM module achieves real-time positioning accuracy of ±2cm, accurately determining the device's position in space. The binocular module can capture visitor gestures; when a visitor points to an exhibit, the system recognizes the gesture and displays detailed information about the exhibit, providing an immersive guided tour experience.

[0073] Industrial Inspection: Equipment is mounted on a mobile robot, using ceiling pipes as SLAM landmarks. During inspection, the robot can locate its position in real time using the SLAM module, while the binocular module can detect operator gestures (such as stop / accelerate), enabling remote control and operation of the robot. This application method can improve the efficiency and safety of industrial inspection.

[0074] XR Virtual Shooting Scenes: In film and television production and other fields, virtual shooting is often required in various indoor or outdoor scenes. This device can be applied to such scenarios, using a SLAM camera for positioning as usual. When in complex indoor scenes where markers are not visible on the ceiling, the front visual camera can provide supplementary positioning; it is also applicable in outdoor scenes. Through multimodal visual fusion positioning technology, it can provide high-precision positioning and stable tracking for virtual shooting, ensuring the realism and smoothness of the shooting effect.

[0075] Optionally, the heterogeneous camera collaboration module includes:

[0076] A timestamp synchronization unit is used to align the data acquisition timestamps of the monocular camera, the adjustable baseline binocular RGB camera, and the inertial measurement unit.

[0077] A field-of-view complementarity strategy unit is used to activate or deactivate independently controlled functional sub-modules in the horizontal binocular motion capture module based on the visibility of the fixed features of the target on the ceiling by the monocular camera.

[0078] The data fusion preprocessing unit is used to convert the top-view coordinate system pose coordinates of the ceiling SLAM module and the horizontal coordinate position coordinates of the horizontal binocular motion capture module into global coordinate system pose coordinates.

[0079] In the specific implementation of this specification, a timestamp synchronization unit ensures data timing consistency. A hardware trigger signal combined with a software delay compensation algorithm is used to strictly align the data acquisition timestamps of the monocular camera, the horizontal binocular RGB camera, and the inertial measurement unit. The field-of-view complementarity strategy unit analyzes the monocular camera's capture status of fixed ceiling features in real time and makes intelligent decisions accordingly: when features are stably visible, ceiling SLAM positioning is primarily relied upon, maintaining only the basic perception of the horizontal binocular module or suspending its high-power functions; when feature visibility decreases or is lost, the corresponding functional subset of the horizontal binocular module is dynamically activated to provide supplementary positioning. The data fusion preprocessing unit is responsible for unifying the coordinate system. Using pre-calibrated transformation relationships, it converts the top-view coordinate system pose output by the ceiling SLAM module and the horizontal coordinate system pose output by the horizontal binocular motion capture module into pose data in a unified global coordinate system in real time, providing a consistent spatial reference for subsequent fusion positioning.

[0080] Optionally, the hierarchical positioning fusion module includes:

[0081] The state variables of the extended Kalman filter include the degree of freedom position, degree of freedom attitude, degree of freedom velocity, inertial measurement unit accelerometer bias, and inertial measurement unit gyroscope bias.

[0082] A coupling fusion unit is used to construct an observation model based on the feature point reprojection error output by the horizontal binocular motion capture module, the pre-integration constraint of the inertial measurement unit, and the target fixed feature point projection constraint on the ceiling of the SLAM layer; wherein, the target fixed feature point projection constraint is a projection constraint on fixed features on the ceiling;

[0083] A covariance adjuster is used to adjust the covariance matrix of the observation model based on the variance of the angular velocity and linear acceleration measured by the inertial measurement unit.

[0084] In the specific implementation of this specification, the extended Kalman filter in the hierarchical positioning fusion module constitutes the core state vector based on the device's three-degree-of-freedom position, three-degree-of-freedom attitude, three-degree-of-freedom velocity, and the accelerometer and gyroscope biases of the inertial measurement unit. Its tightly coupled fusion unit jointly constructs a multi-source observation model by optimizing the feature point reprojection error output from the horizontal binocular motion capture module, the inertial measurement unit's pre-integration constraints, and the fixed feature point projection constraints of the ceiling target extracted by the SLAM layer (when visible). The covariance adjuster dynamically evaluates the uncertainty of the device's motion state based on the real-time angular velocity and linear acceleration fluctuation characteristics collected by the inertial measurement unit, and adaptively adjusts the covariance weights of each constraint term in the observation model accordingly. This ensures that the filter prioritizes visual observations when the device is stationary or moving at low speeds, and increases the contribution weight of inertial data during rapid movements, thereby improving the robustness of the positioning system in dynamic environments.

[0085] like Figure 1 As shown, this specification provides a collaborative working method for a multimodal visual fusion positioning and interaction device, including:

[0086] S110: Acquire ceiling image data and perform preliminary positioning using the ceiling SLAM module;

[0087] S120: When the ceiling marker is not visible, activate the horizontal binocular motion capture module for supplementary positioning;

[0088] S130: Directly use horizontal binocular motion capture module for positioning in outdoor scenes;

[0089] S140: Real-time detection of the human skeleton using a lightweight HRNet model;

[0090] S150: Outputs positioning results by integrating inertial measurement unit data, binocular visual odometry, and human skeleton data.

[0091] In the specific implementation of this specification, ceiling image data is acquired through a ceiling SLAM module, and initial positioning is performed using fixed features as anchor points. In special scenarios where ceiling markers are not visible, such as inside a car or above an LED screen, a horizontal binocular motion capture module is activated to supplement positioning. In outdoor scenarios, the horizontal binocular motion capture module is used directly for positioning. The human skeleton is detected in real time using a lightweight HRNet model. The positioning results are output after fusing inertial measurement unit data, binocular visual odometry, and skeleton data and dynamically calibrating the data using an extended Kalman filter.

[0092] Optionally, this includes increasing the weight of circular / linear structures when the improved ORB-SLAM3 framework extracts ceiling features.

[0093] In the specific implementation of this specification, when extracting ceiling features, the weight of circular / linear structural features is increased. Since features such as light fixtures and vents on the ceiling have specific shape characteristics, this optimization improves the effectiveness of feature extraction.

[0094] Optional features include: automatically detecting ceiling QR codes and correcting installation angle deviations when the device starts up.

[0095] In the specific implementation of this specification, the device automatically detects the preset ceiling QR code when it starts up, and dynamically corrects the installation angle deviation based on the marked position and posture information. This online self-calibration method improves positioning accuracy.

[0096] Optionally, it also includes: when applied to XR virtual shooting scenarios, a real-time preview function can be achieved through a horizontal binocular motion capture module.

[0097] In the specific implementation of this specification, when applied to XR virtual shooting scenarios, a real-time preview function is achieved through a horizontal binocular motion capture module, providing stable tracking for virtual shooting.

[0098] Optionally, when the ceiling marker is not visible, enabling the horizontal binocular motion capture module for supplementary positioning includes:

[0099] When the monocular camera fails to detect a valid fixed feature on the ceiling for a preset number of consecutive frames, the visual inertial odometry mode of the horizontal binocular motion capture module is activated to predict motion using data from the inertial measurement unit.

[0100] When the monocular camera detects a valid target fixed feature on the ceiling, the position output by the visual inertial odometry mode is combined with the pose calculated by the SLAM layer for beam adjustment to complete the supplementary positioning.

[0101] In the specific implementation of this specification, when the monocular camera fails to capture effective fixed ceiling features for multiple consecutive frames, the system automatically switches to supplementary positioning mode: immediately activating the visual inertial odometry function of the horizontal binocular motion capture module, and maintaining continuous pose tracking by combining motion prediction data from the inertial measurement unit; when the ceiling features become visible again, by jointly optimizing the pose trajectory accumulated by the visual inertial odometry and the pose estimation calculated in real time by the SLAM layer based on the ceiling features, the drift error caused by motion prediction is eliminated, achieving seamless positioning recovery.

[0102] In this invention, a multimodal visual fusion positioning and interaction device and method can significantly improve positioning accuracy in scenarios such as AR navigation, industrial inspection, and XR virtual photography. Compared with pure visual SLAM, using fixed features on the ceiling as anchor points, combined with inertial measurement unit data and binocular visual odometry, effectively reduces environmental interference and error accumulation, making positioning more accurate and stable. It exhibits strong anti-interference capabilities in complex indoor and outdoor environments. Through a heterogeneous camera collaborative architecture and hierarchical positioning fusion algorithm, even if some sensor data is interfered with or missing in certain situations, the system can still rely on other sensor information for accurate positioning and tracking. Through hardware acceleration and multi-task learning strategies in the embedded AI processing pipeline, real-time human skeleton detection, gesture recognition, and object interaction detection functions are realized. In practical applications, the system has a fast response speed, which can meet the user's real-time interaction needs.

[0103] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the present invention is not inherently related to any specific computer, virtual device, or electronic device, and various general-purpose devices can also implement the present invention. The above descriptions are merely specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

[0104] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

[0105] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A multimodal visual fusion positioning and interaction device, characterized in that, include: Ceiling SLAM module, horizontal binocular motion capture module, heterogeneous camera collaboration module, embedded AI processing module, and hierarchical localization fusion module; The ceiling SLAM module includes a monocular camera and an inertial measurement unit, and constructs a top-view coordinate system using fixed features of the target on the ceiling as anchor points for SLAM. The horizontal binocular motion capture module includes an adjustable baseline binocular RGB camera and an active ring light; The heterogeneous camera collaboration module is used to coordinate with the ceiling SLAM module and the horizontal binocular motion capture module; The heterogeneous camera collaboration module includes: A timestamp synchronization unit is used to align the data acquisition timestamps of the monocular camera, the adjustable baseline binocular RGB camera, and the inertial measurement unit. A field-of-view complementarity strategy unit is used to activate or deactivate independently controlled functional sub-modules in the horizontal binocular motion capture module based on the visibility of the fixed features of the target on the ceiling by the monocular camera. The data fusion preprocessing unit is used to convert the top-view coordinate system pose coordinates of the ceiling SLAM module and the horizontal coordinate position coordinates of the horizontal binocular motion capture module into global coordinate system pose coordinates. The embedded AI processing module includes a neural network processing unit and a lightweight HRNet model for real-time monitoring of the human skeleton. The hierarchical localization fusion module includes a SLAM layer that uses an improved ORB-SLAM3 framework to optimize the extraction of fixed features of the target on the ceiling, and a dynamic calibration layer that fuses inertial measurement data with binocular visual odometry through an extended Kalman filter.

2. The multimodal visual fusion positioning and interaction device as described in claim 1, characterized in that, The hierarchical positioning fusion module includes: The state variables of the extended Kalman filter include the degree of freedom position, degree of freedom attitude, degree of freedom velocity, inertial measurement unit accelerometer bias, and inertial measurement unit gyroscope bias. A coupling fusion unit is used to construct an observation model based on the feature point reprojection error output by the horizontal binocular motion capture module, the pre-integration constraint of the inertial measurement unit, and the target fixed feature point projection constraint of the ceiling of the SLAM layer; wherein, the target fixed feature point projection constraint is a projection constraint on fixed features on the ceiling; A covariance adjuster is used to adjust the covariance matrix of the observation model based on the variance of the angular velocity and linear acceleration measured by the inertial measurement unit.

3. A collaborative working method for a multimodal visual fusion positioning and interaction device, applied to the multimodal visual fusion positioning and interaction device as described in any one of claims 1-2, characterized in that, include: Ceiling image data is acquired and preliminary positioning is performed using a ceiling SLAM module; When the ceiling markers are not visible, the horizontal binocular motion capture module is activated for supplementary positioning. In outdoor scenes, horizontal binocular motion capture modules are used directly for positioning. Real-time detection of the human skeleton using a lightweight HRNet model; The system integrates inertial measurement unit data, binocular visual odometry, and human skeleton data to output positioning results.

4. The collaborative working method of the multimodal visual fusion positioning and interaction device as described in claim 3, characterized in that, include: The improved ORB-SLAM3 framework increases the weight of circular / linear structures when extracting ceiling features.

5. The collaborative working method of the multimodal visual fusion positioning and interaction device as described in claim 4, characterized in that, include: When the equipment starts up, it automatically detects the QR code on the ceiling and corrects for installation angle deviations.

6. The collaborative working method of the multimodal visual fusion positioning and interaction device as described in claim 5, characterized in that, Also includes: When applied to XR virtual shooting scenarios, a real-time preview function is achieved through a horizontal binocular motion capture module.

7. The collaborative working method of the multimodal visual fusion positioning and interaction device as described in claim 6, characterized in that, When the ceiling markers are not visible, the horizontal binocular motion capture module is activated for supplementary positioning, including: When the monocular camera fails to detect a valid fixed feature on the ceiling for a preset number of consecutive frames, the visual inertial odometry mode of the horizontal binocular motion capture module is activated to predict motion using data from the inertial measurement unit. When the monocular camera detects a valid target fixed feature on the ceiling, the position output by the visual inertial odometry mode is combined with the pose calculated by the SLAM layer for beam adjustment to complete the supplementary positioning.