Multi-modal visual fusion positioning and interaction device and cooperative working method thereof

Through the multimodal visual fusion positioning and interaction device, using the ceiling SLAM module, horizontal binocular motion capture module and embedded AI processing module, the data synchronization and information extraction problems in the heterogeneous camera collaborative architecture are solved, and high-precision and stable positioning and interaction are achieved. It is suitable for scenarios such as AR guidance, industrial inspection and XR virtual shooting.

CN120807641AActive Publication Date: 2025-10-17AI TUER
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510967555.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-10-17
Estimated Expiration
2045-07-14

AI Technical Summary

Technical Problem

In the heterogeneous camera collaborative architecture, the time synchronization and spatial alignment problems of sensor data lead to inaccurate fusion data, affecting positioning and interaction effects; the layered positioning fusion algorithm has difficulty accurately extracting and utilizing information in complex environments, resulting in low positioning accuracy; voice interaction has low recognition accuracy in noisy environments, affecting user experience.

Method used

A multimodal visual fusion positioning and interaction device is adopted, including a ceiling SLAM module, a horizontal binocular motion capture module, a heterogeneous camera collaboration module, an embedded AI processing module and a layered positioning fusion module. It uses a monocular camera, an inertial measurement unit, an adjustable baseline binocular RGB camera, an active ring fill light, a neural network processing unit and an improved ORB-SLAM3 framework to perform data fusion and calibration through timestamp synchronization, field of view complementarity strategy and extended Kalman filter.

Benefits of technology

It improves positioning accuracy and stability, enhances the naturalness and smoothness of user interaction, realizes efficient and accurate multimodal visual fusion positioning and interaction in complex environments, and supports real-time interaction needs in scenarios such as AR navigation, industrial inspection and XR virtual shooting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120807641A_ABST
    Figure CN120807641A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal visual fusion positioning and interaction device and a cooperative working method thereof, and the device comprises a ceiling SLAM module which employs a target fixed feature on a ceiling as an anchor point of SLAM to construct an overlook coordinate system; a horizontal binocular motion capture module; the heterogeneous camera cooperation module is used for cooperating with the ceiling SLAM module and the horizontal binocular motion capture module; the embedded AI processing module is a lightweight HRNet model used for monitoring a human skeleton in real time; the SLAM layer is used for carrying out target fixed feature optimization extraction on a ceiling by using an improved ORB-SLAM 3 framework; and the dynamic calibration layer is used for fusing inertial measurement data and a binocular visual odometer through an extended Kalman filter. According to the invention, a multi-modal visual fusion positioning and interaction function which is more accurate and efficient and has a wide application prospect is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer vision, in particular to a multi-modal visual fusion positioning and interaction device and a collaborative working method thereof. BACKGROUND

[0002] In a heterogeneous camera collaborative architecture, different types of sensor data may have problems of time synchronization and spatial alignment in the fusion process, resulting in inaccurate fused data and affecting the positioning and interaction effect. For example, the data update frequency of an inertial measurement unit and a binocular camera is different, and if the synchronization processing is not proper, it may cause positioning delay or error.

[0003] The hierarchical positioning fusion algorithm may not accurately extract and utilize the effective information of each sensor due to the limitations of the algorithm model when processing complex sensor data, resulting in low positioning accuracy. For example, in a complex indoor environment, the algorithm may have difficulty in accurately distinguishing and fusing information from different sensors.

[0004] In AR tour and XR virtual shooting scenarios, the interaction mode of users with virtual objects may not be natural and smooth. For example, the accuracy and response speed of gesture recognition may be insufficient, causing delays or misrecognition when users operate virtual objects.

[0005] The function of voice interaction may not be perfect, with problems such as low recognition accuracy and inaccurate voice command understanding, affecting the user's interaction experience. For example, in a noisy environment, the voice recognition system may not be able to accurately recognize the user's instructions.

[0006] Therefore, a multi-modal visual fusion positioning and interaction device and a collaborative working method thereof are proposed. SUMMARY

[0007] The present application provides a multi-modal visual fusion positioning and interaction device and a collaborative working method thereof, which realizes more accurate, efficient and widely applicable multi-modal visual fusion positioning and interaction functions.

[0008] The present application provides a multi-modal visual fusion positioning and interaction device, comprising: a ceiling SLAM module, a horizontal binocular motion capture module, a heterogeneous camera collaborative module, an embedded AI processing module, and a hierarchical positioning fusion module. The ceiling SLAM module comprises a monocular camera and an inertial measurement unit, and a target fixed feature on the ceiling is used as an anchor point of SLAM to construct a top-down coordinate system. The horizontal binocular motion capture module comprises an adjustable baseline binocular RGB camera and an active ring-shaped light supplement lamp. The isomeric camera coordination module is configured to coordinate the ceiling SLAM module and the horizontal binocular motion capture module. The embedded AI processing module includes a neural network processing unit and a lightweight HRNet model for real-time human skeleton monitoring. The hierarchical positioning fusion module includes a SLAM layer for optimizing and extracting the target fixed features on the ceiling using an improved ORB-SLAM3 framework, and a dynamic calibration layer for fusing inertial measurement data and binocular visual odometry through an extended Kalman filter.

[0009] Optionally, the isomeric camera coordination module includes: A timestamp synchronization unit is configured to align the data acquisition timestamps of the monocular camera, the adjustable-baseline binocular RGB camera, and the inertial measurement unit. A field-of-view complementary strategy unit is configured to activate or sleep independently controlled function sub-modules in the horizontal binocular motion capture module based on the visibility of the target fixed features on the ceiling by the monocular camera. A data fusion preprocessing unit is configured to convert the overhead coordinate system pose coordinates of the ceiling SLAM module and the horizontal coordinate position coordinates of the horizontal binocular motion capture module into global coordinate system pose coordinates.

[0010] Optionally, the hierarchical positioning fusion module includes: The state quantities of the extended Kalman filter include degrees of freedom position, degrees of freedom posture, degrees of freedom velocity, inertial measurement unit accelerometer bias, and inertial measurement unit gyroscope bias. A coupling fusioner is configured to construct an observation model based on the feature point re-projection error output by the horizontal binocular motion capture module, the inertial measurement unit pre-integration constraint, and the target fixed feature point projection constraint of the ceiling SLAM layer. A covariance adjuster is configured to adjust the covariance matrix of the observation model based on the variances of the angular velocity and linear acceleration measured by the inertial measurement unit.

[0011] The present specification provides a collaborative working method of a multi-modal visual fusion positioning and interaction device, including: Obtaining ceiling image data through a ceiling SLAM module and performing preliminary positioning; When the ceiling marker points are not visible, enabling a horizontal binocular motion capture module for supplementary positioning; Directly using the horizontal binocular motion capture module for positioning in an outdoor scene; Real-time detection of human skeletons through a lightweight HRNet model; Fusing inertial measurement unit data, binocular visual odometry, and human skeleton data outputs positioning results.

[0012] Optionally, it comprises: when the improved ORB-SLAM3 framework extracts ceiling features, the weight of circular / linear structures is increased.

[0013] Optionally, it comprises: automatically detecting the ceiling QR code and correcting the installation angle deviation when the device starts.

[0014] Optionally, it further comprises: when applied to an XR virtual shooting scene, a horizontal binocular motion capture module is used to realize real-time preview function.

[0015] Optionally, when the ceiling marker point is invisible, the horizontal binocular motion capture module is enabled for supplementary positioning, comprising: When the monocular camera fails to detect valid target fixed features of the ceiling for a preset number of consecutive frames, the visual inertial odometry mode of the horizontal binocular motion capture module is started, and motion prediction is performed through inertial measurement unit data; When the monocular camera detects valid target fixed features of the ceiling, the position output by the visual inertial odometry mode is combined with the pose calculated by the SLAM layer for joint bundle adjustment, and supplementary positioning is completed.

[0016] In the present application, in AR guide, industrial inspection and XR virtual shooting scenes, the multi-modal visual fusion positioning and interaction device and method can significantly improve the positioning accuracy. Compared with pure visual SLAM, the fixed features on the ceiling are used as anchor points, combined with inertial measurement unit data and binocular visual odometry, which effectively reduces environmental interference and error accumulation, making the positioning more accurate and stable. It has strong anti-interference ability in complex indoor and outdoor environments. Through the heterogeneous camera cooperative architecture and layered positioning fusion algorithm, even if some sensor data is disturbed or missing in some cases, the system can still rely on other sensor information for accurate positioning and tracking. Through hardware acceleration in the embedded AI processing pipeline and multi-task learning strategy, functions such as real-time human skeleton detection, gesture recognition and object interaction detection are realized. In practical application, the system has fast response speed and can meet the needs of real-time interaction of users. BRIEF DESCRIPTION OF DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0018] Figure 1A flowchart of a cooperative working method of a multi-modal visual fusion positioning and interaction device provided by the present application. DETAILED DESCRIPTION

[0019] The following description is provided to enable those skilled in the art to practice the present application. The preferred embodiments described herein are only examples of the present application and it is contemplated that those skilled in the art can devise numerous other variations that are appropriate, given the benefit of this disclosure. The principles defined herein can be applied to other embodiments, variations, improvements, equivalents and other technical solutions without departing from the spirit and scope of the present application.

[0020] The following detailed description is provided to provide a better understanding of the present application. The detailed description includes specific details for the purpose of providing a thorough understanding of various exemplary embodiments of the present application. However, it will be apparent to those skilled in the art that the present application can be practiced without these specific details. In some instances, well-known structures and devices are shown in block diagram form, rather than in detail, in order to avoid obscuring the present application. Additionally, some well-known structures and devices are not shown in order to avoid obscuring the present application. Figure 1 Exemplary embodiments of the present application are described more fully hereinafter with reference to the accompanying drawings. However, the exemplary embodiments can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these exemplary embodiments are provided as a full and enabling disclosure of the present application, and are presented to give those skilled in the art a complete disclosure and comprehensive understanding of the present application and the principles and aspects of the present application as defined herein. Like reference numerals refer to like elements throughout the several views of the drawings.

[0021] In the case of specific embodiments described in accordance with the technical concept of the present application, features, structures, characteristics or other details described in the embodiments do not exclude the possibility of being combined in one or more other embodiments in a suitable manner.

[0022] In the description of specific embodiments, features, structures, characteristics or other details described in the embodiments are intended to enable a person skilled in the art to fully understand the embodiments. However, it does not exclude the possibility that one or more of the specific features, structures, characteristics or other details can be practiced without the technical solution of the present application.

[0023] The flowchart shown in the accompanying drawings is only an exemplary illustration and does not necessarily include all contents and operations / steps, nor does it necessarily have to be executed in the order described. For example, some operations / steps can be further decomposed, and some operations / steps can be combined or partially combined, so the actual execution order can be changed according to the actual situation.

[0024] The block diagram shown in the accompanying drawings is only a functional entity and does not necessarily have to correspond to a physically independent entity. That is, the functional entity can be implemented in the form of software, or in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0025] The term "and / or" or "and / or" includes all combinations of any one or more of the associated listed items.

[0026] In AR navigation and industrial inspection scenarios, when in an environment with complex textures or uneven distribution of feature points, such as densely packed exhibits in a museum or complex equipment layout in a factory workshop, pure visual SLAM may result in decreased positioning accuracy due to difficulty in accurately extracting and matching feature points. For example, in a museum, the material and texture of some exhibits may make it difficult for the camera to accurately capture features, affecting the accuracy of positioning.

[0027] In the case of outdoor strong light or uneven light, such as direct sunlight in outdoor areas or areas with more shadows, the performance of visual sensors may be affected, leading to a decrease in image quality and affecting the stability of positioning. For example, in a strong light environment such as a construction site, the camera may not be able to clearly capture images, causing positioning deviations.

[0028] In the AR navigation scenario, dynamic factors such as the flow of visitors and the movement of exhibits (such as temporary exhibition adjustments in a museum) may cause the pre-constructed map or model to fail, resulting in inaccurate positioning. For example, in a museum, the movement of visitors may block or change the feature points in the environment, affecting the positioning effect of SLAM.

[0029] In the industrial inspection scenario, dynamic factors such as the operation of production equipment and the activities of personnel may interfere with the positioning and navigation of inspection robots. For example, when an automated production line in a factory is running, the environment around the robot is constantly changing, which may cause positioning errors.

[0030] Existing SLAM cameras may not work effectively in certain special scenarios, such as low light, high reflectivity, or transparent object surfaces, making it difficult to extract feature points and affecting positioning accuracy. For example, in a night or dim indoor environment, the camera may not be able to obtain enough image information for positioning.

[0031] Visual sensors are easily affected by factors such as dust and stains, leading to a decrease in image quality and affecting the accuracy and stability of positioning. For example, in a factory workshop environment with a lot of dust, the camera lens is easily contaminated, affecting the positioning effect.

[0032] In the XR virtual shooting scenario, when encountering situations where the ceiling cannot see the marker points, such as special scenarios such as the interior of a car or the upper part of an LED screen, traditional SLAM positioning methods may not be effectively applied. The metal material and complex electronic equipment inside the car may interfere with signal transmission, making it difficult to position; while the high brightness and reflective properties of the upper part of the LED screen may make it difficult for the visual sensor to accurately capture feature points.

[0033] In virtual filming scenarios, frequent scene switching and rapid camera motion can lead to positioning loss or unstable tracking. For example, when shooting an action movie, rapid camera switching and actor movements can prevent the system from updating position information in a timely manner, affecting the filming quality.

[0034] In XR virtual shooting scenarios, achieving real-time preview and interaction requires processing large amounts of image and sensor data, placing high demands on the system's computing power and transmission speed. If system performance is insufficient, problems such as lag and delay may occur, affecting shooting efficiency and quality.

[0035] Compatibility issues can arise between different virtual filming software and hardware devices, leading to poor data transmission and interaction. For example, some virtual filming devices may not be compatible with specific software platforms, limiting creative flexibility.

[0036] The embodiments of this specification provide a multimodal visual fusion positioning and interaction device, including: Ceiling SLAM module, horizontal binocular motion capture module, heterogeneous camera collaboration module, embedded AI processing module, and layered positioning fusion module; The ceiling SLAM module includes a monocular camera and an inertial measurement unit, and constructs a top-down coordinate system using the target fixed features on the ceiling as the anchor points of the SLAM; The horizontal binocular motion capture module includes an adjustable baseline binocular RGB camera and an active ring fill light; The heterogeneous camera collaboration module is used to coordinate with the ceiling SLAM module and the horizontal binocular motion capture module; The embedded AI processing module includes a neural network processing unit and a lightweight HRNet model for real-time monitoring of the human skeleton; The hierarchical positioning fusion module includes a SLAM layer that uses an improved ORB-SLAM3 framework to optimize the extraction of the target fixed features on the ceiling, and a dynamic calibration layer that fuses inertial measurement data with the binocular visual odometry through an extended Kalman filter.

[0037] In the detailed description of the present specification, the ceiling SLAM module adopts a monocular camera combined with an inertial measurement unit, and a camera with a larger FOV is selected to have a larger field of view range and can capture rich feature information on the ceiling. The fixed features on the ceiling (such as lamps, vents, and manually attached marker points) are used as anchor points for SLAM to construct a top-down coordinate system; the horizontal binocular motion capture module is configured with an adjustable baseline binocular RGB camera and an active ring-shaped light supplement, and the depth map and RGB stream are synchronously output through binocular stereo vision; the heterogeneous camera cooperation module makes the two work together; the embedded AI processing module integrates a neural network processing unit (NPU) and deploys a lightweight HRNet model to realize real-time detection of human skeleton; the hierarchical positioning fusion module includes a SLAM layer using an improved ORB-SLAM3 framework and a dynamic calibration layer fusing an inertial measurement unit and a binocular vision odometer through an extended Kalman filter (EKF).

[0038] Optionally, the baseline length adjustment range of the adjustable baseline binocular RGB camera is 10cm-50cm.

[0039] In the detailed description of the present specification, the baseline length of the adjustable baseline binocular RGB camera is flexibly adjusted in the range of 10-50cm, and the stereo vision principle is used to provide data support for motion capture and interaction.

[0040] Optionally, the size of the lightweight HRNet model is compressed from the original 245MB to 85MB.

[0041] In the detailed description of the present specification, the lightweight HRNet model is pruned, quantized and compressed, and after deployment on the NPU, it realizes real-time detection of a 22-keypoint human skeleton.

[0042] Optionally, the ceiling SLAM module and the horizontal binocular motion capture module are kept at a distance of not less than 20cm through a split bracket.

[0043] In the detailed description of the present specification, the ceiling SLAM module and the horizontal binocular motion capture module are physically isolated by a magnesium alloy split head, and the distance between them is not less than 20cm to avoid electromagnetic interference, and the magnesium alloy material ensures the structural stability and reduces the weight.

[0044] The double target calibration adopts Zhang method, which is a classic camera calibration method with high precision and reliability. For the external parameter calibration between the SLAM module and the motion capture module, a chessboard joint optimization method is used. Through accurate calibration before leaving the factory, the coordinate relationship between the modules can be ensured to be accurate, providing a basis for subsequent positioning and motion capture work.

[0045] Optionally, the improved ORB-SLAM3 framework includes a weight adjustment mechanism for circular / linear structure features.

[0046] In the detailed description of the present specification, the improved ORB-SLAM3 framework optimizes feature point extraction for ceiling features, and increases the weight of circular / linear structures during extraction. Since ceiling lamps, vents and other features have specific shape characteristics, this optimization improves positioning accuracy and stability.

[0047] In places such as museums, the device is suspended from the top of the corridor. The real-time positioning accuracy of the SLAM module can reach ±2cm, which can accurately determine the position of the device in space. The binocular module can capture the gestures of visitors. When a visitor points to a certain exhibit with his fingers, the system can recognize the visitor's gesture instructions and display detailed information about the corresponding exhibit, providing an immersive tour experience for visitors.

[0048] Industrial inspection: the device is mounted on a mobile robot, and the ceiling pipeline is used as a SLAM landmark. During the inspection process, the robot can locate its position in real time through the SLAM module, and the binocular module can detect the gesture instructions of the operator (such as stop / accelerate), thereby realizing remote control and operation of the robot. This application can improve the efficiency and safety of industrial inspection.

[0049] XR virtual shooting scene: in the field of film and television production, virtual shooting is often required in various indoor or outdoor scenes. The device can be applied to such scenes, and the normal SLAM camera is used for positioning. When in a complex indoor ceiling scene where no marker points can be seen, the front visual camera can be used for supplementary positioning; it can also be applied in outdoor scenes. Through multi-modal visual fusion positioning technology, high-precision positioning and stable tracking can be provided for virtual shooting, ensuring the authenticity and smoothness of the shooting effect.

[0050] Optionally, the heterogeneous camera coordination module includes: a timestamp synchronization unit for aligning the data acquisition timestamps of the monocular camera, the adjustable baseline binocular RGB camera, and the inertial measurement unit; a field of view complementary strategy unit for activating or hibernating independently controlled function submodules in the horizontal binocular motion capture module based on the visibility of the target fixed features on the ceiling by the monocular camera; a data fusion preprocessing unit for converting the overhead coordinate system pose coordinates of the ceiling SLAM module and the horizontal coordinate position coordinates of the horizontal binocular motion capture module into global coordinate system pose coordinates.

[0051] In the detailed description of the present specification, the data timing consistency is ensured by the time stamp synchronization unit, the data acquisition time stamps of the monocular camera, the horizontal binocular RGB camera and the inertial measurement unit are strictly aligned by using the hardware trigger signal combined with the software delay compensation algorithm. The field of view complementary strategy unit analyzes the capture state of the monocular camera to the fixed features on the ceiling in real time, and intelligently decides that when the features are stable and visible, the ceiling SLAM positioning is mainly relied on, and only the basic perception of the horizontal binocular module or the high-power function of the horizontal binocular module is maintained; when the feature visibility decreases or is lost, the corresponding function subset of the horizontal binocular module is dynamically activated to provide supplementary positioning. The data fusion preprocessing unit is responsible for the unified coordinate system, and uses the conversion relationship determined by the pre-calibration to convert the overhead coordinate system pose output by the ceiling SLAM module and the horizontal coordinate system pose output by the horizontal binocular motion capture module into pose data in the unified global coordinate system in real time, providing a consistent spatial reference for subsequent fusion positioning.

[0052] Optionally, the hierarchical positioning fusion module comprises: The state quantity of the extended Kalman filter comprises a degree of freedom position, a degree of freedom posture, a degree of freedom velocity, an inertial measurement unit accelerometer bias, and an inertial measurement unit gyroscope bias. The coupling fusioner is configured to construct an observation model based on the feature point re-projection error output by the horizontal binocular motion capture module, the inertial measurement unit pre-integration constraint, and the target fixed feature point projection constraint of the ceiling of the SLAM layer; wherein the target fixed feature point projection constraint is a projection constraint on the fixed features on the ceiling. The covariance adjuster is configured to adjust the covariance matrix of the observation model based on the variances of the angular velocity and linear acceleration measured by the inertial measurement unit.

[0053] In the detailed description of the present specification, the extended Kalman filter in the hierarchical positioning fusion module takes the three-degree-of-freedom position, the three-degree-of-freedom attitude, the three-degree-of-freedom velocity of the device, and the accelerometer bias and the gyroscope bias of the inertial measurement unit as the core state vector. The tight coupling fusioner jointly optimizes the feature point re-projection error output by the horizontal binocular motion capture module, the inertial measurement unit pre-integration constraint, and the target fixed feature point projection constraint of the ceiling extracted by the SLAM layer (when visible), to jointly construct a multi-source observation model. The covariance adjuster dynamically assesses the uncertainty of the device motion state according to the fluctuation characteristics of the angular velocity and linear acceleration collected by the inertial measurement unit in real time, and adaptively adjusts the covariance weights of the constraint terms in the observation model, so that the filter prioritizes visual observations when the device is stationary or moving at low speed, and increases the contribution weight of inertial data when the device is moving violently, thereby improving the robustness of the positioning system in dynamic environments.

[0054] As Figure 1As shown, the present specification provides a multi-modal visual fusion positioning and interaction device cooperation method, comprising: S110: Obtain the ceiling image data through the ceiling SLAM module and perform preliminary positioning; S120: When the ceiling marker point is invisible, enable the horizontal binocular motion capture module for supplementary positioning; S130: Directly use the horizontal binocular motion capture module for positioning in an outdoor scene; S140: Real-time detection of human skeleton through a lightweight HRNet model; S150: Fusion of inertial measurement unit data, binocular vision odometry, and human skeleton data to output the positioning result.

[0055] In the specific embodiments of the present specification, the ceiling image data is obtained through the ceiling SLAM module, and the fixed features are used as anchor points for preliminary positioning; when the ceiling marker point is invisible in special scenes such as inside a car and on the upper part of an LED screen, the horizontal binocular motion capture module is enabled for supplementary positioning; the horizontal binocular motion capture module is directly used for positioning in an outdoor scene; the human skeleton is real-time detected through a lightweight HRNet model; the inertial measurement unit data, binocular vision odometry, and skeleton data are fused, and the positioning result is output through dynamic calibration of an extended Kalman filter.

[0056] Optionally, it comprises: when the improved ORB-SLAM3 framework extracts ceiling features, the weight of circular / linear structures is increased.

[0057] In the specific embodiments of the present specification, when the ceiling features are extracted, the weight of circular / linear structure features is increased, because the features such as lamps and ventilation openings on the ceiling have specific shape characteristics, and this optimization improves the effectiveness of feature extraction.

[0058] Optionally, it comprises: automatically detecting the ceiling QR code and correcting the installation angle deviation when the device starts.

[0059] In the specific embodiments of the present specification, the preset ceiling QR code is automatically detected when the device starts, and the installation angle deviation is dynamically corrected according to the marker position and attitude information, and this online self-calibration method improves the positioning accuracy.

[0060] Optionally, it further comprises: when applied to an XR virtual shooting scene, a real-time preview function is realized through the horizontal binocular motion capture module.

[0061] In the specific embodiments of the present specification, when applied to an XR virtual shooting scene, a real-time preview function is realized through the horizontal binocular motion capture module, and stable tracking is provided for virtual shooting.

[0062] Optionally, when the ceiling marker point is invisible, a horizontal binocular motion capture module is enabled for supplementary positioning, comprising: When the monocular camera fails to detect valid target fixed features of the ceiling for a preset number of continuous frames, a visual-inertial odometer mode of the horizontal binocular motion capture module is started to perform motion prediction through inertial measurement unit data; When the monocular camera detects valid target fixed features of the ceiling, the position output by the visual-inertial odometer mode is jointly adjusted with the pose calculated by the SLAM layer to complete supplementary positioning.

[0063] In the specific embodiments of the present specification, when the monocular camera fails to capture valid fixed features of the ceiling for multiple continuous frames, the system automatically switches to a supplementary positioning mode: the visual-inertial odometer function of the horizontal binocular motion capture module is immediately enabled to maintain continuous pose tracking in combination with motion prediction data of the inertial measurement unit; when the ceiling features become visible again, the drift error generated by motion prediction is eliminated by jointly optimizing the pose trajectory accumulated by the visual-inertial odometer and the pose estimate calculated in real time by the SLAM layer based on the ceiling features, realizing seamless positioning recovery.

[0064] In the present application, in the scenarios of AR tour guide, industrial inspection and XR virtual shooting, etc., the multi-modal visual fusion positioning and interaction device and method can significantly improve the positioning accuracy. Compared with pure visual SLAM, the fixed features on the ceiling are used as anchor points, combined with inertial measurement unit data and binocular visual odometer, which effectively reduces environmental interference and error accumulation, making the positioning more accurate and stable. It has strong anti-interference ability in complex indoor and outdoor environments. Through the heterogeneous camera collaborative architecture and layered positioning fusion algorithm, even if some sensor data is disturbed or missing in some cases, the system can still rely on other sensor information for accurate positioning and tracking. Through hardware acceleration in the embedded AI processing pipeline and multi-task learning strategy, functions such as real-time human skeleton detection, gesture recognition and object interaction detection are realized. In practical applications, the system has fast response speed and can meet the needs of real-time interaction of users.

[0065] The above specific embodiments further illustrate the purpose, technical solutions and beneficial effects of the present application. It should be understood that the present application is not inherently related to any specific computer, virtual device or electronic device, and various general-purpose devices can also implement the present application. The above description is only for specific embodiments of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

[0066] The various embodiments in the specification are described in progressive manner, and the same or similar parts between the various embodiments can be mutually referred to, and each embodiment focuses on the difference from other embodiments.

[0067] The above merely provides an example of the present application, and is not intended to limit the present application. The present application can have various modifications and changes for those skilled in the art. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the scope of the claims of the present application.

Claims

1. A multimodal visual fusion positioning and interaction device, characterized in that: include: Ceiling SLAM module, horizontal binocular motion capture module, heterogeneous camera collaboration module, embedded AI processing module, and layered positioning fusion module; The ceiling SLAM module includes a monocular camera and an inertial measurement unit, and constructs a top-down coordinate system using the target fixed features on the ceiling as the anchor points of the SLAM; The horizontal binocular motion capture module includes an adjustable baseline binocular RGB camera and an active ring fill light; The heterogeneous camera collaboration module is used to coordinate with the ceiling SLAM module and the horizontal binocular motion capture module; The heterogeneous camera collaboration module includes: A timestamp synchronization unit, configured to align the data acquisition timestamps of the monocular camera, the adjustable baseline binocular RGB camera, and the inertial measurement unit; A field of view complementation strategy unit, configured to activate or deactivate independently controlled functional submodules in the horizontal binocular motion capture module based on the visibility of the target fixed feature on the ceiling by the monocular camera; A data fusion preprocessing unit, configured to convert the overhead coordinate system pose coordinates of the ceiling SLAM module and the horizontal coordinate position coordinates of the horizontal binocular motion capture module into global coordinate system pose coordinates; The embedded AI processing module includes a neural network processing unit and a lightweight HRNet model for real-time monitoring of the human skeleton; The hierarchical positioning fusion module includes a SLAM layer that uses an improved ORB-SLAM3 framework to optimize the extraction of the target fixed features on the ceiling, and a dynamic calibration layer that fuses inertial measurement data with the binocular visual odometry through an extended Kalman filter.

2. The multimodal visual fusion positioning and interaction device according to claim 1, characterized in that: The layered positioning fusion module includes: The state quantities of the extended Kalman filter include degree of freedom position, degree of freedom posture, degree of freedom velocity, inertial measurement unit accelerometer bias, and inertial measurement unit gyroscope bias; A coupled fuser is configured to construct an observation model based on the feature point reprojection error output by the horizontal binocular motion capture module, the inertial measurement unit pre-integration constraint, and the target fixed feature point projection constraint of the ceiling of the SLAM layer; wherein the target fixed feature point projection constraint is a projection constraint on features fixed on the ceiling; A covariance adjuster is used to adjust the covariance matrix of the observation model based on the variance of the angular velocity and linear acceleration measured by the inertial measurement unit.

3. A collaborative working method of a multimodal visual fusion positioning and interaction device, characterized in that: include: Acquire ceiling image data and perform preliminary positioning through the ceiling SLAM module; When the ceiling markers are not visible, the horizontal binocular motion capture module is used for supplementary positioning; Directly use the horizontal binocular motion capture module for positioning in outdoor scenes; Real-time detection of human skeleton using lightweight HRNet model; The positioning result is output by fusing inertial measurement unit data, binocular visual odometry and human skeleton data.

4. The collaborative working method of the multimodal visual fusion positioning and interaction device according to claim 3, characterized in that: include: The improved ORB-SLAM3 framework increases the weight of circular / linear structures when extracting ceiling features.

5. The collaborative working method of the multimodal visual fusion positioning and interaction device according to claim 4, characterized in that: include: When the device starts up, it automatically detects the ceiling QR code and corrects the installation angle deviation.

6. The collaborative working method of the multimodal visual fusion positioning and interaction device according to claim 5, characterized in that: Also includes: When applied to XR virtual shooting scenes, real-time preview function is achieved through the horizontal binocular motion capture module.

7. The collaborative working method of the multimodal visual fusion positioning and interaction device according to claim 6, characterized in that: When the ceiling markers are not visible, the horizontal binocular motion capture module is enabled for supplementary positioning, including: When the monocular camera fails to detect a valid target fixed feature on the ceiling for a preset number of consecutive frames, activating the visual inertial odometry mode of the horizontal binocular motion capture module to perform motion prediction using inertial measurement unit data; When the monocular camera detects the effective target fixed feature of the ceiling, the position output by the visual inertial odometry mode is jointly bundle adjusted with the pose calculated by the SLAM layer to complete the supplementary positioning.

Citation Information

Patent Citations

  • Video data acquisition method, video creation method and related products

    CN113570730A

  • Positioning and mapping method based on multi-sensor fusion and two-dimensional code correction

    CN113706626A

  • Hand-tracking with IR camera for XR systems

    US20240193982A1