Directing interaction method based on fusion of vision and inertial measurement units

By using a fusion method of vision and inertial measurement units, and dynamically calling multi-path algorithms and data fusion technology, the problems of cumulative error and hardware dependence of directional positioning technology are solved, and high-precision, low-latency, and easy-to-use directional human-computer interaction positioning is achieved.

CN121807160APending Publication Date: 2026-04-07旌翔(北京)科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing directional positioning technologies suffer from cumulative errors and dependence on specific external hardware deployments, making it impossible to achieve high-precision, low-latency, and easily accessible directional human-computer interaction positioning.

Method used

By employing a fusion method of vision and inertial measurement units, real-time data acquisition is achieved through image acquisition units and inertial measurement units. Combined with multi-path algorithms and data fusion technology, vision and inertial data are dynamically invoked to achieve high-precision, low-latency directional positioning.

Benefits of technology

It achieves high-precision directional positioning without the need for external base station deployment and terminal devices without dedicated receiving hardware, reducing system costs and deployment complexity. It also features robustness across all scenarios, low power consumption, and excellent dynamic response performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121807160A_ABST
    Figure CN121807160A_ABST
Patent Text Reader

Abstract

The invention discloses a directional interaction method based on fusion of vision and inertial measurement units, and relates to the technical field of man-machine interaction and space positioning. The method comprises the following steps: S1, data acquisition and preprocessing: acquiring and aligning data through an image acquisition unit and an inertial measurement unit; s2, data processing and fusion: dynamically calling a plurality of algorithm paths to generate visual observation data based on a motion state or environment characteristics, and deeply fusing the visual observation data with inertial data; and S3, coordinate mapping and application. The method has the following beneficial effects: 1, the dependence of special hardware is eliminated, high-precision pointing can be realized only by a camera and an IMU, and the universality is high; 2, full-scene robustness and high precision are realized, different environments and motion states are adapted through a multi-path dynamic calling mechanism, and accumulated drift is eliminated; 3, resource efficiency and low power consumption are realized, and the power consumption is reduced through dynamic frame rate adjustment and load scheduling; 4, the dynamic response performance is excellent, IMU compensation delay and roller shutter correction are utilized, and smooth following of a cursor is ensured; and 5, targeted screen pointing optimization is carried out, and the pointing precision is improved by using a screen locking path.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of human-computer interaction and spatial positioning technology, specifically to a pointing interaction method based on the fusion of vision and inertial measurement units. Background Technology

[0002] Currently, in smart TVs, AR / VR and other smart display terminals, the technologies for human-computer interaction are mainly divided into non-directional remote control solutions and directional positioning solutions. The current status of these two solutions is as follows: 1. Non-directional remote control solution Such solutions (such as infrared, radio frequency (RF), and Bluetooth technologies) aim to achieve universal control without directional limitations. Their core is to solve the problem of control signal coverage, rather than precise spatial pointing. For example, traditional infrared remote control is inexpensive but requires strict alignment with the device; radio frequency remote control achieves omnidirectional control, but its essence is to solve the problem of command transmission and it has no ability to precisely position the cursor on the screen.

[0003] 2. Directional positioning scheme Such solutions aim to achieve high-precision spatial pointing.

[0004] Pure Inertial Measurement Unit (IMU) approach: This approach relies on the IMU to measure angular velocity and acceleration, and uses integration calculations to infer the device's attitude and position changes. It is an autonomous calculation system that does not depend on external signals.

[0005] UWB / NearLink (UWB) and IMU Fusion Solution: UWB technology achieves centimeter-level absolute positioning accuracy by measuring the time-of-flight of wireless pulse signals. NearLink, as an emerging short-range wireless communication technology, also supports high-precision ranging. The fusion approach is similar: combining the absolute spatial coordinates provided by UWB or NearLink with high-frequency relative motion data provided by the IMU. UWB / NearLink is used to correct the accumulated errors of the IMU over time, while the IMU provides continuous motion tracking during the intervals between wireless signal updates.

[0006] However, the above solutions have some drawbacks, such as the technical limitation that non-directional remote control cannot support modern directional interactive functions such as cursor following and precise clicking.

[0007] In directional positioning schemes, the pure IMU approach calculates pose through integration, and its error (especially with low-cost MEMS-IMUs) accumulates and diverges over time. Small measurement deviations over a short period can lead to meter-level positioning errors after integration, making it unsuitable for independent directional applications requiring stable accuracy over long periods.

[0008] The core drawback of the UWB / Starlight and IMU convergence solution lies in its heavy reliance on external, specialized hardware infrastructure. It requires the deployment of multiple fixed positioning base stations (anchor points) in the environment (such as a room), and terminal devices (such as televisions and monitors) must pre-integrate dedicated UWB or Starlight receiver modules. This results in complex and costly system deployment, and makes it difficult to widely adopt on the vast number of existing consumer electronics devices that do not have this hardware pre-installed.

[0009] In summary, existing directional positioning technologies either suffer from inherent, insurmountable cumulative errors (pure IMU) or are limited by specific external hardware deployments and terminal device support (UWB / Starflash + IMU). Therefore, the industry urgently needs a new solution that achieves high-precision, low-latency directional positioning without relying on specific external base station deployments or requiring pre-integrated dedicated receiving hardware in terminal devices. Summary of the Invention

[0010] The purpose of this invention is to provide a directional interaction method based on the fusion of vision and inertial measurement units. This method does not rely on the deployment of specific external base stations, nor does it require the pre-integration of dedicated receiving hardware in terminal devices. It is a new paradigm of directional positioning based on the tight coupling of vision and inertial. Specifically, this invention aims to solve the cumulative drift problem of pure IMU solutions and the dependence of technologies such as UWB / starflash on specific infrastructure. This will enable a low-cost, high-precision, low-latency directional human-computer interaction positioning method that is easy to popularize and apply on existing consumer electronics terminals (such as TVs and AR / VR devices), thereby solving the problems mentioned in the background art.

[0011] This invention provides a pointing interaction method based on the fusion of vision and inertial measurement unit, comprising the following steps: S1: Data acquisition and preprocessing; Visual data is obtained by acquiring environmental image sequences in real time through the image acquisition unit of the human-computer interaction device; The inertial measurement unit of the human-computer interaction device collects angular velocity and acceleration data in real time to obtain inertial measurement unit data; The visual data and inertial measurement unit data are timestamped to establish a unified time reference. S2: Data Processing and Fusion; Data processing includes visual data processing, inertial measurement unit data preprocessing, and inertial motion tracking processing; The visual data processing involves performing multi-path operations on the acquired image sequences. Based on the motion state or environmental characteristics of the human-computer interaction device, at least one path is dynamically called from multiple preset algorithm paths to generate visual observation data for pointing and positioning. The inertial measurement unit data preprocessing is used to perform signal error suppression processing on the inertial measurement unit data; The inertial motion tracking processing is used to calculate at least one of the short-term attitude increment, velocity increment, or displacement increment of the human-computer interaction device based on angular velocity and acceleration data, to obtain inertial motion tracking data; The data fusion involves deeply fusing visual observation data with inertial motion tracking data; The validity of visual observation data is determined. The validity determination includes: when the number of valid features is lower than a preset threshold, the matching inlier rate is lower than a preset threshold, the optical flow tracking success rate is lower than a preset threshold, or the reprojection residual is higher than a preset threshold, the visual observation data is determined to be invalid. The system uses visual observation data to correct the cumulative error of the inertial measurement unit (IMU), while using high-frequency data from the IMU to compensate for the delay in visual processing, and maintains the system's tracking capability when visual observation data fails. S3: Coordinate Mapping and Applications; The pose after data fusion is mapped onto the coordinate system of the target interaction area and converted into cursor coordinates; The cursor coordinates are transmitted to the smart device in real time so that the smart device can update the cursor position on the target interactive area.

[0012] Preferably, the preset multiple algorithm paths include a feature point matching path, a screen locking path, an optical flow tracing path, a manual marker tracing path, and a line feature tracing path; the dynamic invocation includes independently, serially, or in parallel fusion invocation of the multiple algorithm paths based on the processor load, frame processing latency or available computing power indicators, motion speed, and the number of effective features, matching inlier rate, optical flow tracing success rate, or reprojection residual of the current image frame.

[0013] Preferably, the data fusion employs a nonlinear filtering algorithm based on a prediction-correction mechanism; using an extended Kalman filter or a nonlinear optimization framework, it includes the following steps: Construct a state vector, which includes at least the attitude, velocity, position of the human-computer interaction device and the zero bias error of the inertial measurement unit; The prior state estimate is obtained by recursively calculating the state using the data collected by the inertial measurement unit. The visual observation data is used as the observed values ​​to calculate the observation residuals; The calculated filter gain is used to correct the prior state estimate, update the state vector and error covariance matrix, output the fused pose data, and use the updated zero bias error to correct the subsequent measurement data of the inertial measurement unit.

[0014] Preferably, the screen locking path includes the following steps: Identify the physical screen regions in the image sequence and calculate the pointing offset vector of the center of the physical screen region relative to the optical center of the image in the image coordinate system; When a physical screen area is detected in the image sequence, the screen locking path is activated, and the absolute pointing deviation is corrected using the pointing offset vector.

[0015] Preferably, the screen locking path further includes the following steps: The system identifies static reference objects around the physical screen area in real time and establishes spatial topological constraints between the static reference objects and the physical screen area. The spatial topological constraints include one or more of the following: relative orientation, relative distance, and relative angle. When the physical screen area is missing or obscured in the image, the current position of the center of the physical screen area is predicted and completed using the spatial topological constraints.

[0016] Preferably, the screen locking path further includes the following steps: During the period when the physical screen area completely disappears or the visual observation data fails, the pose is maintained using the zero bias of the calibrated inertial measurement unit; When the physical screen area is re-identified, the screen locking path is used for rapid repositioning to correct the spatial accumulation error caused by pure inertial recursion.

[0017] Preferably, the feature point matching path includes the following steps: Extract key feature points from the current frame image and calculate the feature descriptors of the key feature points; Match the feature descriptors of the current frame with the feature descriptors in the preset reference frame or local map to obtain an initial set of matching point pairs; The random sampling consensus algorithm is used to remove mismatched points from the initial set of matching point pairs to obtain valid matching point pairs; Based on the effective matching point pairs, the pose change of the human-computer interaction device relative to the reference frame is calculated using epipolar geometric constraints or the PnP algorithm.

[0018] Preferably, the optical flow tracing path includes the following steps: Select corner points or regions with significant texture gradients in the previous frame as tracking feature points; The corresponding position of the tracking feature point in the current frame image is calculated using the pyramid optical flow algorithm to obtain the pixel displacement vector of the feature point; A reprojection error function is constructed based on the pixel displacement vector, and the rotation and translation motion parameters of the human-computer interaction device between consecutive frames are solved iteratively by minimizing the reprojection error. When the number of tracking feature points is lower than a preset threshold or the tracking error is greater than a preset threshold, new tracking feature points are extracted again in the current frame.

[0019] Preferably, the manual identification tracking path includes the following steps: Detect whether there are preset artificial markers in the current frame image. The artificial markers include specific geometric patterns, coded marks, or active light-emitting dot matrices. When the artificial marker is detected, the coordinates of the corner point or center point of the artificial marker are extracted; Based on the correspondence between the preset physical size of the artificial marker and the image coordinates, the PnP algorithm is used to directly calculate the six-degree-of-freedom pose of the human-computer interaction device relative to the artificial marker.

[0020] Preferably, the line feature tracing path includes the following steps: Extract the features of line segments in the current frame image and calculate the geometric parameters of the line segments; Match the line segments in the current frame with the line features in the reference frame or local map to construct line feature matching pairs; Construct point-to-line distance constraints or line-to-line angle constraints, and solve the pose changes of the human-computer interaction device through nonlinear optimization to assist in localization in weak texture environments.

[0021] Preferably, it also includes an interactive calibration step before data acquisition: In response to a calibration command, enter calibration mode; In calibration mode, instantaneous visual data and inertial measurement unit attitude data are acquired when the human-computer interaction device points to a specific position in the target interaction area. The feature center in the instantaneous visual data or the current pose of the inertial measurement unit is forcibly mapped to the reference coordinates of the target interaction area, and the coordinate mapping relationship is updated.

[0022] Preferably, the inertial measurement unit further includes a magnetometer; The data fusion further includes: generating heading constraint information based on geomagnetic data collected by a geomagnetometer, and inputting the heading constraint information as an observation into the data fusion algorithm to suppress the cumulative drift of the heading angle.

[0023] Preferably, the generation of the heading constraint information further includes a magnetic interference detection step: Calculate the magnitude and magnetic inclination of the current geomagnetic data and compare them with the preset geomagnetic reference range or reference vector; When the difference is less than a preset threshold, the magnetic environment is determined to be stable, and the heading constraint information is activated. When the difference exceeds a preset threshold, magnetic interference is determined to exist, the current geomagnetic data is discarded, and heading estimation is performed solely based on vision and gyroscope.

[0024] Preferably, the visual data processing further includes an image distortion correction step based on inertial data: Using inertial measurement unit data aligned with image frame time, the angular velocity of the human-computer interaction device during image exposure is calculated; The pixel offset of each row in the image is calculated based on the angular velocity and the shutter readout time of the image acquisition unit. The original image is reverse-sampled or geometrically transformed to eliminate geometric distortion caused by the rolling shutter effect.

[0025] Preferably, the image distortion correction step is triggered only when the angular velocity of the human-computer interaction device exceeds a preset high-speed motion threshold.

[0026] Preferably, the data acquisition further includes a dynamic frame rate adjustment step: Real-time monitoring of the angular velocity and acceleration amplitude output by the inertial measurement unit; When the amplitude is lower than the preset static threshold for a certain period of time, the image acquisition unit is controlled to enter a low frame rate mode to reduce system power consumption. When the amplitude exceeds the preset motion threshold, the image acquisition unit is immediately triggered to switch to high frame rate mode to capture fast motion features.

[0027] Preferably, when the human-computer interaction device is equipped with a laser emission module, the preset multiple algorithm paths further include an active spot tracking path, which includes: During the period when the laser emitting module is activated and projects a laser spot onto the target interaction area, the laser spot in the visual data is identified and used as the main tracking feature. Based on the motion trajectory of the laser spot in the image sequence, the pointing motion of the human-computer interaction device is calculated to maintain the continuity of visual positioning.

[0028] Preferably, the preset multiple algorithm paths also include an active spot tracking path; the dynamic invocation includes independently, serially, or in parallel fusion invocation of the active spot tracking path based on the processor load, frame processing latency or available computing power index, motion speed, and the number of effective features, matching inlier rate, optical flow tracking success rate, or reprojection residual of the current image frame.

[0029] Preferably, the dynamic invocation further includes: When the human-computer interaction device is in a fast motion state and the image blur is higher than a preset threshold, the inertial measurement unit data is used first for motion tracking, and the weights of the feature point matching path, screen lock path, optical flow tracking path, manual mark tracking path, line feature tracking path and active spot tracking path are reduced. When the human-computer interaction device is in a low-speed motion or hovering state, the feature point matching path, screen locking path, optical flow tracing path, manual mark tracing path, line feature tracing path, or active spot tracing path are preferentially invoked to suppress the zero-bias drift of the inertial measurement unit.

[0030] Preferably, data preprocessing also includes a rolling shutter distortion correction algorithm.

[0031] Compared with the prior art, the beneficial effects of the present invention are: 1. Eliminates dependence on dedicated hardware and has strong universality: No external base station (such as UWB anchor point) needs to be deployed, and terminal devices do not need to integrate dedicated receiving modules in advance. High-precision pointing can be achieved using only a camera and IMU, which greatly reduces system cost and deployment complexity and is easy to popularize on a large number of existing consumer electronics products.

[0032] 2. Robustness and High Precision Across All Scenes: Through a multi-path dynamic invocation mechanism, the system can automatically switch to the optimal algorithm based on environmental characteristics (such as texture richness and the presence of a screen) and motion state. For example, it utilizes line features or manual markings in weak texture environments, and relies on IMU and optical flow in fast motion, effectively overcoming the problem of easy tracking loss in complex scenes with a single vision algorithm. At the same time, it completely eliminates the cumulative drift of the pure IMU scheme by utilizing absolute visual observation.

[0033] 3. Resource efficiency and low power consumption: Through dynamic frame rate adjustment and load-based algorithm scheduling, the system can significantly reduce the computing burden and power consumption of the embedded processor while ensuring tracking accuracy, thus extending the battery life of the human-computer interaction device.

[0034] 4. Excellent dynamic response performance: Utilizing high-frequency data from the IMU to compensate for visual processing delays, and combined with rolling shutter distortion correction technology, it ensures that the cursor can still follow smoothly and accurately without noticeable lag or jitter when the user performs rapid flicking or clicking operations.

[0035] 5. Targeted screen pointing optimization: The unique screen locking path can use the geometric features of the screen itself to perform absolute pose correction, so that when facing typical interactive objects such as TVs and large screens, it can achieve pointing accuracy far exceeding that of general SLAM algorithms.

[0036] A pointing interaction method based on the fusion of vision and inertial measurement units Attached Figure Description

[0037] Figure 1 This is a flowchart of the pointing interaction method that integrates vision and inertial measurement units as described in this invention; Figure 2 This is a system flowchart of the pointing interaction method that integrates vision and inertial measurement units as described in this invention; Figure 3 This is a schematic diagram of the hardware system of the human-computer interaction device based on the fusion of vision and inertial measurement units according to the present invention; Figure 4 This is a schematic diagram of the hardware structure of an embodiment of the human-computer interaction device based on the fusion of vision and inertial measurement units according to the present invention; Figure 5 This is a schematic diagram illustrating the working principle of data fusion as described in this invention; Figure 6 This is a flowchart of the coordinate mapping described in this invention; Figure 7 This is a schematic diagram illustrating the working principle of the coordinate mapping described in this invention; Figure 8 This is a schematic diagram illustrating an application scenario of an embodiment of the pointing interaction method based on the fusion of vision and inertial measurement units described in this invention. Figure 9 This is a logical diagram illustrating the dynamic multi-path invocation in the visual data processing described in this invention. Detailed Implementation

[0038] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0039] like Figures 1-9 As shown, this invention provides a pointing interaction method based on the fusion of vision and inertial measurement units, comprising the following steps: S1: Data acquisition and preprocessing; like Figure 2 , Figure 5 and Figure 8As shown, a user holds a human-computer interaction device (also known as a directional interaction device) and makes pointing movements. For example, the human-computer interaction device can be a spatial pointing device or a handheld controller. The target of the pointing activity is a smart device, such as a smart display terminal or an extended reality (XR) device. During the pointing movement, the image acquisition unit set on the human-computer interaction device captures images in real time to obtain a sequence of environmental images and obtain visual data (i.e., visual pose). At the same time, the inertial measurement unit (IMU) built into the human-computer interaction device collects the angular velocity and acceleration data of the human-computer interaction device in real time to obtain inertial measurement unit data. For example, the image acquisition unit module captures environmental images at 30fps, and the IMU collects motion data of the human-computer interaction device at 100Hz or higher frequency. During the data preprocessing stage, the visual data and inertial measurement unit data are timestamped to establish a unified time reference and ensure data synchronization.

[0040] S2: Data Processing and Fusion; like Figure 2 , Figure 3 and Figure 5 As shown, after the data acquisition and preprocessing stage, the data processing and fusion stage begins, which includes visual data processing, inertial measurement unit data preprocessing, and inertial motion tracking processing. Visual data processing involves performing multi-path operations on acquired image sequences, such as... Figure 9 As shown, based on the motion state or environmental characteristics of the human-computer interaction device, at least one path is dynamically called from multiple preset algorithm paths to generate visual observation data for pointing and positioning.

[0041] The preset multiple algorithm paths include: 1. Feature Point Matching Path: Extract key feature points (such as ORB features) from the current frame image and calculate the feature descriptors of the key feature points; match the feature descriptors of the current frame with the feature descriptors in the preset reference frame or local map to obtain an initial set of matching point pairs; use the Random Sample Consensus Algorithm (RANSAC) to remove mismatched points from the initial set of matching point pairs to obtain valid matching point pairs; based on the valid matching point pairs, use epipolar geometry constraints or the PnP algorithm to calculate the pose change of the human-computer interaction device relative to the reference frame.

[0042] 2. Screen Lock Path: Identify physical screen regions in the image sequence and calculate the pointing offset vector of the center of the physical screen region relative to the optical center of the image in the image coordinate system. When a physical screen region is detected in the image sequence, the screen lock path is activated, and the pointing offset vector is used to correct the absolute pointing deviation. Furthermore, static reference objects around the physical screen region are identified in real time, and spatial topological constraints between the static reference objects and the physical screen region are established. When a physical screen region is missing or occluded in the image, the spatial topological constraints are used to predict and complete the current position of the center of the physical screen region.

[0043] 3. Optical flow tracing path: Select corner points or regions with significant texture gradients in the previous frame as tracking feature points; use the pyramid optical flow algorithm to calculate the corresponding positions of the tracking feature points in the current frame to obtain the pixel displacement vectors of the feature points; construct a reprojection error function based on the pixel displacement vectors, and iteratively solve the rotation and translation motion parameters of the human-computer interaction device between consecutive frames by minimizing the reprojection error.

[0044] 4. Artificial Marker Tracking Path: Detect whether there are preset artificial markers in the current frame image. Artificial markers include specific geometric patterns, coded marks, or actively emitting dot matrices. When an artificial marker is detected, extract the coordinates of the corner point or center point of the artificial marker. Based on the correspondence between the preset physical size of the artificial marker and the image coordinates, use the PnP algorithm to directly calculate the six-degree-of-freedom pose of the human-computer interaction device relative to the artificial marker.

[0045] 5. Line Feature Tracking Path: Extract the line segment features in the current frame image and calculate the geometric parameters of the line segments; match the line segments in the current frame with the line features in the reference frame or local map to construct line feature matching pairs; construct point-line distance constraint or line-line angle constraint equations, and solve the pose changes of the human-computer interaction device through nonlinear optimization to assist in localization in weak texture environments.

[0046] 6. Active spot tracking path (optional): When the human-computer interaction device is equipped with a laser emission module, during the period when the laser emission module is activated and projects a laser spot onto the target interaction area, the laser spot in the visual data is identified and used as the main tracking feature; based on the motion trajectory of the laser spot in the image sequence, the pointing motion of the human-computer interaction device is calculated.

[0047] Dynamic invocation includes: independently, serially, or in parallel fusion invocation of multiple algorithm paths based on the processor load, frame processing latency or available computing power indicators of the human-computer interaction device, motion speed, and the number of effective features, matching inlier rate, optical flow tracking success rate, or reprojection residual of the current image frame.

[0048] For example, when the human-computer interaction device is in a fast-moving state and the image blur is higher than a preset threshold, the inertial measurement unit data is used first for motion tracking, and the weights of feature point matching path, screen lock path, optical flow tracking path, manual marker tracking path, line feature tracking path, and active spot tracking path are reduced; when the human-computer interaction device is in a low-speed movement or hovering state, the feature point matching path, screen lock path, optical flow tracking path, manual marker tracking path, line feature tracking path, or active spot tracking path are used first to suppress the zero-bias drift of the inertial measurement unit.

[0049] Inertial measurement unit (IMU) data preprocessing is used to suppress signal errors in IMU data; inertial motion tracking processing is used to calculate at least one of the short-term attitude increment, velocity increment, or displacement increment of the human-machine interaction device based on angular velocity and acceleration data, to obtain inertial motion tracking data.

[0050] Data fusion involves deeply integrating visual observation data with inertial motion tracking data. The validity of the visual observation data is determined by criteria such as: the number of valid features being lower than a preset threshold, the matching inlier rate being lower than a preset threshold, the optical flow tracking success rate being lower than a preset threshold, or the reprojection residual being higher than a preset threshold. The visual observation data is used to correct the cumulative error of the inertial measurement unit (IMU), while the high-frequency data from the IMU is used to compensate for the delay in visual processing. The system's tracking capability is maintained even when the visual observation data fails.

[0051] S3: Coordinate Mapping and Applications; like Figures 6-8 As shown, coordinate mapping maps the pose after data fusion to the coordinate system of the target interaction area, transforming it into cursor coordinates; Then, the cursor coordinate data is transmitted to the smart device in real time through the wireless communication module, so that the smart device can update the cursor position on the target interaction area and realize the pointing interaction between the human-computer interaction device and the control smart device.

[0052] For example, when a smart device has a display screen, such as a smart TV, the target interaction area is the smart device screen; when a smart device does not have a display screen, such as an AR / VR device, the target interaction area is the virtual screen formed within the user's field of vision while wearing the AR / VR device.

[0053] To further improve the system's accuracy, robustness, and power consumption performance, the present invention also includes the following preferred embodiments: 1. Data fusion algorithm based on prediction-correction mechanism Data fusion employs extended Kalman filtering (EKF) or nonlinear optimization (such as sliding window optimization) frameworks.

[0054] The specific steps are as follows: State vector construction: Constructing the system state vector Where p is position, v is velocity, q is attitude quaternion, and b is position. a Accelerometer zero bias, b g This is for zero bias of the gyroscope.

[0055] State recursion (prediction): Using angular velocity and acceleration data collected by the IMU, the state at the previous moment is recursively integrated using the kinematic equations to obtain the prior state estimate for the current moment. and prior covariance matrix .

[0056] Observation Update (Correction): The pose (or feature point coordinates) output by the vision processing module is used as the observation value Z. k Calculate the observed residuals. ,in This is the observation model.

[0057] State Correction: Calculate Kalman Gain K k And use residuals to correct prior states: Simultaneously update the error covariance matrix and utilize the updated zero bias b. a , b g Real-time compensation is performed on subsequent IMU measurements.

[0058] 2. Interactive calibration mechanism To eliminate installation errors or drift accumulated over long periods of operation, the system allows users to perform active calibration before data acquisition.

[0059] Trigger: The user triggers the calibration mode through a specific gesture (such as drawing an "8") or pressing a button.

[0060] Execution: In calibration mode, the user points the device at the center of the screen or a specific calibration point. The system acquires the instantaneous visual feature center or IMU pose at that moment.

[0061] Update: The system will force the current measurement value to be aligned to the reference coordinates of the target interactive area (such as the screen center coordinates (0,0)), calculate and update the coordinate mapping matrix, thereby eliminating systematic bias.

[0062] 3. Geomagnetic confinement and interference immunity When the IMU includes a magnetometer, geomagnetic data is used to suppress long-term drift of the heading angle (Yaw).

[0063] Magnetic interference detection: Real-time calculation of the magnitude and magnetic inclination of geomagnetic data. If the difference between the current measurement value and the local geomagnetic reference vector exceeds a preset threshold, it is determined that there is magnetic interference (such as near a speaker magnet). At this time, the geomagnetic data is automatically discarded, relying only on vision and gyroscope.

[0064] Heading fusion: When the magnetic environment is stable, the heading angle calculated by the magnetometer is used as the observation value and input into the data fusion algorithm to constrain the heading drift of the gyroscope.

[0065] 4. Roller shutter distortion correction To address the rolling shutter effect commonly found in low-cost image sensors—namely, motion distortion caused by different exposure times for different rows of an image—the following correction measures are implemented: Line offset calculation: Using IMU high-frequency angular velocity data that is strictly time-aligned with the image frame, the amount of rotation of the device relative to the beginning of the frame is calculated for each line of the image exposure time.

[0066] Resampling correction: Based on the calculated rotation amount, each row of pixels in the original image undergoes an inverse geometric transformation or resampling to restore the true geometric shape of the image at the same moment. This function is triggered only when the device angular velocity exceeds a high-speed threshold (e.g., 2 rad / s) to save computing power.

[0067] 5. Dynamic frame rate adjustment (power management) Static detection: Real-time monitoring of IMU data; when the angular velocity and acceleration amplitude are below the static threshold for a certain period of time (e.g., 2 seconds), the device is determined to be in a static or slightly moving state.

[0068] Low power mode: Controls the image acquisition unit to reduce the sampling rate (e.g., from 60fps to 5fps) or enter sleep mode, retaining only the IMU low-frequency wake-up detection.

[0069] Quick wake-up: Once the IMU detects that the motion amplitude exceeds the threshold, it immediately triggers the image acquisition unit to switch back to high frame rate mode to ensure the capture of fast motion features.

[0070] Corresponding to the above method, the present invention also provides a human-computer interaction device based on the fusion of vision and inertial measurement units. One embodiment of this device includes a housing, within which a motherboard is disposed. The motherboard is equipped with an image acquisition unit, an inertial measurement unit, a data processing and control unit, a data communication unit, a storage unit, and a power supply unit. The housing has mounting holes and function buttons. Furthermore, a laser emission module can be optionally added to the device as needed.

[0071] The image acquisition unit employs an image sensor assembly or other image acquisition components, such as a low-power CMOS image sensor, and connects to the processing and control unit via a data interface. This unit captures environmental images during the movement of the human-computer interaction device, providing a sequence of visual feature points and an absolute displacement reference for signal error suppression processing. Preferably, the image acquisition unit supports dynamic frame rate adjustment, switching between low frame rates (e.g., 5fps) and high frame rates (e.g., 60fps) based on the motion state to balance power consumption and performance. Depending on the product form, the unit can be designed as a built-in or modular, detachable type. Built-in type: It is directly fixed inside the housing, and the lens is used to view the image through the opening in the housing.

[0072] Modular and detachable: As an independent module, it is connected to the main body of the device via a magnetic interface, buckle or ribbon cable socket, making it easy for users to replace sensor modules of different specifications (such as different field of view and different photosensitive bands).

[0073] An inertial measurement unit (IMU) is a sensor assembly responsible for high-frequency acquisition of the angular velocity and linear acceleration of a human-machine interface (HMI) in three-dimensional space. This is used for initial attitude tracking of the HMI. The IMU includes a three-axis accelerometer, a three-axis gyroscope, and optionally a three-axis magnetometer (geomagnetic meter). It is responsible for high-frequency acquisition of the HMI's angular velocity, linear acceleration, and geomagnetic field strength in three-dimensional space, used for initial attitude tracking and heading correction. This unit can be designed as a built-in or modular, detachable unit. Built-in: Directly fixed inside the outer casing; Modular and detachable: It is an independent module that connects to the human-machine interface device through a magnetic interface, buckle or ribbon cable socket, making it easy for users to replace sensor modules of different specifications (such as different field of view and different photosensitive bands).

[0074] The data processing and control unit employs a high-performance microcontroller (MCU) or SoC as its processor, responsible for running the aforementioned data acquisition, processing, fusion, and coordinate mapping algorithms. Internally or externally, the processor is equipped with a DSP or NPU unit to accelerate visual feature extraction (such as ORB features and optical flow calculation) and matrix operations. Specifically, this unit executes dynamic scheduling of multi-path algorithms, including feature point matching, screen locking, and optical flow tracing, as well as data fusion algorithms based on extended Kalman filtering (EKF) or nonlinear optimization. It also calculates and corrects the integration error of the inertial measurement unit in real time, ensuring that the cursor pointing remains stable and accurate.

[0075] The laser emission module (optional) is used to project a laser spot onto the target interaction area, and works with the image acquisition unit to achieve active spot tracking path to enhance the positioning capability in weak texture environment.

[0076] Data communication unit: It uses a wireless communication module, such as Bluetooth (Bluetooth Low Energy), Wi-Fi or 2.4G radio frequency, to transmit cursor coordinate data to the smart device, and the smart device displays the optical coordinates on its screen.

[0077] The storage unit uses built-in or external non-volatile memory to store program firmware, algorithm model parameters, and calibration data.

[0078] The power supply unit includes a power management module (PMIC) and power components. The power management module is used for power management.

[0079] The power management module is responsible for converting the voltage of the power components to the stable voltage required by each unit and implementing low-power management strategies.

[0080] The power supply component can be either built-in or detachable, depending on the product form factor. When using a built-in structure, the power supply is a fixed-mount rechargeable battery (such as a lithium polymer battery) that is charged via a charging interface on the device (such as USB Type-C) or a wireless charging coil.

[0081] When a modular, detachable structure is used, the power supply components are installed in the battery compartment, and the types of batteries include replaceable batteries (such as AA / AAA alkaline batteries).

[0082] Interactive input components are disposed on the surface of the casing and are used to receive user operation commands. Their specific forms include, but are not limited to: Physical buttons: such as mechanical switches and micro switches, are used for power control or operations with clear tactile feedback.

[0083] Touch module: This includes capacitive touch buttons, a touchpad, or a touch display. When a touch display is used, it can not only realize virtual button functions, but also support gesture operations such as swiping and multi-touch, as well as display device status or interactive menus.

[0084] Multi-dimensional input devices, such as trackballs, joysticks, or opto-pads (OFNs), provide continuous two-dimensional control signals to assist in fine cursor adjustments or page scrolling.

[0085] The interactive input component connects to the data processing and control unit to enable mode switching (such as calibration mode triggering), confirmation / return operations, cursor-assisted control, or custom shortcut functions.

[0086] The supplementary lighting component is located at the front end of the human-computer interaction device. This component includes LED beads and driving circuits, and is connected to the data processing and control unit. When the ambient light intensity is lower than the preset threshold, it will be turned on automatically or manually by the user to provide auxiliary lighting for the image acquisition unit, ensuring that effective visual features can be extracted even in low-light environments.

[0087] Working principle The working principle of this invention is based on a complementary fusion mechanism of vision and inertia. When the user operates and controls the human-computer interaction device to move in three-dimensional space: Motion capture and fusion: The inertial measurement unit (IMU) acquires the device's angular velocity and acceleration at high frequencies (e.g., above 100Hz), and rapidly responds to the user's minute hand movements through integration calculations, ensuring smooth cursor movement and low latency. Simultaneously, the image acquisition unit captures environmental images at lower frequencies (e.g., 30-60fps), extracting natural features, screen borders, or actively projected laser spots to provide the system with an absolute pose reference. The internal processor utilizes algorithms such as the extended Kalman filter (EKF) to deeply fuse the visual "absolute position" with the inertial "relative motion," correcting the IMU's bias and cumulative drift in real time to ensure stable operation even during extended use.

[0088] Adaptive scene switching: The device intelligently schedules the optimal algorithm path (such as enabling active spot tracking under weak textures and enabling screen lock when facing the screen) based on the current environmental characteristics (such as texture richness and lighting conditions) and motion state (such as being stationary or rapidly swinging) to adapt to various complex usage scenarios.

[0089] Coordinate Mapping and Interaction: The precise 3D pose (6DoF) after fusion correction is converted into cursor coordinates (x, y) on a target 2D plane (such as a TV screen) using a coordinate mapping algorithm. This coordinate data is transmitted in real time to a smart device via a wireless communication module, such as... Figure 2 , Figure 8 As shown, the smart device updates the cursor position on the screen accordingly, enabling the user to accurately point to, select, and manipulate icons, thereby achieving a natural human-computer interaction experience of "what you point to is what you get".

Claims

1. A pointing interaction method based on the fusion of vision and inertial measurement units, characterized in that, Includes the following steps: S1: Data acquisition and preprocessing; The image acquisition unit of the human-computer interaction device acquires environmental image sequences in real time to obtain visual data; The inertial measurement unit of the human-computer interaction device collects angular velocity and acceleration data in real time to obtain inertial measurement unit data; The visual data and inertial measurement unit data are timestamped to establish a unified time reference. S2: Data Processing and Fusion; Data processing includes visual data processing, inertial measurement unit data preprocessing, and inertial motion tracking processing; The visual data processing involves performing multi-path operations on the acquired image sequences. Based on the motion state or environmental characteristics of the human-computer interaction device, at least one path is dynamically called from multiple preset algorithm paths to generate visual observation data for pointing and positioning. The inertial measurement unit data preprocessing is used to perform signal error suppression processing on the inertial measurement unit data; The inertial motion tracking processing is used to calculate at least one of the short-term attitude increment, velocity increment, or displacement increment of the human-computer interaction device based on angular velocity and acceleration data, to obtain inertial motion tracking data; The data fusion involves deeply fusing visual observation data with inertial motion tracking data; The validity of visual observation data is determined. The validity determination includes: when the number of valid features is lower than a preset threshold, the matching inlier rate is lower than a preset threshold, the optical flow tracking success rate is lower than a preset threshold, or the reprojection residual is higher than a preset threshold, the visual observation data is determined to be invalid. The system uses visual observation data to correct the cumulative error of the inertial measurement unit (IMU), while using high-frequency data from the IMU to compensate for the delay in visual processing, and maintains the system's tracking capability when visual observation data fails. S3: Coordinate Mapping and Applications; The pose after data fusion is mapped onto the coordinate system of the target interaction area and converted into cursor coordinates; The cursor coordinates are transmitted to the smart device in real time so that the smart device can update the cursor position on the target interactive area.

2. The pointing interaction method according to claim 1, characterized in that: The preset multiple algorithm paths include feature point matching path, screen locking path, optical flow tracing path, manual marker tracing path, and line feature tracing path; the dynamic invocation includes independently, serially, or in parallel fusion invocation of the multiple algorithm paths based on the processor load, frame processing latency or available computing power index, motion speed, and the number of effective features, matching inlier rate, optical flow tracing success rate, or reprojection residual of the current image frame.

3. The pointing interaction method according to claim 2, characterized in that: The data fusion employs a nonlinear filtering algorithm based on a prediction-correction mechanism; it uses an extended Kalman filter or a nonlinear optimization framework, and includes the following steps: Construct a state vector, which includes at least the attitude, velocity, position of the human-computer interaction device and the zero bias error of the inertial measurement unit; The prior state estimate is obtained by recursively calculating the state using the data collected by the inertial measurement unit. The visual observation data is used as the observed values ​​to calculate the observation residuals; The calculated filter gain is used to correct the prior state estimate, update the state vector and error covariance matrix, output the fused pose data, and use the updated zero bias error to correct the subsequent measurement data of the inertial measurement unit.

4. The pointing interaction method according to claim 3, characterized in that: The screen lock path includes the following steps: Identify the physical screen regions in the image sequence and calculate the pointing offset vector of the center of the physical screen region relative to the optical center of the image in the image coordinate system; When a physical screen area is detected in the image sequence, the screen locking path is activated, and the absolute pointing deviation is corrected using the pointing offset vector.

5. The pointing interaction method according to claim 4, characterized in that: The screen lock path also includes the following steps: The system identifies static reference objects around the physical screen area in real time and establishes spatial topological constraints between the static reference objects and the physical screen area. The spatial topological constraints include one or more of the following: relative orientation, relative distance, and relative angle. When the physical screen area is missing or obscured in the image, the current position of the center of the physical screen area is predicted and completed using the spatial topological constraints.

6. The pointing interaction method according to claim 5, characterized in that: The screen lock path also includes the following steps: During the period when the physical screen area completely disappears or the visual observation data fails, the pose is maintained using the zero bias of the calibrated inertial measurement unit; When the physical screen area is re-identified, the screen locking path is used for rapid repositioning to correct the spatial accumulation error caused by pure inertial recursion.

7. The pointing interaction method according to claim 6, characterized in that: The feature point matching path includes the following steps: Extract key feature points from the current frame image and calculate the feature descriptors of the key feature points; Match the feature descriptors of the current frame with the feature descriptors in the preset reference frame or local map to obtain an initial set of matching point pairs; The random sampling consensus algorithm is used to remove mismatched points from the initial set of matching point pairs to obtain valid matching point pairs; Based on the effective matching point pairs, the pose change of the human-computer interaction device relative to the reference frame is calculated using epipolar geometric constraints or the PnP algorithm.

8. The pointing interaction method according to claim 7, characterized in that: The optical flow tracing path includes the following steps: Select corner points or regions with significant texture gradients in the previous frame as tracking feature points; The corresponding position of the tracking feature point in the current frame image is calculated using the pyramid optical flow algorithm to obtain the pixel displacement vector of the feature point; A reprojection error function is constructed based on the pixel displacement vector, and the rotation and translation motion parameters of the human-computer interaction device between consecutive frames are solved iteratively by minimizing the reprojection error. When the number of tracking feature points is lower than a preset threshold or the tracking error is greater than a preset threshold, new tracking feature points are extracted again in the current frame.

9. The pointing interaction method according to claim 8, characterized in that: The manual identification tracking path includes the following steps: Detect whether there are preset artificial markers in the current frame image. The artificial markers include specific geometric patterns, coded marks, or active light-emitting dot matrices. When the artificial marker is detected, the coordinates of the corner point or center point of the artificial marker are extracted; Based on the correspondence between the preset physical size of the artificial marker and the image coordinates, the PnP algorithm is used to directly calculate the six-degree-of-freedom pose of the human-computer interaction device relative to the artificial marker.

10. The pointing interaction method according to claim 9, characterized in that: The line feature tracing path includes the following steps: Extract the features of line segments in the current frame image and calculate the geometric parameters of the line segments; Match the line segments in the current frame with the line features in the reference frame or local map to construct line feature matching pairs; Construct point-to-line distance constraints or line-to-line angle constraints, and solve the pose changes of the human-computer interaction device through nonlinear optimization to assist in localization in weak texture environments.

11. The pointing interaction method according to claim 1, characterized in that: It also includes an interactive calibration step before data acquisition: In response to a calibration command, enter calibration mode; In calibration mode, instantaneous visual data and inertial measurement unit attitude data are acquired when the human-computer interaction device points to a specific position in the target interaction area. The feature center in the instantaneous visual data or the current pose of the inertial measurement unit is forcibly mapped to the reference coordinates of the target interaction area, and the coordinate mapping relationship is updated.

12. The pointing interaction method according to claim 1 or 11, characterized in that: The inertial measurement unit also includes a magnetometer; The data fusion further includes: generating heading constraint information based on geomagnetic data collected by a geomagnetometer, and inputting the heading constraint information as an observation into the data fusion algorithm to suppress the cumulative drift of the heading angle.

13. The pointing interaction method according to claim 12, characterized in that: The generation of the heading constraint information also includes a magnetic interference detection step: Calculate the magnitude and magnetic inclination of the current geomagnetic data and compare them with the preset geomagnetic reference range or reference vector; When the difference is less than a preset threshold, the magnetic environment is determined to be stable, and the heading constraint information is activated. When the difference exceeds a preset threshold, magnetic interference is determined to exist, the current geomagnetic data is discarded, and heading estimation is performed solely based on vision and gyroscope.

14. The pointing interaction method according to claim 13, characterized in that: The visual data processing also includes an image distortion correction step based on inertial data: Using inertial measurement unit data aligned with image frame time, the angular velocity of the human-computer interaction device during image exposure is calculated; The pixel offset of each row in the image is calculated based on the angular velocity and the shutter readout time of the image acquisition unit. The original image is reverse-sampled or geometrically transformed to eliminate geometric distortion caused by the rolling shutter effect.

15. The pointing interaction method according to claim 14, characterized in that: The image distortion correction step is triggered only when the angular velocity of the human-computer interaction device exceeds a preset high-speed motion threshold.

16. The pointing interaction method according to claim 15, characterized in that: The data acquisition also includes a dynamic frame rate adjustment step: Real-time monitoring of the angular velocity and acceleration amplitude output by the inertial measurement unit; When the amplitude is lower than the preset static threshold for a certain period of time, the image acquisition unit is controlled to enter a low frame rate mode to reduce system power consumption. When the amplitude exceeds the preset motion threshold, the image acquisition unit is immediately triggered to switch to high frame rate mode to capture fast motion features.

17. The pointing interaction method according to claim 16, characterized in that: When the human-computer interaction device is equipped with a laser emission module, the preset multiple algorithm paths also include an active spot tracking path, which includes: During the period when the laser emitting module is activated and projects a laser spot onto the target interaction area, the laser spot in the visual data is identified and used as the main tracking feature. Based on the motion trajectory of the laser spot in the image sequence, the pointing motion of the human-computer interaction device is calculated to maintain the continuity of visual positioning.

18. The pointing interaction method according to claim 1, characterized in that: The preset multiple algorithm paths also include an active spot tracking path; the dynamic invocation includes independently, serially, or in parallel fusion invocation of the active spot tracking path based on the processor load, frame processing latency or available computing power index, motion speed, and the number of effective features, matching inlier rate, optical flow tracking success rate, or reprojection residual of the current image frame.

19. The pointing interaction method according to claim 18, characterized in that: The dynamic invocation also includes: When the human-computer interaction device is in a fast motion state and the image blur is higher than a preset threshold, the inertial measurement unit data is used first for motion tracking, and the weights of the feature point matching path, screen lock path, optical flow tracking path, manual mark tracking path, line feature tracking path and active spot tracking path are reduced. When the human-computer interaction device is in a low-speed motion or hovering state, the feature point matching path, screen locking path, optical flow tracing path, manual mark tracing path, line feature tracing path, or active spot tracing path are preferentially invoked to suppress the zero-bias drift of the inertial measurement unit.

20. The pointing interaction method according to claim 19, characterized in that: Data preprocessing also includes a rolling shutter distortion correction algorithm.