Lightweight space positioning method and device based on binocular six-axis gyroscope
By employing a lightweight spatial positioning method based on binocular six-axis gyroscopes, and utilizing temporal analysis and guided perception processing of inertial data and visual information, spatial anchor point information is filtered and an adaptive positioning update mode is selected. This solves the problems of high computational complexity and inertial error accumulation in existing technologies, achieving high-precision, low-energy real-time positioning.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIANGSU BLACK MIRROR STONE TECHNOLOGY CO LTD
- Filing Date
- 2026-01-05
- Publication Date
- 2026-05-29
AI Technical Summary
Existing visual-inertial positioning technology suffers from high computational complexity and energy consumption in resource-constrained equipment and complex environments. It is also susceptible to changes in lighting and the accumulation of inertial errors, leading to a decline in positioning accuracy and stability.
By acquiring environmental image data and inertial measurement data for time-series analysis, motion description information is generated. Combined with binocular vision for guided perception processing, spatial anchor point information is filtered, and the positioning update mode is selected based on attitude reliability parameters to achieve lightweight real-time attitude capture.
It improves the stability and computational efficiency of high-precision positioning in complex environments, reduces the terminal's computing power and energy consumption requirements, suppresses the accumulation of inertial errors, and improves the reliability and real-time performance of positioning.
Smart Images

Figure CN122108099A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of positioning and navigation technology, and in particular to a lightweight spatial positioning method and device based on a binocular six-axis gyroscope. Background Technology
[0002] With the development of visual inertial positioning technology, spatial positioning schemes based on the fusion of visual sensors and inertial measurement units are widely used in scenarios such as mobile terminals, robot navigation, and augmented reality to achieve posture capture of devices in complex environments.
[0003] In practical engineering deployments, existing visual inertial positioning technology typically requires high-frequency processing of continuously acquired image data and combining it with inertial data to complete state prediction and correction in order to ensure the accuracy and stability of positioning.
[0004] However, in practical engineering applications, existing visual inertial positioning technologies generally require continuous high-frequency processing of the entire image area and adopt fixed data fusion and update strategies, resulting in high computational complexity and strong dependence on terminal computing power and energy consumption. At the same time, visual information is easily affected by changes in lighting, insufficient scene texture, and dynamic interference, while inertial data suffers from error accumulation during long-term operation. Without evaluation and adjustment of positioning reliability, it is easy to cause a decrease in positioning accuracy and unstable results. Therefore, it is difficult to balance computational efficiency, environmental adaptability, and positioning stability in resource-constrained equipment and complex environments. Summary of the Invention
[0005] This application provides a lightweight spatial positioning method and device based on a binocular six-axis gyroscope. The core of this method lies in: acquiring environmental image data and inertial measurement data during the operation of the device being positioned; performing time-series analysis on the inertial data to generate motion description information; and using the motion description information to perform guided perception processing on the environmental image data to determine the visual perception area. Within the visual perception area, spatial anchor point information is selected by combining environmental features and geometric structural features. The spatial anchor point information is matched with the motion description information to calculate attitude reliability parameters, and a positioning update mode is selected based on these parameters to update the spatial positioning information of the device being positioned in real time. Using the initial position and attitude as a reference, the continuously updated spatial positioning information is recorded, ultimately realizing a lightweight real-time attitude capture technology based on the fusion of inertial sensing and binocular vision. This achieves high-precision, low-computational-load spatial positioning and attitude tracking through inertial data prediction and visual anchor point matching.
[0006] To achieve the above objectives, this application adopts the following technical solution: This application provides a lightweight spatial positioning method based on a binocular six-axis gyroscope, the method comprising: During the operation of the device being located, spatial perception data of the device being located within the positioning cycle is acquired, and the spatial perception data includes environmental image data and inertial data; The inertial data is subjected to time series analysis to extract the attitude change trend of the positioned device over a set time period, and the attitude change trend is converted into corresponding motion description information. Based on the motion description information, guided perception processing is performed on the environmental image data to map the posture change trend of the located device to the binocular vision image space and determine the corresponding visual perception area. Within the visual perception area, the environmental image data is jointly identified by environmental features and geometric structure features to filter out corresponding spatial anchor point information, which is used as a spatial reference for the positioned device in the environment. The spatial anchor point information is matched with the motion description information to obtain attitude change information, and the corresponding attitude reliability parameter is calculated based on the attitude change information. Based on the attitude confidence parameter, select the corresponding positioning update mode, and update the spatial positioning information of the positioned device based on the positioning update mode and the attitude change information. Using the initial position and attitude of the device being located as a spatial reference starting point, the spatial positioning information obtained in the positioning update mode is recorded to form the coordinates and attitude information of the device being located at the corresponding time.
[0007] In some possible implementations, the environmental image data includes left-eye image data and right-eye image data. The step of performing guided perception processing on the environmental image data based on the motion description information, mapping the attitude change trend of the located device to the binocular visual image space, and determining the corresponding visual perception region includes: The direction and magnitude of the attitude change of the positioned device during the positioning cycle are obtained based on the motion description information. Based on the direction and magnitude of the posture change, the corresponding pixel offsets in the left and right eye image data are predicted respectively, and the preliminary perception area is determined based on the pixel offsets. The left and right eye image data within the initial perception area are locally processed to form the corresponding visual perception area.
[0008] In some possible implementations, the step of predicting the corresponding pixel offsets in the left and right eye image data based on the direction and magnitude of the pose change, and determining the preliminary perception region based on the pixel offsets, includes: In the left eye image data, the corresponding left eye pixel offset is calculated based on the direction and magnitude of the posture change, and a local area is calibrated as the initial perception area of the left eye based on the left eye pixel offset. In the right eye image data, the corresponding right eye pixel offset is calculated based on the direction and magnitude of the posture change, and a local area is calibrated as the initial perception area of the right eye based on the right eye pixel offset. The preliminary perception areas of the left eye and the right eye are merged to form a preliminary perception area.
[0009] In some possible implementations, the joint identification of environmental features and geometric structural features of the environmental image data within the visual perception area to filter out corresponding spatial anchor point information includes: Within the visual perception area, feature extraction is performed on the left eye image data and the right eye image data respectively to obtain the corresponding local visual features and structural features; The local visual features and the structural features are correlated and evaluated to select anchor feature points that can be repeatedly identified under different viewpoints and time periods; The anchoring feature points are used as spatial anchor point information.
[0010] In some possible implementations, the step of matching the spatial anchor point information with the motion description information to obtain attitude change information, and calculating the corresponding attitude reliability parameter based on the attitude change information, includes: Based on the motion description information, the predicted attitude change state of the positioned device in the current positioning cycle is determined; Based on the predicted attitude change state, the spatial anchor point information is matched and analyzed to obtain the matching result of the spatial anchor point information under the predicted attitude change state. Based on the matching result, the degree of deviation between the spatial anchor point information and the predicted attitude change state is calculated, and attitude change information is formed based on the degree of deviation. Based on the comparison between the degree of deviation and the preset deviation threshold, the attitude change information is evaluated, and corresponding attitude reliability parameters are generated.
[0011] In some possible implementations, calculating the degree of deviation between the spatial anchor point information and the predicted attitude change state based on the matching result, and forming attitude change information based on the degree of deviation, includes: Calculate the expected position of the spatial anchor point information in the image plane or three-dimensional space under the predicted pose change state; The expected position is compared with the observed anchor position to obtain the corresponding deviation value; The deviation values are weighted according to a preset weighting factor to form a deviation index; Attitude change information is generated based on the deviation index.
[0012] In some possible implementations, the positioning update mode includes a first positioning update mode and a second positioning update mode. The step of selecting the corresponding positioning update mode based on the attitude confidence parameter, and updating the spatial positioning information of the positioned device based on the positioning update mode and in conjunction with the attitude change information, includes: If the attitude confidence parameter is greater than or equal to a set threshold, the first positioning update mode is selected; In the first positioning update mode, the attitude change information is used as an initial reference, and the spatial anchor point information and the motion description information are fused and calculated to update the spatial positioning information of the positioned device. If the attitude confidence parameter is less than the set threshold, the second positioning update mode is selected; In the second positioning update mode, the spatial positioning information of the positioned device is updated based on the attitude change information accumulated during the positioning period and with the help of the spatial anchor point information for correction.
[0013] Among some possible implementation methods, the following are also included: When the positioning update mode is switched, the spatial positioning information and corresponding attitude change information output in the positioning update mode before the switch are used as the initial state input of the positioning update mode after the switch. Based on the initial state input, the spatial positioning information is updated under the new positioning update mode.
[0014] In some possible implementations, the inertial data includes acceleration data and angular velocity data. The step of performing time-series analysis on the inertial data to extract the attitude change trend of the positioned device over a set time period, and converting the attitude change trend into corresponding motion description information, includes: Based on the amplitude change of acceleration data within a set time period, a first index is generated, which is used to characterize the linear motion trend. Based on the cumulative change and direction of angular velocity data over a set time period, a second index is generated, which is used to characterize the rotational motion trend. The first and second indicators are combined to form motion description information.
[0015] A lightweight spatial positioning device based on a binocular six-axis gyroscope, the device comprising: a binocular vision acquisition component, a six-axis inertial measurement component, a computing and processing module, and a positioning module; The binocular vision acquisition component is used to acquire environmental image data of the device being located during the positioning period during the operation of the device being located. The six-axis inertial measurement unit is used to acquire the inertial data of the positioned device during the positioning cycle during the operation of the positioned device; The computing processing module includes a first computing layer, a second computing layer, a third computing layer, and a fourth computing layer; The first computing layer is used to perform time-series analysis on the inertial data, extract the attitude change trend of the positioned device over a set time period, and convert the attitude change trend into corresponding motion description information. The second computing layer is used to perform guided perception processing on the environmental image data according to the motion description information, and to map the posture change trend of the located device to the binocular vision image space to determine the corresponding visual perception area. The third computing layer is used to perform joint recognition of environmental features and geometric structure features on the environmental image data within the visual perception area, and filter out the corresponding spatial anchor point information. The spatial anchor point information is used as a spatial reference for the positioning device in the environment. The fourth calculation layer is used to match the spatial anchor point information with the motion description information to obtain attitude change information, and calculate the corresponding attitude confidence parameter based on the attitude change information. The positioning module is used to select the corresponding positioning update mode according to the attitude confidence parameter, and update the spatial positioning information of the positioned device based on the positioning update mode and in combination with the attitude change information. Using the initial position and attitude of the device being located as a spatial reference starting point, the spatial positioning information obtained in the positioning update mode is recorded to form the coordinates and attitude information of the device being located at the corresponding time.
[0016] As can be seen from the above technical solution, this application has the following beneficial effects: 1. This application integrates inertial measurement data with binocular vision information and performs time-series analysis on the inertial data to generate motion description information, enabling the positioned device to achieve continuous and high-precision pose estimation in complex environments, thereby improving the accuracy and stability of spatial positioning.
[0017] 2. This application performs regional processing on environmental images through a guided perception mechanism and dynamically selects the positioning update mode in combination with attitude confidence parameters to achieve adaptive adjustment of positioning computation, thereby reducing terminal computing power and energy consumption requirements while ensuring positioning accuracy and improving deployment feasibility on resource-constrained devices.
[0018] 3. This application achieves real-time correction and continuous updating of attitude changes by matching spatial anchor point information with motion description information, enabling the positioned device to effectively suppress the accumulation of inertial errors during long-term operation, improve positioning reliability, and ensure the stability and real-time performance of the spatial positioning process. Attached Figure Description
[0019] The present application will be further described below with reference to the accompanying drawings.
[0020] Figure 1 A flowchart illustrating a lightweight spatial positioning method based on a binocular six-axis gyroscope provided in this application; Figure 2 An example diagram of a lightweight spatial positioning device based on a binocular six-axis gyroscope provided in this application. Detailed Implementation
[0021] The terms "first," "second," and "third," etc., used in this application specification, claims, and drawings are used to distinguish different objects, not to limit a specific order.
[0022] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0023] To ensure clarity and conciseness in the description of the following embodiments, a brief introduction to the related technologies is given first: Visual-inertial spatial positioning technology is a spatial positioning method that integrates data from visual sensors and inertial measurement units (IMUs). Its basic principle is to acquire environmental images using a camera, analyze the features or texture information in the images to estimate the device's motion relative to the environment, and simultaneously use acceleration and angular velocity data recorded by the IMU to supplement and predict the motion state. By fusing visual observation and inertial measurement data, pose estimation of the device can be achieved during continuous motion, and the impact of single sensor errors on the positioning results can be reduced to some extent.
[0024] Research has shown that while this method can reduce the impact of single-sensor errors to some extent, it still faces several limitations in practical engineering applications. This method often requires high-frequency processing of the entire image region and the use of a fixed data fusion strategy, resulting in high computational load and energy consumption, making it difficult to run continuously on devices with limited computing power. Simultaneously, visual information is sensitive to changes in lighting, insufficient scene texture, and dynamic interference; inertial data is prone to accumulating drift errors over long-term operation; and the lack of dynamic evaluation and adjustment mechanisms makes the positioning results susceptible to instability or accuracy degradation in complex environments. Based on the above analysis, it is evident that existing technologies struggle to achieve a balance between computational efficiency, environmental adaptability, and positioning stability.
[0025] Example 1: To solve the above problems, this application provides a lightweight spatial positioning method based on a binocular six-axis gyroscope. Please refer to [link to example]. Figure 1 .
[0026] S101: During the operation of the device being located, acquire spatial perception data of the device being located within the positioning cycle.
[0027] During the operation of the device being located, spatial perception data of the device within the positioning cycle is acquired. The spatial perception data includes environmental image data and inertial measurement data. The above data is stored in time sequence to ensure the continuity and traceability of subsequent processing.
[0028] To ensure clarity and conciseness in the description of the following embodiments, a brief introduction to the relevant terms is given first: The device being located refers to a mobile terminal, drone, or robotic platform equipped with a binocular vision acquisition unit and a six-axis inertial measurement unit (IMU), used for real-time spatial positioning and navigation. The binocular vision acquisition unit includes two synchronously triggered camera modules that acquire left and right eye image data. The six-axis IMU includes a three-axis accelerometer and a three-axis gyroscope, used to measure linear acceleration and angular velocity.
[0029] The positioning cycle refers to the time interval between continuous acquisition and processing of spatial sensing data. It can be a fixed duration (e.g., 50ms to 200ms) or dynamically adjusted according to the actual movement speed or environmental complexity to ensure positioning accuracy and real-time performance.
[0030] Spatial perception data includes environmental image data and inertial measurement data. Environmental image data is used to extract visual features and geometric structures, while inertial measurement data is used to capture motion trends and help predict device attitude.
[0031] Environmental image data: A continuous sequence of images acquired by the binocular camera in each positioning cycle, including high-resolution grayscale or color images, used for subsequent feature extraction, matching, and spatial anchor point identification.
[0032] Inertial measurement data: Acceleration and angular velocity data acquired by the six-axis inertial measurement unit in each positioning cycle are used to generate attitude change trends and provide motion prior information for visual feature matching.
[0033] In some possible implementation methods, comprehensive perception of the surrounding environment and the motion state of the positioned device can be achieved by simultaneously acquiring visual and inertial information. Environmental image data reflects the structural information and texture features of the surrounding space, which can be used for feature point extraction, scene matching, and spatial geometric analysis; the inertial measurement unit provides acceleration and angular velocity data, which can be used to predict short-term motion trends and attitude changes, providing prior motion constraints for visual information.
[0034] Visual and inertial data are stored and calibrated synchronously according to time series, enabling the association of each frame with its corresponding motion state in subsequent computational processing. This time-series processing ensures the accurate generation of motion description information and supports the dynamic selection of guided visual perception regions. By mapping the attitude change trend of the positioned device to the binocular image space, feature analysis and anchor point recognition are performed only on key regions, thereby significantly reducing computational load and improving processing efficiency.
[0035] In terms of binocular vision acquisition, simultaneous acquisition of left and right eye image data enables stereo vision depth estimation, thereby accurately generating spatial anchor point positions in three-dimensional space. These spatial anchor points are matched and verified in conjunction with motion description information to calculate attitude reliability parameters, providing a basis for dynamically selecting the positioning update mode.
[0036] S102, perform time-series analysis on the inertial data to obtain motion description information.
[0037] During the operation of the positioned device, the collected inertial measurement data is analyzed over time to extract the attitude change trend of the positioned device within a set time period. The attitude change trend is then converted into corresponding motion description information to provide information for subsequent visual perception area selection and spatial anchor point matching.
[0038] Temporal analysis refers to processing and calculating inertial data according to the order of acquisition time to reflect the motion change trend of the device over a continuous time period. Attitude change trend refers to the variation of the rotation angle and direction of motion of the positioned device over a continuous time period, which can be used to predict short-term motion states. Motion description information is index data that jointly represents the linear motion trend and rotational motion trend of the positioned device over a set time period, used to guide visual information processing and spatial anchor point matching.
[0039] In some possible implementations, inertial data is filtered and denoised in chronological order to eliminate sensor noise and external interference, thereby ensuring the accuracy of attitude change trends. Amplitude variation analysis is performed on acceleration data to generate a first index reflecting the linear motion trends in the X, Y, and Z directions. Simultaneously, cumulative change calculations are performed on angular velocity data to generate a second index reflecting the rotation angle around each axis and the direction and magnitude of angular velocity changes. The first and second indices are jointly characterized to form a complete motion description, including linear components, rotational components, and their temporal correspondences, providing constraints for short-term motion prediction.
[0040] The amplitude changes of acceleration data can reflect the linear motion trend of the device in space. By calculating the average rate of change or cumulative displacement in each axis direction, a first indicator of linear motion can be obtained. The cumulative changes of angular velocity data can reflect the rotational trend of the device. By integrating the angular velocities of each axis and combining them with rotational direction information, a second indicator of rotational motion can be obtained. Combining the first and second indicators to construct motion description information can provide short-term constraints for visual processing, allowing feature extraction and spatial anchor matching to be performed only in image regions where motion is predicted, avoiding high-frequency calculations on the entire image.
[0041] S103, based on motion description information, perform guided perception processing on environmental image data to determine the corresponding visual perception area.
[0042] During the operation of the located device, the direction and amplitude of its attitude change within the current positioning cycle are obtained based on motion description information. Based on these directions and amplitudes, the corresponding pixel offsets in the left and right eye image data are calculated, and the corresponding preliminary perception areas are calibrated according to these pixel offsets. Local image processing is then performed on the preliminary perception areas to form the visual perception areas. The size, position, and shape of the visual perception areas can be dynamically adjusted to adapt to different motion amplitudes and rotation directions, thereby ensuring that feature recognition and matching are performed only on key image regions, reducing computational load and improving processing efficiency.
[0043] The visual perception region refers to a local image region calibrated based on motion description information, covering the area most likely to contain key visual features during the movement of the located device, and is used to efficiently extract environmental information.
[0044] In some possible implementations, based on the calculated pose change information, corresponding pixel offsets are predicted in the left and right eye images, respectively. The calculation of pixel offsets can be combined with camera intrinsic parameters or binocular baseline distance to map spatial motion onto the image plane, obtaining the horizontal and vertical pixel movement ranges. Based on the pixel offsets, local rectangular or elliptical regions are marked in the left and right eye images as preliminary sensing regions to cover key areas that may contain moving targets or feature points. The size of the preliminary sensing region can be dynamically adjusted according to the current movement speed and rotation amplitude of the device being located; for example, the region size is increased when the movement amplitude is large to ensure feature coverage, and the region is reduced when the movement amplitude is small or the rotation is slow to reduce the computational load.
[0045] Local image processing is performed on the initial perception region, including image pyramid construction, multi-scale feature extraction, region enhancement and denoising, and geometric correction and alignment of the left and right eye images. The final visual perception region is formed by fusing the initial perception regions of the left and right eyes. Feature points extracted within this region are used for subsequent spatial anchor point generation and matching analysis. The position, size, and shape of the visual perception region can be dynamically updated with device movement, ensuring that the processed area always covers key visual information. Furthermore, a weighted strategy or priority ranking can be used to prioritize high-confidence regions for further improvement in processing efficiency.
[0046] In some possible implementations, the visual perception region generation process can incorporate historical frame information to achieve temporal continuity and feature tracking. By predicting and matching feature points of the visual perception region from the previous frame, and using successfully matched points as references for the next frame, motion compensation and feature tracking can be achieved, thereby reducing mismatches and drift accumulation, and improving the accuracy of spatial anchor point recognition. Simultaneously, by applying different resolutions to image regions of different scales, the computational load can be further reduced while maintaining positioning accuracy.
[0047] S104, within the visual perception area, jointly identify environmental image data and filter out the corresponding spatial anchor point information.
[0048] Specifically, within the visual perception area, the left and right eye image data are processed separately to extract corresponding local visual and structural features. The extracted local visual and structural features are then evaluated for correlation. Through stereo matching and continuous time frame tracking, feature points that can be stably and repeatedly identified under different viewpoints (left and right eyes) and different time periods (multiple consecutive frames) are selected from the evaluated features and determined as anchor feature points, serving as the output spatial anchor point information.
[0049] In the above embodiments, local visual features refer to descriptors extracted from an image that characterize the appearance information of local pixel blocks.
[0050] Structural features refer to information extracted from an image that characterizes the macroscopic geometric contours of a scene. Examples include edge segments extracted by a line segment detector and planar regions extracted by a planar detection algorithm.
[0051] Association evaluation refers to the process of matching, verifying and judging the consistency of various features extracted from left and right eye images and possible historical frames. It aims to eliminate erroneous matches (outside points) and confirm those reliable feature points that are consistent in both three-dimensional space and time series.
[0052] Anchored feature points refer to feature points that, after the aforementioned association evaluation, are confirmed to be geometrically stable and visually significant. They are a concrete manifestation of spatial anchor point information, possessing the ability to be rediscovered and tracked in subsequent frames, providing reliable observation constraints for localization.
[0053] Spatial anchor information is a set of information consisting of a group of anchored feature points, their associated three-dimensional spatial location (or depth), feature descriptors, and their semantic or structural categories.
[0054] In some possible implementations, motion description information is used to map the pose change trend of the device being located to the binocular image space, thereby determining the visual perception region in the image. Feature extraction and matching are performed only within the visual perception region, reducing computational and memory consumption, achieving lightweight processing, and making it suitable for resource-constrained devices.
[0055] Within the visual perception area, local visual features and geometric structural features are extracted from the left and right eye image data, respectively. Specifically, local visual features are extracted through corner detection and texture descriptor generation. Simultaneously, line segments, planes, or curved surfaces within the area are fitted, and local geometric descriptors are generated to form structural features. Local visual features and structural features are jointly correlated and evaluated, and a candidate set of feature points is selected based on spatial geometric consistency and the correspondence between texture and structure. Based on the candidate set, stereo matching is performed to calculate the 3D position, and the feature point positions are tracked over multiple consecutive positioning cycles. Time-series verification is used to eliminate feature points that appear only in a single frame or have excessive positional drift, ensuring that the selected feature points can be stably identified across different viewpoints and time periods. The final spatial anchor point information includes 3D spatial coordinates, 2D pixel positions, timestamps, visual feature descriptors, and structural feature descriptors, which can serve as a reference for the positioned device in the environment, providing reliable data support for subsequent attitude change calculations and spatial positioning information updates.
[0056] S105, match the spatial anchor point information with the motion description information to obtain the attitude change information, and calculate the corresponding attitude reliability parameters based on the attitude change information.
[0057] The system uses motion description information to predict the attitude change state of the located device within the current positioning cycle, including the device's rotation angle and linear displacement direction. The predicted attitude change state is applied to spatial anchor point information. By comparing the predicted position with the actual observed spatial anchor point position, a matching result is obtained, and a deviation value is calculated. The deviation values are weighted according to preset weights to generate a deviation index. Based on the comparison result of the deviation index and a preset deviation threshold, the attitude change information is evaluated, and an attitude reliability parameter is output.
[0058] Among them, attitude change information characterizes the difference between the actual observed position of the spatial anchor point and the predicted position based on motion description, reflecting the deviation between the actual attitude change and the prediction of the positioned device within the current positioning cycle. The attitude reliability parameter is a value calculated based on the deviation index, used to quantify the reliability of the predicted attitude change and provide a decision-making basis for subsequent positioning updates.
[0059] In some possible implementations, spatial anchor point information is matched with motion description information to quantify and reliably assess the attitude changes of the located device. Based on the motion description information, the attitude change state of the located device within the current positioning cycle is predicted, including rotation angle, pitch and yaw directions, and linear translation vector. The predicted attitude changes are applied to the spatial anchor point information to calculate the expected position of each anchor feature point in the image plane and 3D space. These expected positions are matched and analyzed with the actual observed anchor feature point positions within the visual perception area, and the rotational deviation, translational deviation, and 3D position deviation are obtained by comparing the deviations between the two.
[0060] During deviation calculation, a weighting factor is introduced to weight and merge different types of deviations (such as rotation angle deviation, translation distance deviation, and depth deviation) according to their importance, forming a deviation index. Matching analysis can employ a multi-anchor-point weighted averaging strategy, which statistically fuses the deviation results from multiple spatial anchor points, eliminating outliers and errors caused by transient occlusion to ensure the reliability of the obtained attitude change information. Based on the comparison between the deviation index and a preset deviation threshold, an attitude reliability parameter is generated. This parameter quantifies the degree of matching between the predicted attitude and the actual attitude, providing a basis for selecting subsequent positioning update modes. For example, when the attitude reliability parameter is greater than the set threshold, a lightweight update mode assisted by inertial prediction can be directly adopted; when it is less than the set threshold, a repositioning based on spatial anchor points or an enhanced matching mode can be triggered to ensure positioning accuracy.
[0061] In some possible implementations, a corresponding deviation value is calculated for each spatial anchor point, and the deviation value is compared with a preset deviation threshold one by one. When the deviation value corresponding to a certain spatial anchor point is less than or equal to the preset deviation threshold, it is determined that the spatial anchor point has a high consistency with the predicted attitude change state, and the spatial anchor point is recorded as a valid matching anchor point; when the deviation value is greater than the preset deviation threshold, it is determined that the spatial anchor point has a risk of unreliable matching, and the spatial anchor point is recorded as a low-confidence anchor point or an abnormal anchor point.
[0062] Based on this, the overall attitude change information is comprehensively evaluated according to the proportion of effective matching anchor points in all spatial anchor point information, the statistical distribution characteristics of the corresponding deviation values, and the weighted results of the deviation values of each anchor point, thereby generating an attitude reliability parameter. The attitude reliability parameter quantitatively reflects the reliability of the attitude change information relative to the predicted attitude change state within the current positioning cycle. Its value is inversely related to the degree of deviation; that is, the smaller the deviation and the more effective matching anchor points, the higher the generated attitude reliability parameter. Conversely, when the overall deviation increases or the number of effective matching anchor points is insufficient, the generated attitude reliability parameter decreases accordingly.
[0063] In some possible implementations, the attitude confidence parameter can be represented as a continuous numerical form for subsequent comparison with a set threshold to trigger different positioning update modes. Alternatively, the attitude confidence parameter can be discretized into multiple confidence levels to distinguish between high-confidence, medium-confidence, and low-confidence attitude change states. Through this evaluation mechanism, the reliability of attitude change information can be effectively reflected without introducing additional complex calculations, providing a basis for the adaptive selection of subsequent positioning update strategies.
[0064] In some possible implementations, time-series constraints can be introduced. By tracking anchor point matching results over multiple consecutive positioning cycles, the cumulative deviation trend can be calculated, and pose change information can be smoothed and filtered to further reduce the impact of single-frame anomalies or dynamic occlusion on pose assessment. The deviation calculation is then processed hierarchically by combining visual and structural features at different scales.
[0065] S106. Based on the attitude reliability parameters, select the corresponding positioning update mode and update the spatial positioning information of the positioned device in conjunction with the attitude change information.
[0066] Based on the evaluation results of the attitude reliability parameters, the corresponding positioning update mode is adaptively selected, and the spatial positioning information of the positioned device is updated in combination with the attitude change information.
[0067] Specifically, the attitude confidence parameter is compared with a set threshold. The attitude confidence parameter is used to characterize the consistency between attitude change information and spatial anchor point information within the current positioning cycle. When the attitude confidence parameter is greater than or equal to the set threshold, the current attitude estimation result is determined to be reliable, and the first positioning update mode is selected; when the attitude confidence parameter is less than the set threshold, the current attitude estimation is determined to have uncertainty, and the second positioning update mode is selected.
[0068] The first positioning update mode is a lightweight fusion update method. In this mode, attitude change information is used as the initial reference state input, and spatial anchor point information and motion description information are fused and calculated based on this. Through constraint optimization or state estimation, the spatial position and attitude parameters of the positioned device within the current positioning cycle are directly updated, thereby achieving low computational overhead and continuous spatial positioning updates. The second positioning update mode is an enhanced update method based on historical accumulated information. In this mode, it does not directly rely on attitude change information from a single cycle, but integrates and analyzes attitude change information accumulated over multiple frames within the positioning cycle, and uses spatial anchor point information to assist in correcting the accumulated attitude changes, thereby suppressing error propagation caused by sensor drift, visual interference, or environmental changes, and thus obtaining more stable spatial positioning results.
[0069] When switching positioning update modes, to ensure the continuity and stability of spatial positioning information, the spatial positioning information output from the previous positioning update mode and the corresponding attitude change information are used as the initial state input for the new positioning update mode. This prevents abrupt errors from being introduced during the switching process and avoids jumps in positioning results. Through this method, the positioning update strategy can be adaptively adjusted under different environmental conditions and different attitude confidence levels, achieving a balance between positioning accuracy, robustness, and computational efficiency.
[0070] For example, when the device being located is in an indoor environment with a clear structural outline and stable lighting conditions, if the calculated attitude reliability parameter is greater than a set threshold for multiple consecutive positioning cycles, it is determined that there is a high degree of consistency between the current attitude change information and the spatial anchor point information, thus selecting the first positioning update mode. In this first positioning update mode, the predicted attitude change state obtained based on motion description information is used as the initial reference state, and the initial reference state and the successfully matched spatial anchor point information within the current positioning cycle are used together as constraint inputs to fuse and update the spatial position parameters and attitude parameters of the device being located. Specifically, in the predicted attitude change state, the consistency of the three-dimensional position constraints of the spatial anchor point information is checked, correcting small errors accumulated in the predicted attitude, thereby quickly obtaining the updated spatial positioning information. This update process has a short computation path and stable dependent information, making it suitable for achieving high-frequency, low-latency positioning updates.
[0071] When the device being located is in a scenario with dynamic occlusion, rapid changes in viewpoint, or violent movement, if the attitude reliability parameter falls below a set threshold within a certain positioning cycle or multiple consecutive positioning cycles, the reliability of the attitude change information in the current cycle is deemed insufficient, and the system automatically switches to a second positioning update mode. In this second mode, instead of directly using attitude change information from a single positioning cycle for updates, the system integrates and processes the accumulated attitude change information from multiple consecutive positioning cycles to form a temporally continuous attitude change sequence. Simultaneously, spatial anchor point information is introduced as an external geometric constraint to correct the attitude change sequence, ensuring that the attitude change trend remains smooth and consistent over time. This approach effectively reduces the impact of instantaneous mismatches, sensor noise, or short-term environmental interference on spatial positioning results, preventing abrupt changes in positioning results within consecutive cycles.
[0072] S107, using the initial position and attitude of the device being positioned as the spatial reference starting point, records the spatial positioning information obtained in the positioning update mode to form the coordinates and attitude information of the device being positioned at the corresponding time.
[0073] The initial position and attitude parameters acquired by the device being located at the start of the positioning task are determined as the spatial reference starting point to construct a unified spatial reference coordinate system. The initial position parameters characterize the initial spatial coordinates of the device being located in the reference coordinate system, and the initial attitude parameters characterize the initial orientation state of the device being located in the reference coordinate system. After completing step S106, based on the spatial positioning information of the current positioning cycle output by the selected positioning update mode, the spatial positioning information is regarded as an incremental positioning result relative to the spatial reference starting point, and the incremental positioning results are mapped to the spatial reference coordinate system in chronological order to obtain the absolute coordinates and attitude information of the device being located at the corresponding time.
[0074] Specifically, within each positioning cycle, the position and attitude changes updated in the current cycle are accumulated or combined with the coordinate and attitude information recorded in the previous positioning cycle to form the current coordinate and attitude states. These current coordinate and attitude states are then associated with the corresponding timestamp and written as a positioning record into the positioning trajectory dataset. In this way, the spatial motion of the positioned device is continuously recorded, forming a time-sequential sequence of coordinates and attitudes, which comprehensively describes the movement trajectory of the positioned device in space.
[0075] Example 2: This example provides a lightweight spatial positioning device based on a binocular six-axis gyroscope. Please refer to [link / reference]. Figure 2 .
[0076] The device includes: a binocular vision acquisition component, a six-axis inertial measurement component, a computing and processing module, and a positioning module; The binocular vision acquisition component is used to acquire environmental image data of the device being located during the positioning cycle while the device is in operation; A six-axis inertial measurement unit is used to acquire inertial data of the positioned device during the positioning cycle while the device is in operation; The computation processing module includes a first computation layer, a second computation layer, a third computation layer, and a fourth computation layer; The first computing layer is used to perform time-series analysis on inertial data, extract the attitude change trend of the positioned device over a set time period, and convert the attitude change trend into corresponding motion description information. The second computing layer is used to perform guided perception processing on environmental image data based on motion description information, and to map the posture change trend of the positioned device to the binocular vision image space to determine the corresponding visual perception area. The third computing layer is used to jointly identify environmental features and geometric structure features of environmental image data within the visual perception area, and filter out the corresponding spatial anchor point information. The spatial anchor point information is used as a spatial reference for the positioned device in the environment. The fourth computational layer is used to match spatial anchor point information with motion description information to obtain attitude change information, and calculate the corresponding attitude confidence parameters based on the attitude change information. The positioning module is used to select the corresponding positioning update mode based on the attitude reliability parameters, and update the spatial positioning information of the positioned device based on the positioning update mode and combined with the attitude change information. Using the initial position and attitude of the device being located as a spatial reference starting point, the spatial positioning information obtained in the positioning update mode is recorded to form the coordinates and attitude information of the device being located at the corresponding time.
[0077] The foregoing has shown and described the basic principles, main features, and advantages of this application. Those skilled in the art should understand that this application is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of this application. Various changes and modifications can be made to this application without departing from the spirit and scope thereof, and all such changes and modifications fall within the scope of this application as claimed. The scope of protection of this application is defined by the appended claims and their equivalents.
Claims
1. A lightweight spatial positioning method based on a binocular six-axis gyroscope, characterized in that, The method includes: During the operation of the device being located, spatial perception data of the device being located within the positioning cycle is acquired, and the spatial perception data includes environmental image data and inertial data; The inertial data is subjected to time series analysis to extract the attitude change trend of the positioned device over a set time period, and the attitude change trend is converted into corresponding motion description information. Based on the motion description information, guided perception processing is performed on the environmental image data to map the posture change trend of the located device to the binocular vision image space and determine the corresponding visual perception area. Within the visual perception area, the environmental image data is jointly identified by environmental features and geometric structure features to filter out corresponding spatial anchor point information, which is used as a spatial reference for the positioned device in the environment. The spatial anchor point information is matched with the motion description information to obtain attitude change information, and the corresponding attitude reliability parameter is calculated based on the attitude change information. Based on the attitude confidence parameter, select the corresponding positioning update mode, and update the spatial positioning information of the positioned device based on the positioning update mode and the attitude change information. Using the initial position and attitude of the device being located as a spatial reference starting point, the spatial positioning information obtained in the positioning update mode is recorded to form the coordinates and attitude information of the device being located at the corresponding time.
2. The method according to claim 1, characterized in that, The environmental image data includes left-eye image data and right-eye image data. The guided perception processing of the environmental image data based on the motion description information maps the attitude change trend of the located device to the binocular visual image space, determining the corresponding visual perception region, including: The direction and magnitude of the attitude change of the positioned device during the positioning cycle are obtained based on the motion description information. Based on the direction and magnitude of the posture change, the corresponding pixel offsets in the left and right eye image data are predicted respectively, and the preliminary perception area is determined based on the pixel offsets. The left and right eye image data within the initial perception area are locally processed to form the corresponding visual perception area.
3. The method according to claim 2, characterized in that, The step of predicting the corresponding pixel offsets in the left and right eye image data based on the direction and magnitude of the posture change, and determining the preliminary perception region based on the pixel offsets, includes: In the left eye image data, the corresponding left eye pixel offset is calculated based on the direction and magnitude of the posture change, and a local area is calibrated as the initial perception area of the left eye based on the left eye pixel offset. In the right eye image data, the corresponding right eye pixel offset is calculated based on the direction and magnitude of the posture change, and a local area is calibrated as the initial perception area of the right eye based on the right eye pixel offset. The preliminary perception areas of the left eye and the right eye are merged to form a preliminary perception area.
4. The method according to claim 2, characterized in that, Within the visual perception area, the environmental image data undergoes joint recognition of environmental features and geometric structure features to filter out corresponding spatial anchor point information, including: Within the visual perception area, feature extraction is performed on the left eye image data and the right eye image data respectively to obtain the corresponding local visual features and structural features; The local visual features and the structural features are correlated and evaluated to select anchor feature points that can be repeatedly identified under different viewpoints and time periods; The anchoring feature points are used as spatial anchor point information.
5. The method according to claim 1, characterized in that, The step of matching the spatial anchor point information with the motion description information to obtain attitude change information, and calculating the corresponding attitude reliability parameter based on the attitude change information, includes: Based on the motion description information, the predicted attitude change state of the positioned device in the current positioning cycle is determined; Based on the predicted attitude change state, the spatial anchor point information is matched and analyzed to obtain the matching result of the spatial anchor point information under the predicted attitude change state. Based on the matching result, the degree of deviation between the spatial anchor point information and the predicted attitude change state is calculated, and attitude change information is formed based on the degree of deviation. Based on the comparison between the degree of deviation and the preset deviation threshold, the attitude change information is evaluated, and corresponding attitude reliability parameters are generated.
6. The method according to claim 5, characterized in that, The step of calculating the degree of deviation between the spatial anchor point information and the predicted attitude change state based on the matching result, and forming attitude change information based on the degree of deviation, includes: Calculate the expected position of the spatial anchor point information in the image plane or three-dimensional space under the predicted pose change state; The expected position is compared with the observed anchor position to obtain the corresponding deviation value; The deviation values are weighted according to a preset weighting factor to form a deviation index; Attitude change information is generated based on the deviation index.
7. The method according to claim 1, characterized in that, The positioning update mode includes a first positioning update mode and a second positioning update mode. The step of selecting the corresponding positioning update mode based on the attitude confidence parameter, and updating the spatial positioning information of the positioned device based on the positioning update mode and in conjunction with the attitude change information, includes: If the attitude confidence parameter is greater than or equal to a set threshold, the first positioning update mode is selected; In the first positioning update mode, the attitude change information is used as an initial reference, and the spatial anchor point information and the motion description information are fused and calculated to update the spatial positioning information of the positioned device. If the attitude confidence parameter is less than the set threshold, the second positioning update mode is selected; In the second positioning update mode, the spatial positioning information of the positioned device is updated based on the attitude change information accumulated during the positioning period and with the help of the spatial anchor point information for correction.
8. The method according to claim 7, characterized in that, Also includes: When the positioning update mode is switched, the spatial positioning information and corresponding attitude change information output in the positioning update mode before the switch are used as the initial state input of the positioning update mode after the switch. Based on the initial state input, the spatial positioning information is updated under the new positioning update mode.
9. The method according to claim 1, characterized in that, The inertial data includes acceleration data and angular velocity data. The step of performing time-series analysis on the inertial data to extract the attitude change trend of the positioned device over a set time period, and converting the attitude change trend into corresponding motion description information, includes: Based on the amplitude change of acceleration data within a set time period, a first index is generated, which is used to characterize the linear motion trend. Based on the cumulative change and direction of angular velocity data over a set time period, a second index is generated, which is used to characterize the rotational motion trend. The first and second indicators are combined to form motion description information.
10. A lightweight spatial positioning device based on a binocular six-axis gyroscope, characterized in that, The device includes: a binocular vision acquisition component, a six-axis inertial measurement component, a calculation and processing module, and a positioning module; The binocular vision acquisition component is used to acquire environmental image data of the device being located during the positioning period during the operation of the device being located. The six-axis inertial measurement unit is used to acquire the inertial data of the positioned device during the positioning cycle during the operation of the positioned device; The computing processing module includes a first computing layer, a second computing layer, a third computing layer, and a fourth computing layer; The first computing layer is used to perform time-series analysis on the inertial data, extract the attitude change trend of the positioned device over a set time period, and convert the attitude change trend into corresponding motion description information. The second computing layer is used to perform guided perception processing on the environmental image data according to the motion description information, and to map the posture change trend of the located device to the binocular vision image space to determine the corresponding visual perception area. The third computing layer is used to perform joint recognition of environmental features and geometric structure features on the environmental image data within the visual perception area, and filter out the corresponding spatial anchor point information. The spatial anchor point information is used as a spatial reference for the positioned device in the environment. The fourth calculation layer is used to match the spatial anchor point information with the motion description information to obtain attitude change information, and calculate the corresponding attitude confidence parameter based on the attitude change information. The positioning module is used to select the corresponding positioning update mode according to the attitude confidence parameter, and update the spatial positioning information of the positioned device based on the positioning update mode and in combination with the attitude change information. Using the initial position and attitude of the device being located as a spatial reference starting point, the spatial positioning information obtained in the positioning update mode is recorded to form the coordinates and attitude information of the device being located at the corresponding time.