A binocular six-axis low cumulative error spatial positioning method and device

CN122083906APending Publication Date: 2026-05-26JIANGSU BLACK MIRROR STONE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JIANGSU BLACK MIRROR STONE TECHNOLOGY CO LTD
Filing Date
2025-12-31
Publication Date
2026-05-26

Smart Images

  • Figure CN122083906A_ABST
    Figure CN122083906A_ABST
Patent Text Reader

Abstract

This application provides a binocular six-axis low-cumulative-error spatial positioning method and apparatus, applied in the field of positioning and navigation technology. When an anomaly is detected in the spatial motion correlation between vision and inertial at continuous time intervals, it is determined that the current moment is in a state of perception history break. The visual pose constraints at the corresponding moment are isolated from the historical visual constraint set and participate in pose calculation only in an independent local visual constraint set, thereby avoiding the continuous accumulation of abnormal visual constraints in the time dimension and interference with inertial pose updates. In the absence of perception history break, the consistency between the historical visual constraint relationship and the inertial pose change sequence and spatial motion self-consistency data is reviewed, and inconsistent visual constraint relationships are dynamically eliminated. This allows the effectiveness of visual constraints in pose fusion to be adaptively adjusted over time, suppressing the amplification process of inertial cumulative error and meeting the requirement of maintaining high-precision positioning under complex environments and long-term operating conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of positioning and navigation technology, and in particular to a binocular six-axis low cumulative error spatial positioning method and device. Background Technology

[0002] As unmanned equipment is increasingly used in complex environments, spatial positioning technology based on the fusion of vision and inertial technology is gradually becoming an important foundation for realizing continuous pose estimation and environmental perception.

[0003] In existing technologies, for complex scenarios such as indoor environments, binocular six-axis spatial positioning typically employs a visual-inertial fusion strategy based on continuous temporal frames. This involves estimating visual pose changes and inertial pose changes separately at consecutive moments, and then jointly optimizing the two types of pose information through a fixed fusion framework. This type of technical solution usually assumes that visual and inertial information remain consistent in terms of time dimension and motion model, and on this basis, historical visual constraints are continuously introduced into subsequent pose calculation processes to improve positioning accuracy and stability.

[0004] However, existing technologies, by accumulating and participating in subsequent pose calculations with visual and inertial pose constraints formed at continuous moments, assume that visual constraints remain valid over time. This leads to instability or abnormal deviations in visual perception results at certain times due to changes in indoor lighting, occlusion, or sudden changes in scene structure. Existing technologies still introduce the corresponding visual constraints along with historical visual constraints into subsequent calculations, causing abnormal visual information to accumulate and participate in optimization over time. Consequently, the pose residuals corresponding to abnormal visual constraints are repeatedly introduced into the fusion process at subsequent moments. During continuous updates, the inertial pose is constantly corrected around visual references that deviate from the true motion state, causing inertial errors to gradually evolve from instantaneous deviations to accumulated deviations over time. This makes the overall pose estimation highly sensitive to initial anomalies, resulting in reduced positioning accuracy and overall trajectory drift during long-term operation or in complex environments. This makes it difficult to meet the requirements of high-precision positioning for long-term consistency and error controllability, further reducing the stability and reliability of spatial positioning in complex conditions.

[0005] To address this, a binocular six-axis spatial positioning method and device with low cumulative error is proposed. Summary of the Invention

[0006] This application provides a binocular six-axis low cumulative error spatial positioning method and device. Its core is that, for complex scenes such as indoor environments, this method innovatively focuses on using angular velocity data and visual data to form a physical consistency verification, thereby entrusting the main task of translation estimation to visual information, thus forming a highly robust positioning strategy. Specifically, in the process of fusing binocular visual positioning and six-axis inertial positioning, a joint judgment mechanism based on parallax temporal stability data, angular velocity direction consistency data, and time offset data is introduced to continuously evaluate the consistency between visual perception and inertial perception in the time dimension. When an anomaly is detected in the spatial motion correlation between vision and inertial at consecutive moments, it is determined that the current moment is in a state of perception history break, and the visual pose constraint at the corresponding moment is isolated from the historical visual constraint set, participating in pose calculation only in an independent local visual constraint set. This avoids the continuous accumulation of abnormal visual constraints in the time dimension and interference with inertial pose updates. In the absence of perception history break, inconsistent visual constraint relationships are dynamically eliminated by reviewing the consistency between historical visual constraint relationships, inertial pose change sequences, and spatial motion self-consistency data. This allows the effectiveness of visual constraints in pose fusion to be adaptively adjusted over time, thereby suppressing the amplification process of inertial accumulation error, reducing the computational power consumption caused by invalid constraints, and improving the stability and reliability of spatial positioning results under complex environments and long-term operating conditions.

[0007] To achieve the above objectives, this application adopts the following technical solution: This application provides a binocular six-axis low cumulative error spatial positioning method, the method may include: During the operation of the device being positioned, binocular image data and six-axis inertial measurement data are acquired at continuous intervals, wherein the six-axis inertial measurement data includes at least angular velocity data; Binocular matching processing is performed on binocular image data at consecutive time points to obtain the disparity distribution results of the corresponding spatial regions. The disparity distribution results are compared at multiple consecutive time points to extract the temporal continuity features of the direction and magnitude of disparity change, and to generate disparity temporal stability data. Directional component analysis is performed on angular velocity data at continuous time intervals to obtain the trajectory of angular velocity direction change in the time dimension, and angular velocity direction consistency data is generated based on the concentration and persistence of the direction change in the trajectory. Alignment processing is performed on the disparity temporal stability data and the angular velocity direction consistency data to generate spatial motion self-consistency data; Based on the binocular image data and the six-axis inertial measurement data, visual pose change sequences and inertial pose change sequences are obtained in chronological order and compared on the same time axis to generate time offset data. Based on the spatial motion self-consistency data and time offset data, determine whether the current moment is in a state of perceived historical breakage. When it is determined that the perception history is in a broken state, the visual constraint relationship formed by the visual pose change sequence corresponding to the current moment is written only into the local visual constraint set that is independent of the historical visual constraint set, and a local pose change sequence is generated. The inertial pose change sequence is updated based on the local pose change sequence to obtain the updated inertial pose change sequence. The spatial pose at the current moment is calculated based on the updated inertial pose change sequence and the local visual constraint relationship, and is used for spatial positioning.

[0008] In some possible implementations, the method may further include: When it is determined that the perception history is not in the state of interruption, the consistency of the visual constraint relationship in the historical visual constraint set is checked, and the participation qualification of the visual constraint relationship that is inconsistent with the inertial pose change sequence or the spatial motion self-consistency data is removed. Based on the historical visual constraints that have not been removed from the participation status, the local visual constraints, and the six-axis inertial measurement data at the corresponding moment, the spatial pose at the current moment is calculated and used for spatial positioning.

[0009] In some possible implementations, the generation of disparity temporal stability data may include: In the binocular image data at consecutive time points, the corresponding disparity distribution results are extracted for the same spatial region, and the disparity distribution results at each time point are correlated in chronological order to form a disparity temporal set. In the disparity temporal set, the direction of change of the disparity distribution results in the spatial dimension is statistically analyzed to obtain the disparity change direction information corresponding to each consecutive time. Consistency determination is made on the disparity change direction information of adjacent time moments, the continuity of the disparity change direction in the time dimension is identified, and disparity direction continuity description data is generated. Based on the disparity time series set, the variation amplitude between disparity distribution results at adjacent times is quantitatively characterized to obtain the evolution information of disparity variation amplitude over time and generate disparity amplitude continuity description data. Based on the disparity direction continuity description data and the disparity amplitude continuity description data, the stability of disparity changes in the time dimension is comprehensively characterized to generate disparity temporal stability data.

[0010] In some possible implementations, generating consistent angular velocity direction data may include: The angular velocity data is decomposed into angular velocity direction components according to three orthogonal directions, and the change trajectory of each direction component in the time dimension is recorded. In the trajectory of the change of the angular velocity direction component, the offset angle of the angular velocity direction at consecutive moments is analyzed, and the concentration of the change of the angular velocity direction is statistically analyzed to obtain the concentration data of the change of direction. Analyze the persistence of the trajectory of angular velocity direction change over a continuous time period, identify the continuous interval of direction change, and generate data on the persistence of direction change; The concentration data of directional change and the persistence data of directional change are comprehensively evaluated to generate angular velocity directional consistency data.

[0011] In some possible implementations, the generation of spatial motion self-consistency data may include: The disparity temporal stability data and the angular velocity direction consistency data are aligned in the time dimension to obtain aligned disparity temporal stability data and aligned angular velocity direction consistency data. For each consecutive moment, the aligned parallax temporal stability data and the aligned angular velocity direction consistency data are correlated and analyzed to calculate the degree of matching between the two in terms of spatial motion trend and direction change. Within a continuous time span, the matching degree at each time point is statistically analyzed and the changing trend of the matching degree within the continuous time span is analyzed to evaluate the consistency and continuity of spatial motion information in the time dimension and obtain the evaluation results. The matching degree and the evaluation results are comprehensively characterized to generate spatial motion self-consistency data.

[0012] In some possible implementations, the generation of time offset data may include: Based on the binocular image data and the six-axis inertial measurement data, respectively, visual pose change sequences and inertial pose change sequences are obtained in chronological order. The visual pose change sequence and the inertial pose change sequence are compared on the same time axis to obtain the pose differences at each consecutive moment. Based on the pose differences at each consecutive moment, time offset data representing the time offset between the visual pose change sequence and the inertial pose change sequence is generated.

[0013] In some possible implementations, determining whether the current moment is in a state of perceived historical discontinuity may include: Based on the spatial motion self-consistency data, a spatial self-consistency judgment result is obtained by comparing it with a preset spatial self-consistency threshold. Based on the time offset data, a time offset determination result is obtained by comparing it with a preset time offset threshold. The spatial self-consistency determination result and the time offset determination result are compared with preset judgment conditions to determine whether the current moment is in a state of perceived historical break. The preset judgment conditions include: When both the spatial self-consistency determination result and the time offset determination result are greater than the corresponding threshold, it is determined that the current moment is in a state of perceived historical discontinuity. If at least one of the spatial self-consistency determination results and the time offset determination results is less than or equal to the corresponding threshold, it is determined that the current moment is not in a state of perceived historical break.

[0014] In some possible implementations, the step of writing the visual constraint relationship formed by the visual pose change sequence corresponding to the current moment only into a local visual constraint set independent of the historical visual constraint set, and generating a local pose change sequence, may include: Extract the visual constraint relationships from the visual pose change sequence corresponding to the current moment; Create a local set of visual constraints that is independent of the historical set of visual constraints; The extracted visual constraint relationships are written into the local visual constraint set; Based on the set of local visual constraints, the local pose changes at consecutive time intervals are calculated to generate a sequence of local pose changes.

[0015] In some possible implementations, the consistency review of visual constraint relationships in the historical visual constraint set, and the removal of participation eligibility for visual constraint relationships inconsistent with the inertial pose change sequence or the spatial motion self-consistency data, may include: Extract the visual constraint relationships and corresponding inertial pose change data and spatial motion self-consistency data from the historical visual constraint set; The historical visual constraint relationship is compared with the inertial pose change sequence at the corresponding moment to obtain the first comparison result; The historical visual constraint relationship is compared with the spatial motion self-consistency data at the corresponding time to obtain a second comparison result. When the first comparison result or the second comparison result is inconsistent, the visual constraint relationship of the consistency of the corresponding inertial pose change sequence or the spatial motion self-consistency data is released from participation eligibility.

[0016] A binocular six-axis low cumulative error spatial positioning device, the device may include an acquisition module and a processing module; The acquisition module includes a visual positioning component and an inertial positioning component, used to acquire binocular image data and six-axis inertial measurement data, wherein the six-axis inertial measurement data includes at least angular velocity data; The processing module includes: The disparity processing unit is used to perform binocular matching processing on binocular image data at consecutive time points, obtain the disparity distribution results of the corresponding spatial region, compare the disparity distribution results at multiple consecutive time points, extract the temporal continuity features of the direction and magnitude of disparity change, and generate disparity temporal stability data. The angular velocity analysis unit is used to analyze the angular velocity data at continuous time intervals, obtain the trajectory of the change of angular velocity direction in the time dimension, and generate angular velocity direction consistency data based on the concentration and persistence of the directional change in the trajectory. A spatial motion self-consistency unit is used to align the parallax temporal stability data with the angular velocity direction consistency data to generate spatial motion self-consistency data. The time offset calculation unit is used to obtain the visual pose change sequence and the inertial pose change sequence based on the binocular image data and the six-axis inertial measurement data respectively in chronological order, and compare them on the same time axis to generate time offset data. The fracture determination unit is used to determine whether the current moment is in a state of sensing historical fracture based on the spatial motion self-consistency data and time offset data. The local visual constraint processing unit is used to write the visual constraint relationship formed by the visual pose change sequence corresponding to the current moment into a local visual constraint set that is independent of the historical visual constraint set when it is determined that the perception history is in a broken state, and to generate a local pose change sequence. An inertial pose update unit is used to update the inertial pose change sequence based on the local pose change sequence to obtain the updated inertial pose change sequence. A spatial positioning calculation unit is used to calculate the current spatial pose based on the updated inertial pose change sequence and the local visual constraint relationship, for spatial positioning. The historical visual constraint verification unit is used to perform a consistency verification of the visual constraint relationships in the historical visual constraint set when it is determined that the perception history is not in a broken state. It removes the participation qualification of visual constraint relationships that are inconsistent with the inertial pose change sequence or the spatial motion self-consistency data, and calculates the spatial pose at the current moment based on the historical visual constraint relationships that have not been removed from the participation qualification, the local visual constraint relationships, and the six-axis inertial measurement data at the corresponding time, for spatial positioning.

[0017] As can be seen from the above technical solution, this application has the following beneficial effects: 1. When this application determines that the current moment is in a state of perceptual history discontinuity, it isolates the visual pose constraints at the corresponding moment from the historical visual constraint set, and only allows them to participate in pose calculation in an independent local visual constraint set. This isolation mechanism can prevent abnormal visual constraints from accumulating continuously over time, prevent interference with inertial pose updates, and ensure that spatial positioning can maintain computational continuity and stability even when visual information is abnormal or discontinuous.

[0018] 2. In the absence of a perceptual history break, this application dynamically eliminates inconsistent visual constraint relationships by performing a consistency check on historical visual constraint relationships, inertial pose change sequences, and spatial motion self-consistency data. This allows the effectiveness of visual constraints in pose fusion to adaptively adjust over time. This mechanism can suppress the amplification process of inertial cumulative errors, reduce the computational burden caused by invalid constraints, and improve the stability and reliability of spatial positioning under complex environments or long-term operating conditions.

[0019] 3. This application utilizes a binocular six-axis low cumulative error spatial positioning device to achieve dynamic evaluation and adaptive intervention of the consistency between visual and inertial information at continuous moments. It integrates visual constraints and inertial pose updates into a unified spatial positioning control system, thereby ensuring the continuity and accuracy of spatial pose estimation, meeting the requirements of high-precision positioning for long-term consistency and error controllability, and improving the stability and reliability of positioning results under complex environments and long-term operating conditions. Attached Figure Description

[0020] The present application will be further described below with reference to the accompanying drawings.

[0021] Figure 1 A flowchart of a binocular six-axis low cumulative error spatial positioning method provided in this application; Figure 2 An example diagram of a binocular six-axis low cumulative error spatial positioning device provided in this application. Detailed Implementation

[0022] The terms "first," "second," and "third," etc., used in this application specification, claims, and drawings are used to distinguish different objects, not to limit a specific order.

[0023] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0024] To ensure clarity and conciseness in the description of the following embodiments, a brief introduction to the related technologies is given first: Visual-Inertial Odometry (VIO) technology is widely used in the spatial positioning of robots, drones and automated equipment. Its core lies in acquiring environmental image data through binocular or monocular vision sensors, while using an inertial measurement unit (IMU) to collect angular velocity, thereby realizing spatial pose estimation and trajectory tracking. VIO technology typically calculates the sequence of position and attitude changes of a device in three-dimensional space by fusing visual pose constraints and inertial pose constraints at consecutive moments, in order to support navigation, positioning and autonomous control functions. Its basic principle is to achieve the positioning and attitude estimation of the device in three-dimensional space by simultaneously acquiring environmental images and inertial measurement data at continuous moments. Visual localization generates visual pose constraints by analyzing the displacement changes of environmental feature points in the image. Inertial localization uses the angular velocity data provided by the IMU to obtain the inertial pose change sequence. Fusion computation then processes the visual pose constraints and the inertial pose change sequence together in the time and space dimensions. Through complementary correction and constraint optimization, high-precision, continuous and robust spatial pose estimation is achieved.

[0025] Research has revealed that existing technologies accumulate and incorporate visual and inertial pose constraints formed over consecutive moments into subsequent pose calculations. By assuming that visual constraints remain valid over time, visual perception results may become unstable or exhibit abnormal deviations at certain moments due to changes in lighting, occlusion, or abrupt changes in scene structure. Existing technologies still introduce the corresponding visual constraints along with historical visual constraints into subsequent calculations, causing abnormal visual information to accumulate over time and participate in optimization. This results in the pose residuals corresponding to abnormal visual constraints being repeatedly introduced into the fusion process. During continuous updates, the inertial pose is constantly corrected around visual references that deviate from the true motion state, causing inertial errors to evolve from instantaneous deviations to accumulated deviations over time. This makes the overall pose estimation highly sensitive to initial anomalies, leading to reduced positioning accuracy and overall trajectory drift during long-term operation or in complex environments. This makes it difficult to meet the requirements of high-precision positioning for long-term consistency and error controllability, further reducing the stability and reliability of spatial positioning under complex conditions.

[0026] Example 1 To address the aforementioned problems, this application provides a binocular six-axis spatial positioning method with low cumulative error. Please refer to [link to relevant documentation]. Figure 1 .

[0027] S1 acquires binocular image data and six-axis inertial measurement data at continuous intervals during the operation of the positioned device.

[0028] During the operation of the device being positioned, binocular image data and six-axis inertial measurement data for spatial positioning are acquired continuously at different times. The six-axis inertial measurement data includes at least angular velocity data.

[0029] To ensure clarity and conciseness in the description of the following embodiments, a brief introduction to the relevant terms is given first: Binocular image data consists of two image frames, one on the left and one on the right, captured by a binocular camera. These frames are used to calculate the disparity information and pose changes of the corresponding spatial regions. Six-axis inertial measurement data refers to the three-axis angular velocity data and three-axis acceleration data collected by the IMU, which are used to describe the motion state and rotational changes of the carrier. Angular velocity data refers to the three-axis angular velocity data acquired by the IMU; A continuous time interval refers to a series of time points arranged in chronological order, with each time point corresponding to the binocular image data and six-axis inertial measurement data collected by the positioning device.

[0030] In some possible implementations, a binocular camera is fixed to the positioning device, and the camera control module automatically captures continuous left and right image frames at a preset frame rate (such as 30 frames / second or higher). An IMU is fixed to the carrier to collect continuous angular velocity data (around three orthogonal axes) and linear acceleration data in real time. The acquisition frequency is synchronized with the frame rate of the binocular camera to ensure that each frame corresponds to a set of inertial data. The collected binocular image data and angular velocity data are stored. The method mainly utilizes angular velocity data.

[0031] S2 processes the binocular image data and the six-axis inertial measurement data respectively to generate disparity temporal stability data and angular velocity direction consistency data.

[0032] Binocular matching processing is performed on binocular image data at consecutive time points to generate disparity distribution results for the corresponding spatial regions. By comparing the disparity distribution results at multiple consecutive time points over time, the evolution of the direction and magnitude of disparity changes over time is identified, thereby extracting the continuous characteristics of disparity in the time dimension and forming disparity temporal stability data. Disparity temporal stability data is used to describe the degree of disparity stability and the trend of change in the same spatial region at consecutive time points. The angular velocity data from continuous six-axis inertial measurement data is decomposed into directions to obtain angular velocity components in each orthogonal direction, and the trajectory of each component over time is statistically analyzed. By analyzing the concentration and persistence of the angular velocity direction shift angle changes over continuous time steps, the consistency characteristics of the angular velocity direction changes are identified, generating angular velocity direction consistency data. Angular velocity direction consistency data is used to reflect the stability and continuity of the angular velocity direction over time in inertial measurement.

[0033] To ensure clarity and conciseness in the description of the following embodiments, a brief introduction to the relevant terms is given first: Parallax temporal stability data refers to a set of data generated by analyzing binocular image data at consecutive time points, used to characterize the trend and stability of parallax in a spatial region over time.

[0034] Angular velocity direction consistency data: refers to a set of data generated by analyzing six-axis inertial measurement data at continuous time intervals, used to characterize the continuity and consistency of the change in angular velocity direction in the time dimension.

[0035] In some possible implementations, for binocular image data at consecutive time points, the spatial region corresponding to each frame can be divided into multiple disparity units, each corresponding to a pixel block or predefined region in the image. For each disparity unit, disparity values ​​at consecutive time points are obtained in chronological order, and the disparity data at different time points are aligned based on pixel correspondence or block matching algorithms. Subsequently, the disparity values ​​at adjacent time points are compared, and the direction (increasing, decreasing, or remaining stable) and magnitude (absolute change or relative rate of change) of disparity changes are statistically analyzed. The change trend of the same disparity unit at consecutive time points is accumulated and averaged. Specifically, a continuity description value for the unit can be generated through a time series stationarity test (such as the enhanced Dickey-Fuller test) to generate a temporal continuity description of the unit. The continuity descriptions of all disparity units are aggregated to form the disparity temporal stability data of the entire spatial region, including the overall change trend, local fluctuation amplitude, and short-term abrupt change information. For continuous six-axis inertial measurement data, the angular velocity data is first decomposed into directional components along three orthogonal directions (X, Y, and Z axes), and each component is continuously recorded on the time axis. For each component, the angular velocity values ​​at adjacent time points are differentially calculated to obtain the amplitude of angular velocity change; the direction of angular velocity change is classified (positive, negative, or no change), and the duration of consistent direction of change within a continuous time period is statistically analyzed. Based on the amplitude and duration of the offset angle of each directional component, the spherical variance or the mean cosine of the direction vector can be used as the concentration of directional change and the continuous interval of directional change. The concentration and continuous interval of the three orthogonal directional components are comprehensively analyzed to generate angular velocity directional consistency data, which is used to quantify the directional stability and continuity of inertial measurement in the time dimension.

[0036] S3 aligns the parallax temporal stability data with the angular velocity direction consistency data to generate spatial motion self-consistency data.

[0037] Aligning the two types of data along the time dimension allows for correlation analysis between the disparity features at each time point and the angular velocity direction features at the corresponding moment. A comprehensive evaluation of the aligned data generates spatial motion self-consistency data, which characterizes the consistency and continuity of spatial motion information in time and direction across consecutive moments.

[0038] To ensure clarity and conciseness in the description of the following embodiments, a brief introduction to the relevant terms is given first: Spatial motion self-consistency data refers to a quantitative indicator that reflects the continuity and consistency of spatial motion state in the time dimension by aligning the parallax temporal stability data and angular velocity direction consistency data at continuous time points on the time axis and then comprehensively analyzing the two types of data.

[0039] Time alignment refers to interpolating or mapping different data sources along the time dimension so that the two types of data can correspond one-to-one at the same time scale.

[0040] Consistency analysis refers to comparing the trend of parallax change after alignment with the characteristics of angular velocity direction change, and evaluating the degree of matching between the two in terms of spatial motion direction, amplitude and trend of change.

[0041] In some possible implementations, for the already generated disparity temporal stability data and angular velocity direction consistency data, continuous time periods can be divided into time windows with fixed time granularity (such as milliseconds or seconds). For each time window, the continuity index and angular velocity direction consistency index of the disparity unit are first interpolated or mapped to nearest neighbors based on the timestamp, so that the two types of data correspond precisely in the time dimension; Joint analysis is performed on the data within each time window, including: Compare the correspondence between the direction of parallax change (increasing / decreasing / stable) and the direction of angular velocity change (positive / negative / no change), and statistically analyze the consistency index of the direction. The relative matching degree is calculated after normalizing the parallax amplitude change and angular velocity offset amplitude, which is used to measure the consistency of motion amplitude. Analyze the maintenance of direction and amplitude within a continuous time window, identify the length of intervals that remain consistent over time, and quantify the continuity index.

[0042] The results of direction matching, amplitude matching, and continuity matching are weighted and synthesized. Specifically, the entropy weight method can be used to calculate the weight of each indicator, and then a weighted summation formula is used to generate the spatial motion self-consistency value corresponding to each time window. The self-consistency values ​​of all time windows are integrated into a time series to obtain the complete spatial motion self-consistency data sequence.

[0043] S4. Based on binocular image data and six-axis inertial measurement data, respectively, visual pose change sequences and inertial pose change sequences are obtained in chronological order and compared on the same time axis to generate time offset data.

[0044] After acquiring binocular image data and six-axis inertial measurement data at consecutive time points, visual pose change sequences are generated based on the binocular image data, and inertial pose change sequences are generated based on the six-axis inertial measurement data. Subsequently, the two types of pose change sequences are compared on the same time axis, and the offset difference between corresponding time points is calculated to generate time offset data, which is used to characterize the degree of synchronization between visual pose change and inertial pose change in the time dimension.

[0045] To ensure clarity and conciseness in the description of the following embodiments, a brief introduction to the relevant terms is given first: Visual pose change sequence: refers to an ordered sequence of visual pose changes at consecutive time points, obtained by binocular image data at consecutive time points through binocular matching, stereo vision measurement, and pose calculation.

[0046] Inertial pose change sequence: refers to an ordered sequence of inertial pose changes at continuous moments, obtained by integrating angular velocity and acceleration information based on six-axis inertial measurement data at continuous moments.

[0047] Time offset data refers to a numerical sequence generated by comparing visual pose change sequences and inertial pose change sequences on the same time axis, used to quantify the time synchronization deviation between the two.

[0048] In some possible implementations, inter-frame matching processing is performed on the binocular image data at consecutive time points to extract the three-dimensional spatial pose corresponding to each frame image, and a visual pose change sequence is constructed in chronological order. For six-axis inertial measurement data, the angular velocity and acceleration data are integrated in time sequence to generate an inertial pose change sequence; Visual pose change sequences and inertial pose change sequences are mapped to a unified time scale. For differences in time intervals or inconsistent sampling on the time axis, linear interpolation or nearest neighbor mapping methods can be used to align the data of each sequence to the same time node.

[0049] After completing the time alignment, for each time node, the difference between the visual pose and the inertial pose is calculated, including position offset and orientation offset, to obtain the time offset value corresponding to each time node. The offset values ​​of all time nodes are integrated in chronological order to form a complete time offset data sequence.

[0050] S5, based on spatial motion self-consistency data and time offset data, determines whether the current moment is in a state of perceived historical rupture.

[0051] After acquiring spatial motion consistency data and time offset data, threshold comparison and comprehensive judgment are performed on the two types of data to determine the continuity state of visual and inertial information in spatial pose perception at the current moment. When both spatial motion consistency and time offset meet the preset threshold conditions, it is determined that the current moment is not in a state of perception history discontinuity; otherwise, when either or both exceed the threshold, it is determined that the current moment is in a state of perception history discontinuity.

[0052] To ensure clarity and conciseness in the description of the following embodiments, a brief introduction to the relevant terms is given first: Perceptual history discontinuity state: refers to the state in which the spatiotemporal information of the visual pose change sequence and the inertial pose change sequence are inconsistent at continuous moments, resulting in the interruption of the continuity of spatial pose perception.

[0053] Spatial self-consistency judgment result: refers to the judgment result obtained by comparing the spatial motion self-consistency data with the preset spatial self-consistency threshold, which is used to characterize whether the spatial motion data meets the requirements in terms of continuity and consistency.

[0054] Time offset determination result: refers to the judgment result obtained by comparing the time offset data with the preset time offset threshold, which is used to characterize the synchronization of the visual pose change sequence and the inertial pose change sequence in the time dimension.

[0055] In some possible implementation methods, the spatial motion self-consistency data is first compared with a preset spatial self-consistency threshold time-by-time to obtain the spatial self-consistency judgment result at each time node. The judgment result can be represented by a binary identifier (e.g., 0 indicates that the threshold has not been reached, and 1 indicates that the threshold has been reached or exceeded) or a continuous quantized score (e.g., the value range of 0 to 1 indicates the degree of self-consistency compliance).

[0056] Simultaneously, the time offset data is compared with a preset time offset threshold to obtain the time offset determination result for each time node. By analyzing the continuity and abnormal peaks of the time offset, delays or abnormal asynchrony intervals between visual and inertial pose changes can be identified.

[0057] After determining spatial consistency and temporal offset, the two sets of results are logically combined or weighted to form a comprehensive judgment result. Logical combination can use an AND operation: if both the spatial consistency and temporal offset results exceed a threshold, the system is judged as experiencing a historical break; otherwise, it is judged as non-break. Weighting assigns different weights to spatial consistency and temporal offset. These weights can be preset based on statistical analysis of historical data or determined using a simple averaging method. Combined with continuity analysis, this yields a smoother and more robust break judgment result.

[0058] S6, perform the corresponding operation based on the judgment result.

[0059] After obtaining the current moment's perceived historical fracture state judgment result, different data processing and constraint update operations are performed based on the judgment result.

[0060] S601, when the judgment result indicates that the current time is in a state of perceptual history break, an independent set of local visual constraints is generated and the local pose change sequence is updated.

[0061] To ensure clarity and conciseness in the description of the following embodiments, a brief introduction to the relevant terms is given first: Local visual constraint set: refers to the set of constraint information independent of the historical visual constraint set, used to record the visual constraint relationships under the state of perceptual history rupture.

[0062] Local pose change sequence: refers to a continuous pose change sequence calculated based on a set of local visual constraints, used to update the inertial pose change sequence.

[0063] In some possible implementations, visual constraint relationships are extracted from the visual pose change sequence corresponding to the current moment; Create an independent set of local visual constraints and write the extracted visual constraint relationships into it; Based on the set of local visual constraints, the local pose changes at consecutive time steps are calculated to generate a sequence of local pose changes. The local pose change sequence is used to update the inertial pose change sequence to ensure the continuity and stability of spatial positioning under fractured conditions.

[0064] When this application determines that the current moment is in a state of perceptual history discontinuity, it isolates the visual pose constraints at the corresponding moment from the historical visual constraint set, and only allows them to participate in pose calculation in an independent local visual constraint set. This isolation mechanism can prevent abnormal visual constraints from accumulating continuously over time, prevent interference with inertial pose updates, and ensure that spatial positioning can maintain computational continuity and stability even when visual information is abnormal or discontinuous.

[0065] S602, when the judgment result indicates that the current time is not in a state of perceptual history break, the consistency of the historical visual constraint set is checked and the spatial pose is calculated by fusion.

[0066] To ensure clarity and conciseness in the description of the following embodiments, a brief introduction to the relevant terms is given first: Historical visual constraint set review: refers to checking the consistency of historical visual constraint relationships and removing constraints that are inconsistent with the inertial pose change sequence or spatial motion self-consistency data from the eligibility to participate.

[0067] In some possible implementations, a consistency check is performed on the visual constraint relationships in the historical visual constraint set; Extract inertial pose change data and spatial motion self-consistency data corresponding to each constraint relationship in the historical visual constraint set at each moment; By comparing historical visual constraints with the corresponding inertial pose change sequence and spatial motion self-consistency data, inconsistent constraints are removed to determine eligibility for participation. Based on the historical visual constraints that have not been removed from the participation status, the local visual constraints at the current moment, and the six-axis inertial measurement data at the corresponding moment, the spatial pose at the current moment is calculated by fusing the data to ensure the accuracy and continuity of spatial positioning.

[0068] It should be noted that the specific algorithms involved in each of the above steps are all well-known technologies in this field.

[0069] This application, without any perceptual history disruption, dynamically eliminates inconsistent visual constraint relationships by performing a consistency check on historical visual constraint relationships, inertial pose change sequences, and spatial motion self-consistency data. This allows the effectiveness of visual constraints in pose fusion to adaptively adjust over time. This mechanism can suppress the amplification process of inertial cumulative errors, reduce the computational burden caused by invalid constraints, and improve the stability and reliability of spatial positioning under complex environments or long-term operating conditions.

[0070] This application utilizes continuous angular velocity data and visual parallax data to construct a physical consistency check, entrusting the main task of translation estimation to visual information, while angular velocity information is used for directional consistency constraints and spatial motion self-consistency calculation. This design verifies pose changes through cross-modal features, enhancing the robustness of the positioning system to abnormal motion or sensor noise, while reducing the impact of single-modal errors on overall spatial pose estimation.

[0071] Example 2 This application provides a binocular six-axis low cumulative error spatial positioning device. The module functions of this device correspond to the specific implementation steps in Embodiment 1, including an acquisition module and a processing module. Please refer to [link to relevant documentation]. Figure 2 .

[0072] The acquisition module includes a visual positioning component and an inertial positioning component, used to acquire binocular image data and six-axis inertial measurement data, the six-axis inertial measurement data including at least angular velocity data; The processing module includes: The disparity processing unit is used to perform binocular matching processing on binocular image data at consecutive time points, obtain the disparity distribution results of the corresponding spatial region, compare the disparity distribution results at multiple consecutive time points, extract the temporal continuity features of the direction and magnitude of disparity change, and generate disparity temporal stability data. The angular velocity analysis unit is used to analyze the angular velocity data at continuous time intervals, obtain the trajectory of the change of angular velocity direction in the time dimension, and generate consistent angular velocity direction data based on the concentration and persistence of the directional change in the trajectory. The spatial motion self-consistency unit is used to align parallax temporal stability data with angular velocity direction consistency data to generate spatial motion self-consistency data. The time offset calculation unit is used to obtain the visual pose change sequence and the inertial pose change sequence based on the binocular image data and the six-axis inertial measurement data respectively in chronological order, and compare them on the same time axis to generate time offset data. The fracture judgment unit is used to determine whether the current moment is in a state of sensing historical fracture based on spatial motion self-consistency data and time offset data. The local visual constraint processing unit is used to write the visual constraint relationship formed by the visual pose change sequence corresponding to the current moment into a local visual constraint set that is independent of the historical visual constraint set when it is determined that the perception history is broken. It also generates a local pose change sequence. The inertial pose update unit is used to update the inertial pose change sequence based on the local pose change sequence to obtain the updated inertial pose change sequence. The spatial positioning calculation unit is used to calculate the current spatial pose based on the updated inertial pose change sequence and local visual constraint relationship, for spatial positioning. The historical visual constraint verification unit is used to verify the consistency of visual constraint relationships in the historical visual constraint set when it is determined that there is no perceptual historical break. It removes the participation qualification of visual constraint relationships that are inconsistent with the inertial pose change sequence or spatial motion self-consistency data. Based on the historical visual constraint relationships that have not been removed from participation, local visual constraint relationships and the six-axis inertial measurement data at the corresponding time, it fuses and calculates the spatial pose at the current time for spatial positioning.

[0073] The functions of the above modules correspond to the method description in Example 1, and will not be described in detail here.

[0074] This application utilizes a binocular six-axis low cumulative error spatial positioning device to achieve dynamic evaluation and adaptive intervention of the consistency between visual and inertial information at continuous moments. It integrates visual constraints and inertial pose updates into a unified spatial positioning control system, thereby ensuring the continuity and accuracy of spatial pose estimation, meeting the requirements of high-precision positioning for long-term consistency and error controllability, and improving the stability and reliability of positioning results under complex environments and long-term operating conditions.

[0075] The foregoing has shown and described the basic principles, main features, and advantages of this application. Those skilled in the art should understand that this application is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of this application. Various changes and modifications can be made to this application without departing from the spirit and scope thereof, and all such changes and modifications fall within the scope of this application as claimed. The scope of protection of this application is defined by the appended claims and their equivalents.

Claims

1. A binocular six-axis spatial positioning method with low cumulative error, characterized in that, The method includes: During the operation of the device being positioned, binocular image data and six-axis inertial measurement data are acquired at continuous intervals, wherein the six-axis inertial measurement data includes at least angular velocity data; Binocular matching processing is performed on binocular image data at consecutive time points to obtain the disparity distribution results of the corresponding spatial regions. The disparity distribution results are compared at multiple consecutive time points to extract the temporal continuity features of the direction and magnitude of disparity change, and to generate disparity temporal stability data. Directional component analysis is performed on angular velocity data at continuous time intervals to obtain the trajectory of angular velocity direction change in the time dimension, and angular velocity direction consistency data is generated based on the concentration and persistence of the direction change in the trajectory. Alignment processing is performed on the disparity temporal stability data and the angular velocity direction consistency data to generate spatial motion self-consistency data; Based on the binocular image data and the six-axis inertial measurement data, visual pose change sequences and inertial pose change sequences are obtained in chronological order and compared on the same time axis to generate time offset data. Based on the spatial motion self-consistency data and time offset data, determine whether the current moment is in a state of perceived historical breakage. When it is determined that the perception history is in a broken state, the visual constraint relationship formed by the visual pose change sequence corresponding to the current moment is written only into the local visual constraint set that is independent of the historical visual constraint set, and a local pose change sequence is generated. The inertial pose change sequence is updated based on the local pose change sequence to obtain the updated inertial pose change sequence. The spatial pose at the current moment is calculated based on the updated inertial pose change sequence and the local visual constraint relationship, and is used for spatial positioning.

2. The method according to claim 1, characterized in that, The method further includes: When it is determined that the perception history is not in the state of interruption, the consistency of the visual constraint relationship in the historical visual constraint set is checked, and the participation qualification of the visual constraint relationship that is inconsistent with the inertial pose change sequence or the spatial motion self-consistency data is removed. Based on the historical visual constraints that have not been removed from the participation status, the local visual constraints, and the six-axis inertial measurement data at the corresponding moment, the spatial pose at the current moment is calculated and used for spatial positioning.

3. The method according to claim 1, characterized in that, The generation of disparity temporal stability data includes: In the binocular image data at consecutive time points, the corresponding disparity distribution results are extracted for the same spatial region, and the disparity distribution results at each time point are correlated in chronological order to form a disparity temporal set. In the disparity temporal set, the direction of change of the disparity distribution results in the spatial dimension is statistically analyzed to obtain the disparity change direction information corresponding to each consecutive time. Consistency determination is made on the disparity change direction information of adjacent time moments, the continuity of the disparity change direction in the time dimension is identified, and disparity direction continuity description data is generated. Based on the disparity time series set, the variation amplitude between disparity distribution results at adjacent times is quantitatively characterized to obtain the evolution information of disparity variation amplitude over time and generate disparity amplitude continuity description data. Based on the disparity direction continuity description data and the disparity amplitude continuity description data, the stability of disparity changes in the time dimension is comprehensively characterized to generate disparity temporal stability data.

4. The method according to claim 1, characterized in that, The generated angular velocity direction consistency data includes: The angular velocity data is decomposed into angular velocity direction components according to three orthogonal directions, and the change trajectory of each direction component in the time dimension is recorded. In the trajectory of the change of the angular velocity direction component, the offset angle of the angular velocity direction at consecutive moments is analyzed, and the concentration of the change of the angular velocity direction is statistically analyzed to obtain the concentration data of the change of direction. Analyze the persistence of the trajectory of angular velocity direction change over a continuous time period, identify the continuous interval of direction change, and generate data on the persistence of direction change; The concentration data of directional change and the persistence data of directional change are comprehensively evaluated to generate angular velocity direction consistency data.

5. The method according to claim 4, characterized in that, The generated spatial motion self-consistency data includes: The disparity temporal stability data and the angular velocity direction consistency data are aligned in the time dimension to obtain aligned disparity temporal stability data and aligned angular velocity direction consistency data. For each consecutive moment, the aligned parallax temporal stability data and the aligned angular velocity direction consistency data are correlated and analyzed to calculate the degree of matching between the two in terms of spatial motion trend and direction change. Within a continuous time span, the matching degree at each time point is statistically analyzed and the changing trend of the matching degree within the continuous time span is analyzed to evaluate the consistency and continuity of spatial motion information in the time dimension and obtain the evaluation results. The matching degree and the evaluation results are comprehensively characterized to generate spatial motion self-consistency data.

6. The method according to claim 5, characterized in that, The generated time offset data includes: Based on the binocular image data and the six-axis inertial measurement data, respectively, visual pose change sequences and inertial pose change sequences are obtained in chronological order. The visual pose change sequence and the inertial pose change sequence are compared on the same time axis to obtain the pose differences at each consecutive moment. Based on the pose differences at each consecutive moment, time offset data representing the time offset between the visual pose change sequence and the inertial pose change sequence is generated.

7. The method according to claim 6, characterized in that, The determination of whether the current moment is in a state of perceived historical discontinuity includes: Based on the spatial motion self-consistency data, a spatial self-consistency judgment result is obtained by comparing it with a preset spatial self-consistency threshold. Based on the time offset data, a time offset determination result is obtained by comparing it with a preset time offset threshold. The spatial self-consistency determination result and the time offset determination result are compared with preset judgment conditions to determine whether the current moment is in a state of perceived historical break. The preset judgment conditions include: When both the spatial self-consistency determination result and the time offset determination result are greater than the corresponding threshold, it is determined that the current moment is in a state of perceived historical discontinuity. If at least one of the spatial self-consistency determination results and the time offset determination results is less than or equal to the corresponding threshold, it is determined that the current moment is not in a state of perceived historical break.

8. The method according to claim 7, characterized in that, The step of writing the visual constraint relationship formed by the visual pose change sequence corresponding to the current moment only into a local visual constraint set independent of the historical visual constraint set, and generating a local pose change sequence, includes: Extract the visual constraint relationships from the visual pose change sequence corresponding to the current moment; Create a local set of visual constraints that is independent of the historical set of visual constraints; The extracted visual constraint relationships are written into the local visual constraint set; Based on the set of local visual constraints, the local pose changes at consecutive time intervals are calculated to generate a sequence of local pose changes.

9. The method according to claim 2, characterized in that, The process of verifying the consistency of visual constraint relationships in the historical visual constraint set and disqualifying visual constraint relationships inconsistent with the inertial pose change sequence or the spatial motion self-consistency data from participation includes: Extract the visual constraint relationships and corresponding inertial pose change data and spatial motion self-consistency data from the historical visual constraint set; The historical visual constraint relationship is compared with the inertial pose change sequence at the corresponding moment to obtain the first comparison result; The historical visual constraint relationship is compared with the spatial motion self-consistency data at the corresponding time to obtain a second comparison result. When the first comparison result or the second comparison result is inconsistent, the visual constraint relationship of the consistency of the corresponding inertial pose change sequence or the spatial motion self-consistency data is released from participation eligibility.

10. A binocular six-axis low cumulative error spatial positioning device, characterized in that, The device includes a data acquisition module and a processing module; The acquisition module includes a visual positioning component and an inertial positioning component, used to acquire binocular image data and six-axis inertial measurement data, wherein the six-axis inertial measurement data includes at least angular velocity data; The processing module includes: The disparity processing unit is used to perform binocular matching processing on binocular image data at consecutive time points, obtain the disparity distribution results of the corresponding spatial region, compare the disparity distribution results at multiple consecutive time points, extract the temporal continuity features of the direction and magnitude of disparity change, and generate disparity temporal stability data. The angular velocity analysis unit is used to analyze the angular velocity data at continuous time intervals, obtain the trajectory of the change of angular velocity direction in the time dimension, and generate angular velocity direction consistency data based on the concentration and persistence of the directional change in the trajectory. A spatial motion self-consistency unit is used to align the parallax temporal stability data with the angular velocity direction consistency data to generate spatial motion self-consistency data. The time offset calculation unit is used to obtain the visual pose change sequence and the inertial pose change sequence based on the binocular image data and the six-axis inertial measurement data respectively in chronological order, and compare them on the same time axis to generate time offset data. The fracture determination unit is used to determine whether the current moment is in a state of sensing historical fracture based on the spatial motion self-consistency data and time offset data. The local visual constraint processing unit is used to write the visual constraint relationship formed by the visual pose change sequence corresponding to the current moment into a local visual constraint set that is independent of the historical visual constraint set when it is determined that the perception history is in a broken state, and to generate a local pose change sequence. An inertial pose update unit is used to update the inertial pose change sequence based on the local pose change sequence to obtain the updated inertial pose change sequence. A spatial positioning calculation unit is used to calculate the current spatial pose based on the updated inertial pose change sequence and the local visual constraint relationship, for spatial positioning. The historical visual constraint verification unit is used to perform a consistency verification of the visual constraint relationships in the historical visual constraint set when it is determined that the perception history is not in a broken state. It removes the participation qualification of visual constraint relationships that are inconsistent with the inertial pose change sequence or the spatial motion self-consistency data, and calculates the spatial pose at the current moment based on the historical visual constraint relationships that have not been removed from the participation qualification, the local visual constraint relationships, and the six-axis inertial measurement data at the corresponding time, for spatial positioning.