State determination methods, apparatus, equipment and storage media

CN116972835BActive Publication Date: 2026-08-14ZHEJIANG SENSETIME TECH DEV CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-20
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

但相关技术中,利用三角化等方式确定特征点的三维信息,存在计算量较大、精度不高等问题

Benefits of technology

[0011]本公开实施例中,首先,通过在视觉惯性定位系统的运行过程中,从视觉惯性定位系统中获取待处理图像、参考跟踪点的跟踪信息和惯性信息;其中,待处理图像是通过视觉惯性定位系统中的视觉传感器采集的,参考跟踪点的跟踪信息是基于视觉传感器采集的历史处理图像确定的,惯性信息是通过视觉惯性定位系统中的惯性传感器采集的;其次,简单地确定待处理图像中与参考跟踪点匹配的目标跟踪点,以及目标跟踪点的位置信息;然后,基于目标跟踪点的位置信息和参考跟踪点的跟踪信息,快速准确地确定视觉传感器对应的位姿约束;这样,相较于相关技术中,通过确定目标跟踪点的三维信息,基于三维信息确定重投影误差等约束,可以通过直接利用目标跟踪点的位置信息确定位姿约束,可以减少确定三维信息的计算量,以及降低三维信息的准确度较低导致状态更新时精度较低的问题;最后,基于位姿约束和惯性信息,确定视觉惯性定位系统中的状态量,以实现快速准确地对视觉惯性定位系统中的状态量进行更新,得到更新后的状态量,有助于简化状态更新的复杂性、提高状态更新的效率等。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116972835B_ABST
    Figure CN116972835B_ABST
Patent Text Reader

Abstract

This disclosure provides a state determination method, apparatus, device, and storage medium. The method includes: during the operation of a visual-inertial positioning system, acquiring a to-be-processed image, tracking information of a reference tracking point, and inertial information from the visual-inertial positioning system; wherein the to-be-processed image is acquired by a visual sensor in the visual-inertial positioning system, the tracking information of the reference tracking point is determined based on historical processed images acquired by the visual sensor, and the inertial information is acquired by an inertial sensor in the visual-inertial positioning system; determining a target tracking point in the to-be-processed image that matches the reference tracking point, and the position information of the target tracking point; determining the pose constraint corresponding to the visual sensor based on the position information of the target tracking point and the tracking information of the reference tracking point; and determining the state variables in the visual-inertial positioning system based on the pose constraint and the inertial information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to, but is not limited to, the field of computer vision technology, and in particular to a state determination method, apparatus, device, and storage medium. Background Technology

[0002] Visual-inertial (tracking) positioning systems are a crucial underlying technology in fields such as computer vision, robotics, autonomous vehicles, 3D reconstruction, and augmented reality. In these technologies, visual-inertial positioning systems perform 3D reconstruction processing on feature points in acquired images to obtain the 3D information of these feature points, and then update the state of the visual-inertial positioning system based on this 3D information. Therefore, the determinism of the state in a visual-inertial positioning system is highly correlated with the accuracy of the 3D information. However, methods such as triangulation used to determine the 3D information of feature points in these technologies suffer from high computational complexity and low accuracy. Summary of the Invention

[0003] In view of the above, embodiments of this disclosure provide at least one state determination method, apparatus, device, and storage medium.

[0004] The technical solution of this disclosure embodiment is implemented as follows:

[0005] On one hand, embodiments of this disclosure provide a state determination method, comprising: during the operation of a visual-inertial positioning system, acquiring a to-be-processed image, tracking information of a reference tracking point, and inertial information from the visual-inertial positioning system; wherein the to-be-processed image is acquired by a visual sensor in the visual-inertial positioning system, the tracking information of the reference tracking point is determined based on historical processed images acquired by the visual sensor, and the inertial information is acquired by an inertial sensor in the visual-inertial positioning system; determining a target tracking point in the to-be-processed image that matches the reference tracking point, and the position information of the target tracking point; determining a pose constraint corresponding to the visual sensor based on the position information of the target tracking point and the tracking information of the reference tracking point; and determining a state variable in the visual-inertial positioning system based on the pose constraint and the inertial information.

[0006] On the other hand, embodiments of this disclosure provide a state determination device, comprising: a first acquisition module, configured to acquire, during the operation of a visual inertial positioning system, an image to be processed, tracking information of a reference tracking point, and inertial information from the visual inertial positioning system; wherein, the image to be processed is acquired by a visual sensor in the visual inertial positioning system, the tracking information of the reference tracking point is determined based on historical processed images acquired by the visual sensor, and the inertial information is acquired by an inertial sensor in the visual inertial positioning system; a first determination module, configured to determine a target tracking point in the image to be processed that matches the reference tracking point, and the position information of the target tracking point; a second determination module, configured to determine the pose constraint corresponding to the visual sensor based on the position information of the target tracking point and the tracking information of the reference tracking point; and a third determination module, configured to determine the state variables in the visual inertial positioning system based on the pose constraint and the inertial information.

[0007] In another aspect, embodiments of this disclosure provide a computer device including a memory and a processor, wherein the memory stores a computer program that can run on the processor, and the processor executes the program to implement some or all of the steps in the above-described method.

[0008] In another aspect, embodiments of this disclosure provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements some or all of the steps in the above-described method.

[0009] In another aspect, embodiments of this disclosure provide a computer program including computer-readable code, which, when executed in a computer device, causes a processor in the computer device to perform some or all of the steps in the above-described method.

[0010] In another aspect, embodiments of this disclosure provide a computer program product, the computer program product including a non-transitory computer-readable storage medium storing a computer program, wherein when the computer program is read and executed by a computer, it implements some or all of the steps in the above method.

[0011] In this embodiment of the disclosure, firstly, during the operation of the visual-inertial positioning system, the image to be processed, tracking information of the reference tracking point, and inertial information are acquired from the visual-inertial positioning system. The image to be processed is acquired by a visual sensor in the visual-inertial positioning system, the tracking information of the reference tracking point is determined based on historical processed images acquired by the visual sensor, and the inertial information is acquired by an inertial sensor in the visual-inertial positioning system. Secondly, the target tracking point in the image to be processed that matches the reference tracking point, and the position information of the target tracking point, are simply determined. Then, based on the position information of the target tracking point and the tracking information of the reference tracking point, the image to be processed is quickly and accurately... Accurately determine the pose constraints corresponding to the visual sensor. Compared to related technologies that determine constraints such as reprojection error based on the 3D information of the target tracking point, this method directly uses the position information of the target tracking point to determine the pose constraints. This reduces the computational load of determining 3D information and mitigates the problem of low accuracy in state updates due to the low accuracy of 3D information. Finally, based on the pose constraints and inertial information, determine the state variables in the visual inertial positioning system to achieve fast and accurate updates of the state variables. Obtaining the updated state variables helps simplify the complexity of state updates and improve their efficiency.

[0012] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and are not intended to limit the technical solutions of this disclosure. Attached Figure Description

[0013] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the specification, serve to illustrate the technical solutions of this disclosure.

[0014] Figure 1 A schematic diagram illustrating the implementation flow of a state determination method provided in an embodiment of this disclosure;

[0015] Figure 2 A schematic diagram illustrating the implementation flow of a state determination method provided in an embodiment of this disclosure;

[0016] Figure 3 A schematic diagram illustrating the implementation flow of a state determination method provided in an embodiment of this disclosure;

[0017] Figure 4 A schematic diagram illustrating the implementation flow of a state determination method provided in an embodiment of this disclosure;

[0018] Figure 5 This is a schematic diagram of the composition structure of a state determination device provided in an embodiment of the present disclosure;

[0019] Figure 6This is a schematic diagram of the hardware entity of a computer device provided in an embodiment of this disclosure. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of this disclosure clearer, the technical solutions of this disclosure are further described in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be regarded as limitations on this disclosure. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0021] In the following description, references to "some embodiments" describe a subset of all possible embodiments; however, it is understood that "some embodiments" may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict. The terms "first / second / third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first / second / third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein.

[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. The terminology used herein is for descriptive purposes only and is not intended to limit this disclosure.

[0023] This disclosure provides a state determination method, which can be executed by a processor of a computer device. The computer device can refer to mobile devices such as robots, unmanned vehicles, and drones, or devices with state determination capabilities such as user equipment, user terminals, cordless phones, personal digital assistants, handheld devices, computing devices, in-vehicle devices, augmented reality setups, virtual reality devices, and visual inertial navigation devices. Figure 1 This is a schematic diagram illustrating the implementation flow of a state determination method provided in an embodiment of this disclosure, as shown below. Figure 1 As shown, the method includes the following steps S101 to S104:

[0024] Step S101: During the operation of the visual inertial positioning system, the tracking information and inertial information of the image to be processed and the reference tracking point are obtained from the visual inertial positioning system.

[0025] Here, a visual-inertial positioning system can refer to a method that fuses image information (also known as visual information, etc.) and inertial information for simultaneous localization and environmental reconstruction. The image to be processed can be acquired by a visual sensor in the visual-inertial positioning system, such as a camera; the image to be processed can be a single frame or multiple frames, which is not limited here.

[0026] The tracking information of the reference tracking points is determined based on historically processed images acquired by the vision sensor. A reference tracking point can refer to a feature point used for feature point tracking in the historically processed images. The tracking information of the reference tracking points can include direct information such as the number of reference tracking points and their position coordinates in each frame of the historically processed images, as well as indirect information such as the inverse depth and 3D coordinates of the reference tracking points. The indirect information, such as the inverse depth and 3D coordinates, can be represented using the direct information, such as the position coordinates of the reference tracking points in multiple frames of historically processed images and preset parameters of the vision sensor (e.g., camera parameters). The image to be processed can refer to the image acquired by the vision sensor at the current moment, while the historically processed images can refer to images acquired by the vision sensor at previous moments.

[0027] For example, feature point detection is performed on multiple historical processing images to obtain a set of feature points corresponding to each historical processing image; based on the position coordinates of each feature point, multiple sets of feature points are matched to obtain matching results (e.g., the first feature point on the first historical processing image matches the third feature point on the second historical processing image and the second feature point on the third historical processing image). The matching results can be used to characterize the tracking results of one or more points (i.e., reference tracking points) on an object in the world coordinate system in multiple historical processing images.

[0028] Inertial sensors can refer to devices such as inertial measurement units (IMUs), which may include components such as gyroscopes and accelerometers. Inertial information is collected by inertial sensors in a visual-inertial positioning system, such as position and attitude information from the inertial sensors. It may also include angular velocity collected by gyroscopes and acceleration collected by accelerometers, etc., and is not limited here.

[0029] The operation of a visual-inertial positioning system can include an initialization process and a subsequent running process. The initialization process involves acquiring the initial state variables of the visual-inertial positioning system, including visual and inertial state variables such as absolute scale, gyroscope bias, acceleration bias, and gravitational acceleration. The running process involves updating the state variables of the visual-inertial positioning system to obtain new state variables, which are then used to correct the biases of the visual and inertial sensors to achieve tracking and positioning functions. In this embodiment, the determination and updating of the state variables of the visual-inertial positioning system can be achieved during the running process after initialization.

[0030] The implementation of step S101 may include: acquiring image data using a visual sensor in the visual inertial positioning system to obtain an image to be processed; acquiring inertial data using an inertial sensor in the visual inertial positioning system to obtain inertial information; obtaining historical processed images from the visual inertial positioning system, and obtaining tracking information of the reference tracking point based on the historical processed images.

[0031] Step S102: Determine the target tracking point in the image to be processed that matches the reference tracking point, and the position information of the target tracking point.

[0032] Here, the position information of the target tracking point can be the pixel coordinates of the target tracking point relative to the pixel coordinate system (with the upper left corner of the image to be processed as the origin), or the normalized coordinates of the target tracking point relative to the normalized coordinate system, etc., and is not limited here; for example, a ray of light originating from a spatial point passes through the optical center of the vision sensor and is projected onto the plane in front of the vision sensor to obtain the normalized coordinates of the spatial point; where pixel coordinates and normalized coordinates can be converted based on the preset parameters of the vision sensor.

[0033] Taking the tracking information of multiple reference tracking points, including their position coordinates on multiple frames of historical processed images, as an example, feature point detection is performed on the image to be processed to obtain multiple feature points. Based on the position information of these multiple feature points, they are matched with multiple reference tracking points to obtain matching results, such as the first feature point matching the first reference tracking point, the second feature point matching the second reference tracking point, etc. The feature point that matches the reference tracking point is determined as the target tracking point, and the position information of the target tracking point is obtained.

[0034] Step S103: Based on the position information of the target tracking point and the tracking information of the reference tracking point, determine the pose constraint corresponding to the visual sensor.

[0035] Here, pose constraint can refer to the projection error of the target tracking point. Pose constraint can include the movement constraint and rotation constraint of the vision sensor. In this embodiment, the projection error of the target tracking point can be understood as the direct constraint of the target tracking point's observation on the vision sensor's pose across multiple frames of images. For example, based on the projection method of the vision sensor, a transformation relationship can be constructed between the target tracking point's position information, inverse depth of the target tracking point, the position and orientation of the vision sensor relative to the world coordinate system, and the three-dimensional information of the target tracking point. This transformation relationship can be processed by deformation or other methods to obtain the projection error of the target tracking point, thereby determining the movement constraint of the vision sensor relative to its position and the rotation constraint relative to its orientation.

[0036] The inverse depth and 3D information of the target tracking point can be represented using the tracking information of the reference tracking point without direct calculation. For example, the 3D information of the reference tracking point can be represented using the position information, inverse depth, and position and orientation of the vision sensor relative to the world coordinate system from historical processed images. Since the reference tracking point and the target tracking point are matched, the 3D information of the reference tracking point is equal to that of the target tracking point. The position and orientation of the vision sensor relative to the world coordinate system can be estimated. For example, the position of the vision sensor relative to the world coordinate system can be set as a first estimated value, and the orientation can be set as a second estimated value.

[0037] Step S104: Based on the pose constraints and the inertial information, determine the state variables in the visual inertial positioning system.

[0038] Here, the state variables in the visual inertial positioning system can be used to correct the system, helping to improve tracking and positioning accuracy. These state variables can include visual state variables (e.g., position state variables) and inertial state variables (e.g., acceleration state variables). Since pose constraints apply to the motion and rotation constraints of the visual sensor, the motion and rotation state variables in this embodiment can be determined. For example, a trained model can be used to determine these variables. The trained model can be understood as a pre-set machine learning model, such as a neural network model, used to determine the state variables in the visual inertial positioning system.

[0039] In some embodiments, pose constraints can also be solved based on inertial information to obtain translational and rotational state variables. For example, nonlinear optimization methods such as extended Kalman filtering can be used to solve pose constraints to obtain translational and rotational state variables, as well as the covariance matrices corresponding to the translational and rotational state variables.

[0040] In some embodiments, after obtaining the state variables in the visual inertial positioning system, the method may further include: updating the state variables in the visual inertial positioning system based on the determined state variables in the visual inertial positioning system to obtain updated state variables.

[0041] For example, based on the determined motion and rotation state variables in the visual inertial positioning system, the motion and rotation state variables in the visual inertial positioning system are updated to obtain the updated motion and rotation state variables.

[0042] The updated motion and rotation state variables can be used to correct the currently acquired position and state of the visual inertial positioning system. For example, the corrected position can be obtained by adding the currently acquired position and the updated motion state variable; similarly, the corrected attitude can be obtained by adding the currently acquired attitude and the updated rotation state variable. The covariance matrices corresponding to the updated motion and rotation state variables can be used for convergence analysis, thereby improving the accuracy of the visual inertial positioning system.

[0043] In this embodiment of the disclosure, firstly, during the operation of the visual-inertial positioning system, the image to be processed, tracking information of the reference tracking point, and inertial information are acquired from the visual-inertial positioning system. The image to be processed is acquired by a visual sensor in the visual-inertial positioning system, the tracking information of the reference tracking point is determined based on historical processed images acquired by the visual sensor, and the inertial information is acquired by an inertial sensor in the visual-inertial positioning system. Secondly, the target tracking point in the image to be processed that matches the reference tracking point, and the position information of the target tracking point, are simply determined. Then, based on the position information of the target tracking point and the tracking information of the reference tracking point, the image to be processed is quickly and accurately... Accurately determine the pose constraints corresponding to the visual sensor. Compared to related technologies that determine constraints such as reprojection error based on the 3D information of the target tracking point, this method directly uses the position information of the target tracking point to determine the pose constraints. This reduces the computational load of determining 3D information and mitigates the problem of low accuracy in state updates due to the low accuracy of 3D information. Finally, based on the pose constraints and inertial information, determine the state variables in the visual inertial positioning system to achieve fast and accurate updates of the state variables. Obtaining the updated state variables helps simplify the complexity of state updates and improve their efficiency.

[0044] This disclosure provides a state determination method, wherein the position information of the target tracking point is the normalized coordinates of the target tracking point; for example... Figure 2As shown, the method includes the following steps S201 to S206:

[0045] Step S201 corresponds to the aforementioned step S101, and can be implemented with reference to the specific implementation of the aforementioned step S101; Steps S205 to S206 correspond to the aforementioned steps S103 to S104 respectively, and can be implemented with reference to the specific implementation of the aforementioned steps S103 to S104.

[0046] Step S202: Perform feature point detection on the image to be processed to obtain at least two candidate feature points.

[0047] Here, feature point detection algorithms such as Oriented Fast and Rotated BRIEF (ORB) and Optical Flow (Lucas Kanade, LK) can be used to detect feature points in the image to be processed, obtaining at least two candidate feature points and determining their image coordinates. For example, corner detection can be performed on the image to be processed, and the detected corner points can be identified as candidate feature points.

[0048] Step S203: Based on the tracking information of the reference tracking point, determine the target tracking point and the image coordinates of the target tracking point from the at least two candidate feature points.

[0049] Here, based on the tracking information of the reference tracking point and the image coordinates of the candidate feature point, the candidate feature point and the reference feature point can be matched to determine the target tracking point that matches the reference feature point, and the image coordinates of the target feature point can be determined; here, the image coordinates of the target feature point can be pixel coordinates.

[0050] In some embodiments, historical processed images and images to be processed can be acquired, feature points can be extracted from the historical processed images and images to be processed to obtain multiple sets of feature points, and matching processing can be performed on each set of feature points to obtain the matching results of feature points. The feature points that are successfully matched in the historical processed images are determined as reference feature points, and the corresponding feature points that are successfully matched in the images to be processed are determined as target feature points.

[0051] Step S204: Transform the image coordinates of the target tracking point to obtain the normalized coordinates of the target tracking point.

[0052] Here, based on the preset parameters of the vision sensor, the transformation matrix between the pixel coordinate system and the normalized coordinate system of the vision sensor is determined. Using this transformation matrix, the image coordinates of the target tracking point are converted into the normalized coordinates of the target tracking point; wherein, the normalized coordinates of the target tracking point can be represented by a normalized unit vector.

[0053] In this embodiment of the disclosure, the target feature point that matches the reference feature point can be quickly and accurately determined by using the tracking information of the reference tracking point, and then the normalized coordinates of the target feature point can be accurately obtained.

[0054] This disclosure provides a state determination method, wherein the tracking information of the reference tracking point includes preset three-dimensional information of the target tracking point and preset inverse depth of the target tracking point; such as Figure 3 As shown, the method includes the following steps S301 to S304:

[0055] Steps S301 to S302 correspond to the aforementioned steps S101 to S102 respectively, and can be implemented with reference to the specific implementation of the aforementioned steps S101 to S102; step S304 corresponds to the aforementioned step S104, and can be implemented with reference to the specific implementation of the aforementioned step S104.

[0056] Step S303: Using the projection relationship of the vision sensor, and based on the position information, the three-dimensional information, and the inverse depth, determine the pose error function of the vision sensor.

[0057] Here, the preset 3D information of the target tracking point can be characterized by the position information of the reference tracking point on the historical processed image, the inverse depth of the reference tracking point, and the position and orientation of the visual sensor relative to the world coordinate system. The preset inverse depth of the target tracking point can be characterized by the 3D information and normalized coordinates of the reference tracking point, etc., and is not limited here. The pose error function can be used to characterize the pose constraints corresponding to the visual sensor, and the pose error function is used to determine the movement and rotation information of the visual sensor. For example, the movement and rotation information of the visual sensor can be set as estimates, and the position information such as the normalized coordinates of the target tracking point, the 3D information of the target tracking point, and the inverse depth can be determined as known quantities. Based on the imaging principle, the pose error function of the visual sensor can be constructed. In some embodiments, the pose error function can be determined using the following formula:

[0058]

[0059] In formula (1), we can assume that the three-dimensional information of the target tracking point is m, u i Let λ be the normalized unit vector (i.e., normalized coordinates) of the target tracking point in the i-th frame (i.e., the image to be processed). i p is the inverse depth value. ci Let R be the position of the visual sensor relative to the world coordinate system at frame i (i.e., the movement information of the visual sensor). ci Let i be the rotation matrix (i.e., the rotation information of the vision sensor), where i is a positive integer.

[0060] Formula (1) can be rewritten in the following form:

[0061] R ci u i +λ i P ci -λ i m = 0, i = 1, ..., N (2);

[0062] Formula (2) can be converted into matrix form:

[0063]

[0064] In formula (3), I3 is the identity matrix.

[0065] Assuming there are 4 frames of images to be processed where the target tracking points can be observed, a matrix can be constructed as follows:

[0066]

[0067] Dividing the matrix constructed in formula (4) into three parts, it can be rewritten as:

[0068]

[0069] in,

[0070]

[0071]

[0072]

[0073]

[0074] Due to u i If it is a unit vector, then:

[0075] (R ci u i ) T (R ci u i )=u i T (R ci T R ci )u i =u i T u i =I (10);

[0076] From formula (10), we can know that Q T Q = I, construct matrix P = I - QQ T Using the formula (5) to multiply P by the left, PQ = QQ(Q)T If Q) = 0, Q can be eliminated, and formula (5) can be rewritten as:

[0077]

[0078] Construct matrix G = I - WHW T P, where H = (W) T PW) -1 Multiplying P*G by W on the left, PW = PW - PGW((W T PW) -1 (W T PW))=PW-PW=0, then formula (11) can be rewritten as:

[0079] PGp = 0 (12);

[0080] In formula (12), p represents the motion information from the visual sensor. Similarly, using matrix HW... T Multiplying by formula (11) on the left, the three-dimensional coordinates of the target feature point can be obtained as follows:

[0081] m = -HW T Pp (13);

[0082] Among them, PGp=0 in formula (12) can be determined as the pose error function that the visual inertial positioning system needs to optimize.

[0083] In this embodiment of the disclosure, by avoiding the calculation of the three-dimensional information of the target feature points, the pose error function of the visual inertial positioning system can be quickly and accurately determined directly based on the position information of the target tracking point and the tracking information of the reference tracking point. This helps to improve the efficiency and accuracy of determining the state variables in the visual inertial positioning system.

[0084] In some embodiments, steps S311 to S314 may be included before step S304:

[0085] Step S311: Obtain a preset number of historical processed images adjacent to the image to be processed.

[0086] For example: the vision sensor acquires images at a preset acquisition frequency, with a preset number of 2, acquires a first historical processed image adjacent to the image to be processed, and acquires a second historical processed image adjacent to the first historical processed image.

[0087] Step S312: Determine the historical tracking point that matches the target tracking point in each of the historical processed images, and the image coordinates of each of the historical tracking points.

[0088] Here, historical tracking points can refer to feature points on historically processed images that match the target tracking point. Reference tracking points can be determined as historical tracking points. For example, the image coordinates of the historical tracking point on the first historically processed image are (a, b), the image coordinates of the historical tracking point on the second historically processed image are (c, d), and the image coordinates of the target tracking point on the image to be processed are (e, f).

[0089] Step S313: Based on the image coordinates of each historical tracking point and the image coordinates of the target tracking point, obtain image regions of a preset scale from the historical processed image and the image to be processed, respectively.

[0090] Here, the preset scale can be 5*5, 7*7, 8*8, etc., and is not limited to this. Using the image coordinates of the historical tracking points and the image coordinates of the target tracking point as the center, the historical processed image and the image to be processed are cropped to obtain an image region of the preset scale.

[0091] Step S314: Determine the confidence level of the pose constraint based on the pixel values ​​of the pixels in the image region of each historical tracking point and the pixel values ​​of the pixels in the image region of the target tracking point.

[0092] In the actual scene tracking process of a visual inertial positioning system, mismatches between feature points may occur due to various reasons. If the same confidence level (weight) is set for all feature points, the incorrect matches will affect the accuracy of the pose constraint. Therefore, different confidence levels can be assigned to different target feature points to improve the accuracy of the pose constraint. The confidence level of the pose constraint can be used to characterize the weight of the pose constraint. For example: determine the pixel difference matrix between each image region; determine the variance corresponding to each pixel difference matrix; determine the mean of all variances, and use the mean of all variances as the confidence level of the pose constraint.

[0093] In this embodiment of the disclosure, the confidence level of pose constraints can be quickly and accurately determined by using historical processed images and images to be processed, which helps to improve the accuracy of pose constraints.

[0094] In some embodiments, step S314 may include the following steps S3141 to S3144:

[0095] Step S3141: Determine the average pixel value of the image region of each historical tracking point and the image region of the target tracking point.

[0096] For example, the average pixel value of the image area corresponding to the first historical tracking point is 125, the average pixel value of the image area corresponding to the second historical tracking point is 130, and the average pixel value of the image area corresponding to the target tracking point is 127, etc.

[0097] Step S3142: Based on the pixel mean of each image region, normalize the image region of each historical tracking point and the image region of the target tracking point to obtain the corresponding normalized image region.

[0098] Here, the pixel value of each pixel in the image region of the first historical tracking point can be subtracted from the average pixel value of the corresponding image region of the first historical tracking point to obtain the normalized image region of the first historical tracking point; the pixel value of each pixel in the image region of the second historical tracking point can be subtracted from the average pixel value of the corresponding image region of the second historical tracking point to obtain the normalized image region of the second historical tracking point; the pixel value of each pixel in the image region of the target tracking point can be subtracted from the average pixel value of the corresponding image region of the target tracking point to obtain the normalized image region of the target tracking point.

[0099] Step S3143: Determine the root mean square between each of the normalized image regions.

[0100] Here, the normalized image region of the first historical tracking point can be subtracted from the corresponding pixels of the normalized image region of the second historical tracking point to obtain a first result matrix; the first root mean square of the first result matrix is ​​determined; the normalized image region of the first historical tracking point can be subtracted from the corresponding pixels of the normalized image region of the target tracking point to obtain a second result matrix; the second root mean square of the second result matrix is ​​determined; the normalized image region of the second historical tracking point can be subtracted from the corresponding pixels of the normalized image region of the target tracking point to obtain a third result matrix; the third root mean square of the third result matrix is ​​determined, etc.

[0101] Step S3144: Determine the confidence level of the pose constraint based on all the root mean squares.

[0102] Here, the mean of the root mean squares among the first, second, and third root mean squares can be determined, and this mean of the root mean squares can be used as the confidence level of the pose constraint, etc.

[0103] In some embodiments, there can be multiple target tracking points, then the pose error function corresponding to the image to be processed can be expressed as:

[0104]

[0105] In formula (14), α i P represents the confidence level of target tracking point i. i G i p i Formula (12) represents the target tracking point i.

[0106] In this embodiment of the disclosure, the confidence level of the pose constraint corresponding to each target tracking point can be quickly and accurately determined by the pixel value of each image region.

[0107] This disclosure provides a state determination method, wherein the visual-inertial positioning system includes an extended Kalman filter, and the state quantities in the visual-inertial positioning system include motion state quantities and rotation state quantities; such as Figure 4 As shown, the method includes the following steps S401 to S405:

[0108] Steps S401 to S403 correspond to the aforementioned steps S301 to S303 respectively. When implementing these steps, you can refer to the specific implementation methods of the aforementioned steps S301 to S303.

[0109] Step S404: Perform partial derivative processing on the pose error function to obtain the Jacobian matrix corresponding to the movement information and the Jacobian matrix corresponding to the rotation information.

[0110] Here, the Extended Kalman Filter (EKF) is an extension of the standard Kalman filter in the nonlinear case. The EKF algorithm performs a Taylor expansion of the nonlinear function, omitting higher-order terms and retaining only the first-order terms of the expansion, thereby linearizing the nonlinear function. Finally, it approximates the state and variance estimates of the system using the Kalman filter algorithm, and then filters the signal. The purpose of the EKF algorithm can be understood as transforming the calculation of accurate posterior probabilities into effective estimations of the mean (i.e., the updated state variables) and variance (i.e., the covariance matrix corresponding to the state variables).

[0111] The EKF algorithm can include a prediction process and an update process. For example, the prediction process can include: using the posterior estimate determined from the historical processed image acquired at the previous time step, transforming it through a state transition matrix to obtain the prior estimate state at the current time step; using the posterior estimate covariance from the previous time step to calculate the prior estimate covariance at the current time step; and estimating the prior estimate mean and covariance at the current time step using the posterior estimate mean and variance from the previous time step. The update process can include: firstly, calculating the Kalman gain matrix using the prior estimate covariance matrix, the observation matrix, and the measurement state covariance matrix; wherein, the observation matrix can be determined by pose constraints and inertial information, such as based on inertial information... The process involves determining the pose and other state variables of the inertial sensor; determining the movement and rotation state variables of the visual sensor based on pose constraints; determining the observation matrix based on the pose and other state variables of the inertial sensor, as well as the movement and rotation state variables of the visual sensor; then projecting the prior estimates of the Kalman filter onto the measurement space through the observation matrix and calculating the residuals with the measured values; fusing the predicted values ​​of the Kalman filter and the measured values ​​according to the Kalman gain ratio to obtain the posterior estimates; and calculating the posterior estimate covariance of the Kalman filter to update the posterior probability distribution, i.e., determining the state values ​​and the corresponding covariance.

[0112] The implementation process of the EKF algorithm is not limited here. In some embodiments, Kalman filters, iterative Kalman filters, unscented Kalman filters, etc., can also be used to determine the state variables in the visual-inertial positioning system and obtain the updated state variables.

[0113] During the implementation of step S404, the following may be included: determining the partial derivative of the pose error function with respect to the movement information to obtain the Jacobian matrix corresponding to the movement information; determining the partial derivative of the pose error function with respect to the rotation information to obtain the Jacobian matrix corresponding to the rotation information.

[0114] Step S405: Based on the extended Kalman filter and the inertial information, process the Jacobian matrix corresponding to the motion information and the Jacobian matrix corresponding to the rotation information respectively to obtain the motion state variables and rotation state variables in the visual inertial positioning system.

[0115] Here, the Jacobian matrix corresponding to the motion information, the Jacobian matrix corresponding to the rotation information, and the inertial information can be input into the extended Kalman filter to obtain the motion state variables and rotation state variables in the visual inertial positioning system.

[0116] In some embodiments, a sliding window approach can be used with an extended Kalman filter to determine the state variables in the visual inertial positioning system. For example, if an image acquired by a visual sensor is a keyframe image, it can be added to a sliding window and identified as the image to be processed. Then, based on this image, the state variables corresponding to the current sliding window can be determined, etc., without limitation. When adding the image to the sliding window, the previously processed images can be other images within the sliding window. If the image is not a keyframe image, the state variables in the visual inertial positioning system are not determined, and the next frame image is acquired.

[0117] In this embodiment of the disclosure, by using an extended Kalman filter, the motion state variables and rotation state variables in the visual inertial positioning system can be determined quickly and accurately.

[0118] In some embodiments, step S405 may include the following steps S4051 to S4052:

[0119] Step S4051: Based on the extended Kalman filter and the inertial information, process the Jacobian matrix corresponding to the movement information and the Jacobian matrix corresponding to the rotation information respectively to obtain the covariance matrix corresponding to the movement state quantity and the covariance matrix corresponding to the rotation state quantity.

[0120] Here, the Jacobian matrix corresponding to the movement information, the Jacobian matrix corresponding to the rotation information, and the inertial information can be input into the extended Kalman filter to achieve prediction by the extended Kalman filter and obtain the covariance matrix corresponding to the movement state variables and the covariance matrix corresponding to the rotation state variables.

[0121] Step S4052: Determine the movement state quantity and the rotation state quantity based on the covariance matrix corresponding to the movement state quantity and the covariance matrix corresponding to the rotation state quantity.

[0122] Here, the extended Kalman filter can be used to further process the covariance matrices corresponding to the motion state variables and the rotation state variables, so as to update the extended Kalman filter and obtain the motion state variables and rotation state variables in the visual inertial positioning system.

[0123] In this embodiment of the disclosure, by determining the covariance matrix corresponding to the moving state variable and the covariance matrix corresponding to the rotating state variable, the moving state variable and the rotating state variable can be accurately determined, and the convergence analysis of the moving state variable and the rotating state variable can be realized (e.g., determining the posterior probability distribution).

[0124] The state determination method provided in this disclosure is illustrated below using an operational scenario in a visual inertial positioning system as an example.

[0125] Visual inertial positioning systems in related technologies can be divided into three categories: feature point reprojection error optimization schemes based on sliding window mechanisms, which simultaneously optimize the 3D information of the scene and the system state information in the sliding window; feature point reprojection error optimization schemes based on local and / or global images, which solve for the current system state information based on the constructed 3D sparse point cloud; and photometric error optimization schemes based on dense and / or semi-dense depth information of keyframes, which solve for the current system state information based on the depth and photometric information of keyframes.

[0126] However, the state information obtained by visual inertial positioning systems in related technologies is highly correlated with the accuracy of their 3D information. However, 3D information is difficult to obtain in certain scenarios, such as the initialization process of a visual inertial positioning system: each state variable of the system has high uncertainty during initialization, and the solved 3D information of the scene differs significantly from the true value. Incorrect 3D information can further deteriorate the initial information, leading to initialization failure or anomalies. Furthermore, most 3D information in the scene is far from the visual sensor: the triangulation of feature points relies on sufficient observation parallax. If the visual sensor moves within a few meters, it cannot correctly triangulate feature points at a depth of one kilometer. Inaccurate triangulation information of feature points will further lead to the failure of the visual sensor's position determination.

[0127] The visual inertial positioning system in this embodiment is a scheme that determines the constraints of the visual sensor solely based on image information.

[0128] First, compared to visual inertial positioning systems in related technologies, in this embodiment of the present disclosure, the visual inertial positioning system does not need to solve for the three-dimensional information of feature points in the scene during the tracking and positioning process.

[0129] Secondly, the visual inertial positioning system in related technologies has a complex and time-consuming mechanism for constructing and managing three-dimensional points and / or three-dimensional scenes. The visual inertial positioning system in this embodiment of the present disclosure gets rid of the logic of three-dimensional scene construction, thereby further improving the efficiency of state determination and updating, and reducing system complexity.

[0130] Secondly, in related technologies, the visual inertial positioning system exhibits coupling and mutual interference between 3D information and tracking / positioning state information in initialization, distant views, and non-static scenes. If the 3D information is incorrectly calculated, it severely interferes with the tracking / positioning state calculation. The visual inertial positioning system in this embodiment does not involve solving for 3D information; it only solves for the tracking / positioning state information, thus effectively suppressing the mutual interference between the two.

[0131] Then, the visual inertial positioning system in the related technology relies on error functions such as reprojection error and photometric error based on the three-dimensional to two-dimensional projection relationship. The visual inertial positioning system in the embodiments of this disclosure proposes an error function based solely on the pose of the visual sensor (i.e., the pose error function), and directly uses the observation information of the two-dimensional image (i.e. the image to be processed) to solve for the state information of the visual inertial positioning system.

[0132] Finally, if all feature points are applied equally to the pose error function, outliers will greatly affect the accuracy and robustness of the system. The visual inertial positioning system in this embodiment proposes a feature point confidence determination scheme based on image information, which can effectively reduce the impact of outliers on the system.

[0133] The visual inertial positioning system in this embodiment derives a pose error function based on two-dimensional feature tracking points (i.e., target feature points); different confidence levels are applied to different feature points according to the image patch (i.e., different image regions) information of the feature points in different images to prevent incorrect matching points from causing a large negative impact on the system; the pose error function is combined with the existing tracking and positioning system to form a complete visual inertial positioning system.

[0134] In determining the pose error function, intermediate steps can be rewritten in other forms, or other similar pose error functions can be derived; this is not limited. More optimization parameters can be added to the pose error function, such as the system's current speed, the deviation between the visual sensor's and inertial sensor's running time, and the visual sensor's exposure time. Other information can be used to calculate the confidence level of feature points, such as the distance between feature descriptors in the feature space. Feature points (i.e., tracking points) on the image can be composed of descriptors and keypoints. Feature points can also be clustered, with different weights applied to different categories. Classification methods can be based on the location distribution of feature points, the tracking length of feature points, or the similarity of feature descriptors. There are various options for the combined tracking and positioning system, including sliding window optimization, nonlinear optimizers, or filters, as well as nonlinear optimization systems that generate 3D points at the back end and track local maps at the front end.

[0135] This approach helps improve the computational efficiency of state determination and updates, reduces the complexity of visual-inertial positioning systems, and better handles tracking and positioning scenarios where the 3D structure is difficult to recover, such as initialization, distant views, and dynamic scenes. It effectively distinguishes itself from traditional 3D structure recovery-based tracking and positioning systems. This visual-inertial positioning system can be applied to smartphones, virtual reality headsets, augmented reality glasses, autonomous vehicles, and other devices that simultaneously possess visual (or image) sensors and inertial sensor modules.

[0136] Based on the foregoing embodiments, this disclosure provides a state determination device, which includes various units and modules included in each unit, and can be implemented by a processor in a computer device; of course, it can also be implemented by specific logic circuits; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.

[0137] Figure 5 This is a schematic diagram of the composition structure of a state determination device provided in an embodiment of the present disclosure, as shown below. Figure 5 As shown, the state determination device 500 includes: a first acquisition module 510, a first determination module 520, a second determination module 530, and a third determination module 540, wherein:

[0138] The first acquisition module 510 is used to acquire, during the operation of the visual-inertial positioning system, an image to be processed, tracking information of a reference tracking point, and inertial information from the visual-inertial positioning system; wherein, the image to be processed is acquired by a visual sensor in the visual-inertial positioning system, the tracking information of the reference tracking point is determined based on historical processed images acquired by the visual sensor, and the inertial information is acquired by an inertial sensor in the visual-inertial positioning system; the first determination module 520 is used to determine a target tracking point in the image to be processed that matches the reference tracking point, and the position information of the target tracking point; the second determination module 530 is used to determine the pose constraint corresponding to the visual sensor based on the position information of the target tracking point and the tracking information of the reference tracking point; the third determination module 540 is used to determine the state variables in the visual-inertial positioning system based on the pose constraint and the inertial information.

[0139] In some embodiments, the position information of the target tracking point is the normalized coordinates of the target tracking point; the first determining module is further configured to: perform feature point detection on the image to be processed to obtain at least two candidate feature points; determine the target tracking point and the image coordinates of the target tracking point from the at least two candidate feature points based on the tracking information of the reference tracking point; and transform the image coordinates of the target tracking point to obtain the normalized coordinates of the target tracking point.

[0140] In some embodiments, the tracking information of the reference tracking point includes preset three-dimensional information of the target tracking point and preset inverse depth of the target tracking point; the second determining module is further configured to: utilize the projection relationship of the vision sensor, and determine the pose error function of the vision sensor based on the position information, the three-dimensional information and the inverse depth; wherein, the pose error function is used to characterize the pose constraint corresponding to the vision sensor, and the pose error function is used to determine the movement information and rotation information of the vision sensor.

[0141] In some embodiments, the apparatus further includes: a second acquisition module, configured to acquire a preset number of historical processed images adjacent to the image to be processed; a fourth determination module, configured to determine, respectively, a historical tracking point in each historical processed image that matches the target tracking point, and the image coordinates of each historical tracking point; a third acquisition module, configured to acquire image regions of a preset scale from the historical processed images and the image to be processed based on the image coordinates of each historical tracking point and the image coordinates of the target tracking point; and a fifth determination module, configured to determine the confidence level of the pose constraint based on the pixel values ​​of pixels in the image region of each historical tracking point and the pixel values ​​of pixels in the image region of the target tracking point; wherein the confidence level of the pose constraint is used to characterize the weight of the pose constraint.

[0142] In some embodiments, the fifth determining module is further configured to: determine the average pixel value of the image region of each historical tracking point and the image region of the target tracking point; normalize the image region of each historical tracking point and the image region of the target tracking point based on the average pixel value of each image region to obtain the corresponding normalized image region; determine the root mean square between each normalized image region; and determine the confidence level of the pose constraint based on all the root mean squares.

[0143] In some embodiments, the visual-inertial positioning system includes an extended Kalman filter, and the state variables in the visual-inertial positioning system include a motion state variable and a rotation state variable; the third determining module is further configured to: perform partial derivative processing on the pose error function to obtain the Jacobian matrix corresponding to the motion information and the Jacobian matrix corresponding to the rotation information; based on the extended Kalman filter and the inertial information, process the Jacobian matrix corresponding to the motion information and the Jacobian matrix corresponding to the rotation information respectively to obtain the motion state variables and rotation state variables in the visual-inertial positioning system.

[0144] In some embodiments, the third determining module is further configured to: process the Jacobian matrix corresponding to the motion information and the Jacobian matrix corresponding to the rotation information based on the extended Kalman filter and the inertial information, respectively, to obtain the covariance matrix corresponding to the motion state quantity and the covariance matrix corresponding to the rotation state quantity; and determine the motion state quantity and the rotation state quantity based on the covariance matrix corresponding to the motion state quantity and the covariance matrix corresponding to the rotation state quantity.

[0145] The descriptions of the apparatus embodiments above are similar to those of the method embodiments above, and have similar beneficial effects. In some embodiments, the functions or modules included in the apparatus provided in this disclosure can be used to perform the methods described in the method embodiments above. For technical details not disclosed in the apparatus embodiments of this disclosure, please refer to the descriptions of the method embodiments of this disclosure for understanding.

[0146] It should be noted that, in the embodiments of this disclosure, if the above-described state determination method is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of this disclosure, or the part that contributes to related technologies, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this disclosure are not limited to any specific hardware, software, or firmware, or any combination of hardware, software, and firmware.

[0147] This disclosure provides a computer device including a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, it implements some or all of the steps in the above-described method.

[0148] This disclosure provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements some or all of the steps in the above-described method. The computer-readable storage medium may be transient or non-transient.

[0149] This disclosure provides a computer program including computer-readable code, wherein when the computer-readable code is executed in a computer device, a processor in the computer device performs some or all of the steps in the above-described method.

[0150] This disclosure provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, it implements some or all of the steps in the above-described method. This computer program product can be implemented specifically through hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium; in other embodiments, the computer program product is specifically embodied as a software product, such as a software development kit (SDK), etc.

[0151] It should be noted that the descriptions of the various embodiments above tend to emphasize the differences between them, while their similarities or commonalities can be referenced interchangeably. The descriptions of the above embodiments of the device, storage medium, computer program, and computer program product are similar to the descriptions of the above method embodiments and have similar beneficial effects. For technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of this disclosure, please refer to the descriptions of the method embodiments of this disclosure for understanding.

[0152] It should be noted that, Figure 6 This is a schematic diagram of a hardware entity of a computer device in an embodiment of this disclosure, such as... Figure 6 As shown, the hardware entity of the computer device 600 includes: a processor 601, a communication interface 602, and a memory 603, wherein:

[0153] Processor 601 typically controls the overall operation of computer device 600.

[0154] Communication interface 602 enables computer devices to communicate with other terminals or servers via a network.

[0155] The memory 603 is configured to store instructions and applications executable by the processor 601, and can also cache data to be processed or already processed (e.g., image data, audio data, voice communication data, and video communication data) of the processor 601 and various modules in the computer device 600. It can be implemented using flash memory or random access memory (RAM). Data transfer between the processor 601, the communication interface 602, and the memory 603 can be performed via bus 604.

[0156] It should be understood that the phrase "an embodiment" or "one embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this disclosure. Therefore, "in one embodiment" or "one embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this disclosure, the sequence numbers of the above steps / processes do not imply a sequential order of execution; the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this disclosure. The sequence numbers of the above embodiments of this disclosure are merely descriptive and do not represent the superiority or inferiority of the embodiments.

[0157] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0158] In the several embodiments provided in this disclosure, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components may be combined, or integrated into another system, or some features may be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed may be through some interfaces, and the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0159] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.

[0160] In addition, each functional unit in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0161] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.

[0162] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, or the part that contributes to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, magnetic disks, or optical disks.

[0163] The methods disclosed in the several method embodiments provided in this disclosure can be arbitrarily combined without conflict to obtain new method embodiments.

[0164] If the embodiments of this disclosure involve personal information, the products using these embodiments have clearly informed the users of the personal information processing rules and obtained their voluntary consent before processing the personal information. If the embodiments of this disclosure involve sensitive personal information, the products using these embodiments have obtained the individual's separate consent before processing the sensitive personal information, and the requirement of "express consent" is also met.

[0165] The above description is merely an embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for determining a state, characterized in that, include: During the operation of the visual inertial positioning system, the image to be processed, the tracking information of the reference tracking point, and the inertial information are acquired from the visual inertial positioning system. The image to be processed is acquired by the visual sensor in the visual inertial positioning system, the tracking information of the reference tracking point is determined based on the historical processed images acquired by the visual sensor, and the inertial information is acquired by the inertial sensor in the visual inertial positioning system. Determine the target tracking point in the image to be processed that matches the reference tracking point, and the position information of the target tracking point; Based on the position information of the target tracking point and the tracking information of the reference tracking point, the pose constraint corresponding to the visual sensor is determined; the tracking information of the reference tracking point includes the preset three-dimensional information of the target tracking point and the preset inverse depth of the target tracking point; Based on the pose constraints and the inertial information, the state variables in the visual-inertial positioning system are determined; The step of determining the pose constraint corresponding to the visual sensor based on the position information of the target tracking point and the tracking information of the reference tracking point includes: Using the projection relationship of the vision sensor, and based on the position information, the three-dimensional information, and the inverse depth, the pose error function of the vision sensor is determined; The pose error function is used to characterize the pose constraints corresponding to the vision sensor, and the pose error function is used to determine the movement and rotation information of the vision sensor.

2. The method according to claim 1, characterized in that, The location information of the target tracking point is the normalized coordinates of the target tracking point; determining the target tracking point in the image to be processed that matches the reference tracking point, and the location information of the target tracking point, includes: Feature point detection is performed on the image to be processed to obtain at least two candidate feature points; Based on the tracking information of the reference tracking point, the target tracking point and the image coordinates of the target tracking point are determined from the at least two candidate feature points; The image coordinates of the target tracking point are transformed to obtain the normalized coordinates of the target tracking point.

3. The method according to claim 1 or 2, characterized in that, The method further includes: Obtain a preset number of historical processed images adjacent to the image to be processed; Determine the historical tracking point that matches the target tracking point in each of the historical processed images, and the image coordinates of each historical tracking point; Based on the image coordinates of each historical tracking point and the image coordinates of the target tracking point, image regions of a preset scale are obtained from the historical processed image and the image to be processed, respectively. The confidence level of the pose constraint is determined based on the pixel values ​​of the pixels in the image region of each historical tracking point and the pixel values ​​of the pixels in the image region of the target tracking point; wherein the confidence level of the pose constraint is used to characterize the weight of the pose constraint.

4. The method according to claim 3, characterized in that, Determining the confidence level of the pose constraint based on the pixel values ​​of pixels in the image region of each historical tracking point and the pixel values ​​of pixels in the image region of the target tracking point includes: Determine the average pixel value of the image region of each historical tracking point and the image region of the target tracking point; Based on the pixel mean of each image region, the image region of each historical tracking point and the image region of the target tracking point are normalized to obtain the corresponding normalized image region. Determine the root mean square among the normalized image regions; The confidence level of the pose constraint is determined based on all the root mean squares.

5. The method according to claim 1, characterized in that, The visual-inertial positioning system includes an extended Kalman filter, and the state variables in the visual-inertial positioning system include translational state variables and rotational state variables; determining the state variables in the visual-inertial positioning system based on the pose constraints and the inertial information includes: By performing partial derivative processing on the pose error function, we obtain the Jacobian matrix corresponding to the movement information and the Jacobian matrix corresponding to the rotation information; Based on the extended Kalman filter and the inertial information, the Jacobian matrix corresponding to the motion information and the Jacobian matrix corresponding to the rotation information are processed respectively to obtain the motion state variables and rotation state variables in the visual inertial positioning system.

6. The method according to claim 5, characterized in that, The process of processing the Jacobian matrix corresponding to the motion information and the Jacobian matrix corresponding to the rotation information based on the extended Kalman filter and the inertial information respectively to obtain the motion state variables and rotation state variables in the visual-inertial positioning system includes: Based on the extended Kalman filter and the inertial information, the Jacobian matrix corresponding to the movement information and the Jacobian matrix corresponding to the rotation information are processed respectively to obtain the covariance matrix corresponding to the movement state quantity and the covariance matrix corresponding to the rotation state quantity. The movement state variable and the rotation state variable are determined based on the covariance matrix corresponding to the movement state variable and the covariance matrix corresponding to the rotation state variable.

7. A state determination device, characterized in that, include: The first acquisition module is used to acquire, during the operation of the visual inertial positioning system, an image to be processed, tracking information of a reference tracking point, and inertial information from the visual inertial positioning system; wherein, the image to be processed is acquired by a visual sensor in the visual inertial positioning system, the tracking information of the reference tracking point is determined based on historical processed images acquired by the visual sensor, and the inertial information is acquired by an inertial sensor in the visual inertial positioning system. The first determining module is used to determine the target tracking point in the image to be processed that matches the reference tracking point, and the position information of the target tracking point; The second determining module is used to determine the pose constraint corresponding to the visual sensor based on the position information of the target tracking point and the tracking information of the reference tracking point; the tracking information of the reference tracking point includes the preset three-dimensional information of the target tracking point and the preset inverse depth of the target tracking point; The third determining module is used to determine the state variables in the visual-inertial positioning system based on the pose constraints and the inertial information. The second determining module is further configured to utilize the projection relationship of the vision sensor to determine the pose error function of the vision sensor based on the position information, the three-dimensional information, and the inverse depth; wherein the pose error function is used to characterize the pose constraint corresponding to the vision sensor, and the pose error function is used to determine the movement information and rotation information of the vision sensor.

8. A computer device comprising a memory and a processor, the memory storing a computer program executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Positioning method and device adopting visual inertial data deep fusion

    CN109238277A

  • Optimization method and device for instant positioning and map building, medium and electronic equipment

    CN110322500A