Constraint construction method and apparatus for visual odometry, and device and medium

By acquiring the parameters of the supporting plane and image feature points, visual odometry estimation constraints are established, which solves the problem of scale divergence in visual odometry and improves the accuracy and precision of map data.

WO2026001580A1PCT designated stage Publication Date: 2026-01-02AUTONAVI SOFTWARE CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/098784
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-24
Filing Date
2025-06-03
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

In visual odometry, the pose estimation of the camera diverges over time, leading to a decrease in the accuracy of high-precision maps. Existing technologies struggle to effectively constrain scale, thus affecting the accuracy of map generation.

Method used

By acquiring the parameters of the support plane and image feature points, a matching relationship between image feature points is established, and a visual odometry estimation constraint term is generated. The pose of the support plane is used as a scale information that is not easily divergent to constrain the visual odometry estimation.

Benefits of technology

Effective scale constraints improve the accuracy of map data from visual mileage estimation, ensure scale stability, and enhance the generation accuracy of high-precision maps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025098784_02012026_PF_FP_ABST
    Figure CN2025098784_02012026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of vision, and specifically relates to a constraint construction method and apparatus for visual odometry, and a device and a medium. The method comprises: acquiring a parameter of a support plane; acquiring image feature points of at least two frames of images; establishing an image feature point matching relationship between image feature points of different frames of images; and on the basis of the image feature points, the image feature point matching relationship and the parameter of the support plane, generating a visual odometry constraint term. On the basis of the visual odometry constraint term acquired in the solution, the scale can be effectively constrained during visual odometry, thereby ensuring the stability of the scale, and facilitating an improvement in the accuracy of collecting map data on the basis of a visual odometry result.
Need to check novelty before this filing date? Find Prior Art

Description

Constraint construction method, device and equipment of visual odometry and medium

[0001] The present disclosure claims priority to the Chinese patent application No. 202410818552.7, filed on June 24, 2024, entitled "Constraint construction method, device and equipment of visual odometry and medium", the whole content of which is incorporated herein by reference. TECHNICAL FIELD

[0002] The present disclosure relates to the field of vision technology, in particular to a constraint construction method, device and equipment of visual odometry and medium. BACKGROUND

[0003] The scheme of making high-precision maps by relying on point cloud data collected by laser radar equipment has the problem of high cost and cannot meet the high-frequency update demand of high-precision maps. Therefore, the industry has gradually begun to use low-cost visual schemes, that is, relying on image data collected by visual equipment (such as a camera), completing the collection and high-frequency update of high-precision maps through VIO (Visual Inertial Odometry, visual inertial odometry), GPS (Global Positioning System, global positioning system) and three-dimensional reconstruction.

[0004] To ensure the accuracy of high-precision maps, the visual scheme needs to rely on visual odometry algorithm to estimate the pose of the camera, but the projection transformation of camera imaging cannot obtain the scale information of the real world, so the visual odometry algorithm needs to artificially set the scale. The present inventors found that with the passage of time, the scale will gradually diverge, resulting in inaccurate estimation of the pose of the camera, which in turn affects the accuracy of the high-precision map finally generated. Therefore, how to effectively constrain the scale in visual odometry estimation to ensure the stability of the scale is a problem that needs to be solved by those skilled in the art. SUMMARY

[0005] To solve the problems in the related art, the present disclosure provides a constraint construction method, device and equipment of visual odometry and medium.

[0006] In a first aspect, the present disclosure provides a constraint construction method of visual odometry, which comprises:

[0007] Obtaining parameters of a support plane, the support plane being used to carry an image acquisition device;

[0008] Obtaining image feature points of two or more images, the image feature points belonging to the support plane, and the images being collected by the image acquisition device;

[0009] An image feature point matching relationship between image feature points of different frames of images is established;

[0010] A visual odometry constraint term is generated according to the image feature points, the image feature point matching relationship and the parameters of the support plane, and the visual odometry constraint term is used to constrain a scale in visual odometry.

[0011] In an embodiment of the present disclosure, the parameters of the support plane include a height of the image acquisition device to the support plane, a direction of a normal vector of the support plane in a coordinate system of the image acquisition device, and a covariance of the height and the direction.

[0012] In an embodiment of the present disclosure, generating the visual odometry constraint term according to the image feature points, the image feature point matching relationship and the parameters of the support plane comprises:

[0013] An observed depth of the image feature point is obtained according to the parameters of the support plane, the image feature point and an intrinsic parameter of the image acquisition device;

[0014] A depth estimation error constraint of the image feature point is established based on the image feature point, the image feature point matching relationship and the observed depth of the image feature point;

[0015] The visual odometry constraint term is generated, and the visual odometry constraint term at least includes the depth estimation error constraint.

[0016] In an embodiment of the present disclosure, the method further comprises:

[0017] An initial pose when the image acquisition device acquires a first frame of image in the two or more frames of images and an initial spatial position of the image feature point of the first frame of image are obtained;

[0018] The depth estimation error constraint of the image feature point is established based on the image feature point, the image feature point matching relationship and the observed depth of the image feature point, and the depth estimation error constraint comprises:

[0019] The depth estimation error constraint of the image feature point is established based on the initial pose of the image acquisition device, the initial spatial position of the image feature point of the first frame of image, the image feature point, the image feature point matching relationship and the observed depth of the image feature point.

[0020] In an embodiment of the present disclosure, the method further comprises:

[0021] The pose when the image acquisition device acquires the two or more frames of images and the spatial position of the image feature point in the two or more frames of images are converted to a two-dimensional image coordinate system through a projection equation;

[0022] The image coordinate estimation error constraint is obtained based on the image feature points, the pose of the image acquisition device when collecting more than two images in a two-dimensional image coordinate system, and the spatial positions of the image feature points in the more than two images in the two-dimensional image coordinate system.

[0023] The visual odometry estimation constraint term is generated and includes at least the depth estimation error constraint.

[0024] The visual odometry estimation constraint term is generated and includes at least the image coordinate estimation error constraint and the depth estimation error constraint.

[0025] In an embodiment of the present disclosure, before the pose of the image acquisition device when collecting more than two images and the spatial positions of the image feature points in the more than two images are converted to the two-dimensional image coordinate system by the projection equation, the method further includes:

[0026] The projection equation is established according to the initial pose of the image acquisition device, the initial spatial positions of the image feature points, and the intrinsic parameters of the image acquisition device.

[0027] In an embodiment of the present disclosure, the method further includes:

[0028] The visual odometry estimation constraint term is taken as a constraint on the pose of the image acquisition device when collecting corresponding images and the spatial positions of the image feature points in the corresponding images based on a bundle adjustment (BA) problem, to obtain a visual odometry estimation model.

[0029] In a second aspect, an embodiment of the present disclosure provides a constraint construction device for a visual odometry, and the device includes:

[0030] A parameter acquisition module is configured to acquire parameters of a support plane, and the support plane is used to support an image acquisition device.

[0031] A feature point acquisition module is configured to acquire image feature points of more than two images, and the image feature points belong to the support plane, and the images are collected by the image acquisition device.

[0032] A relationship establishment module is configured to establish an image feature point matching relationship between the image feature points of different images.

[0033] A constraint generation module is configured to generate a visual odometry estimation constraint term according to the image feature points, the image feature point matching relationship, and the parameters of the support plane, and the visual odometry estimation constraint term is used to constrain a scale in visual odometry estimation.

[0034] In a third aspect, an electronic device is provided, including a memory and a processor, wherein the memory is configured to store one or more computer instructions, and the one or more computer instructions are configured to be executed by the processor to implement the method according to any one of the first aspect.

[0035] In a fourth aspect, a computer-readable storage medium is provided, and the computer-readable storage medium stores computer instructions, and the computer instructions are configured to be executed by a processor to implement the method according to any one of the first aspect.

[0036] In a fifth aspect, a computer program product is provided, and the computer program product includes computer instructions, and the computer instructions are configured to be executed by a processor to implement the method according to the first aspect.

[0037] According to the technical scheme provided by the embodiments of the present disclosure, by obtaining the parameters of the support plane, obtaining the image feature points of two or more images, establishing the image feature point matching relationship between the image feature points of different images, and generating the visual odometry constraint term according to the image feature points, the image feature point matching relationship and the parameters of the support plane, wherein, considering the characteristics that the distance and direction of the image acquisition device to the support plane remain basically unchanged in a period of time, the pose of the support plane can be taken as the scale information which is not easy to diverge in the visual odometry process, therefore, based on the visual odometry constraint term obtained in the above scheme, the scale can be effectively constrained in the visual odometry process based on the scale information which is not easy to diverge, so as to ensure the stability of the scale, and help to improve the accuracy of collecting map data based on the visual odometry result.

[0038] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0039] Other features, objects, and advantages of the present disclosure will become more apparent from the following detailed description when read in conjunction with the accompanying drawings. In the drawings:

[0040] FIG. 1 shows a flowchart of a constraint construction method of a visual odometer according to an embodiment of the present disclosure.

[0041] FIG. 2 shows a schematic diagram of an image acquisition device according to an embodiment of the present disclosure.

[0042] FIG. 3 shows a structural block diagram of a constraint construction device of a visual odometer according to an embodiment of the present disclosure.

[0043] FIG. 4 shows a structural block diagram of an electronic device according to an embodiment of the present disclosure.

[0044] FIG. 5 shows a structural schematic diagram of a computer system suitable for implementing the method according to the embodiments of the present disclosure. DETAILED DESCRIPTION

[0045] Hereinafter, exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings so as to be easily implemented by those skilled in the art. Also, portions irrelevant to the description of the exemplary embodiments are omitted in the accompanying drawings for the sake of clarity.

[0046] In the present disclosure, it should be understood that terms such as "include" or "have" are intended to indicate that there are features, numbers, steps, actions, components, parts or combinations thereof disclosed in the specification, and do not exclude the possibility that one or more other features, numbers, steps, actions, components, parts or combinations thereof exist or are added.

[0047] It should also be noted that the embodiments in the present disclosure and the features in the embodiments can be combined with each other without conflict. The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.

[0048] In the present disclosure, if the operation of acquiring user information or user data or the operation of showing user information or user data to others is involved, the operation is an operation authorized, confirmed by the user, or actively selected by the user.

[0049] In order to ensure the accuracy of the high-precision map, the visual scheme needs to rely on the visual odometry calculation method to estimate the pose of the camera. In order to obtain the scale information of the real world, the visual odometry calculation method needs to artificially set the scale. For example, the distance between two actual objects in the image can be artificially calibrated, so that the distance serves as the scale information in the visual odometry calculation method.

[0050] However, considering that the artificially given scale information will gradually diverge over time. For example, even if the distance between two actual objects in the image is artificially calibrated, the distance between the two actual objects may change over time, in which case, even if the distance is used as scale information, the accuracy of the estimated pose of the camera based on the scale information cannot be guaranteed, which further affects the accuracy of the finally generated high-precision map.

[0051] In summary, how to effectively constrain the scale in visual odometry estimation and ensure the stability of the scale is a problem that needs to be solved by those skilled in the art.

[0052] To solve the above problems, the embodiments of the present disclosure provide a constraint construction method, device, equipment and medium for visual odometry.

[0053] According to the technical scheme provided by the embodiment of the present disclosure, the parameters of the support plane are obtained, the image feature points of two or more images are obtained, the image feature point matching relationship between the image feature points of different images is established, and the visual odometry constraint term is generated according to the image feature points, the image feature point matching relationship and the parameters of the support plane. Wherein, considering the characteristics that the distance and direction from the image acquisition device to the support plane carrying the image acquisition device remain basically unchanged in a period of time, the pose of the support plane can be taken as the scale information which is not easy to diverge in the visual odometry process, and therefore, the visual odometry constraint term obtained in the above scheme can effectively constrain the scale in the visual odometry process, thereby ensuring the stability of the scale and helping to improve the accuracy of collecting map data based on the visual odometry result.

[0054] FIG. 1 shows a flowchart of a constraint construction method of a visual odometer according to an embodiment of the present disclosure. As shown in FIG. 1, the constraint construction method of the visual odometer includes the following steps S101-S104:

[0055] In step S101, the parameters of the support plane are obtained.

[0056] Wherein, the support plane is used to carry (support) the image acquisition device.

[0057] For example, when a vehicle used for image acquisition is driving on a road, the road surface can be understood as the support plane; when a track vehicle used for image acquisition is driving on a track, the upper surface of the track can be understood as the support plane. In an implementation manner of the present disclosure, the parameters of the support plane can be understood as parameters used to indicate the spatial position of the support plane, and the spatial position can be the spatial position in the world coordinate system.

[0058] In step S102, the image feature points of two or more images are obtained.

[0059] Wherein, the image is collected by the image acquisition device, and the image feature point belongs to the support plane, which can be understood as a pixel point in the image recording the support plane.

[0060] In step S103, the image feature point matching relationship between the image feature points of different images is established.

[0061] In step S104, the visual odometry constraint term is generated according to the image feature points, the image feature point matching relationship and the parameters of the support plane.

[0062] Wherein, the visual odometry constraint term is used to constrain the scale in the visual odometry.

[0063] In an embodiment of the present disclosure, obtaining the image feature points belonging to the support plane can be understood as image processing the image, and determining the image feature points belonging to the support plane according to the image processing result. For example, the image can be processed based on an image pixel classification or a semantic segmentation algorithm, and the semantic segmentation algorithm can be a full-pixel semantic segmentation algorithm.

[0064] In an embodiment of the present disclosure, establishing the image feature point matching relationship between the image feature points of different frames of images can be understood as establishing a feature point matching relationship for the image feature points corresponding to the same object point in the real world in different frames of images. For example, a descriptor vector of the image feature points in different frames of images can be obtained, the descriptor vector can be understood as containing image information of the corresponding image feature point and the surrounding area of the image feature point, and if the similarity rate of the descriptor vectors of two image feature points is greater than or equal to a similarity threshold, a feature point matching relationship can be established for the two image feature points.

[0065] In an embodiment of the present disclosure, generating the visual odometry constraint term according to the image feature points, the image feature point matching relationship, and the parameters of the support plane can be understood as substituting the coordinates of the image feature points, the image feature point matching relationship, and the parameters of the support plane into a pre-obtained algorithm to calculate to generate the visual odometry constraint term. It can also be understood as obtaining a pre-trained visual odometry model, inputting the coordinates of the image feature points, the image feature point matching relationship, and the parameters of the support plane as inputs into the visual odometry model to obtain the visual odometry constraint term output by the visual odometry model.

[0066] According to the technical scheme provided by the embodiments of the present disclosure, the parameters of the support plane are obtained, the image feature points of two or more frames of images are obtained, the image feature point matching relationship between the image feature points of different frames of images is established, and the visual odometry constraint term is generated according to the image feature points, the image feature point matching relationship, and the parameters of the support plane. Considering that the distance and direction of the image acquisition device to the support plane carrying the image acquisition device remain basically unchanged in a period of time, the pose of the support plane can be regarded as scale information that is not easy to diverge in the visual odometry process. Therefore, based on the visual odometry constraint term obtained in the above scheme, the scale implementation can be effectively constrained in the visual odometry process, thereby ensuring the stability of the scale and helping to improve the accuracy of collecting map data based on the visual odometry result.

[0067] In an embodiment of the present disclosure, the parameters of the support plane include the height of the image acquisition device to the support plane, the direction of the normal vector of the support plane in the image acquisition device coordinate system, and the covariance of the height and the direction.

[0068] In one implementation of this disclosure, the height of the image acquisition device from the supporting plane can be understood as the distance from the center of the image acquisition device to the supporting plane. The center of the image acquisition device can be understood as the optical center of the lens in the image acquisition device, or as the center of the photosensitive device of the image acquisition device.

[0069] In one implementation of this disclosure, the coordinate system of the image acquisition device can be understood as a spatial coordinate system with the center of the image acquisition device as the pole.

[0070] In one implementation of this disclosure, the height of the image acquisition device from the supporting plane and the direction of the normal vector of the supporting plane in the coordinate system of the image acquisition device can be obtained through manual calibration or online calibration. Optionally, the height of the image acquisition device from the supporting plane and the direction of the normal vector of the supporting plane in the coordinate system of the image acquisition device can be obtained by using a checkerboard calibration method.

[0071] Figure 2 shows a schematic diagram of an image acquisition device according to an embodiment of the present disclosure. As shown in Figure 2, a vehicle 212 equipped with the image acquisition device 202 can be used to drive on a road and perform image acquisition. A support plane 201 is used to support the vehicle 212 (equivalent to supporting the image acquisition device 202). The parameters of the support plane include the height from the optical center of the lens of the image acquisition device 202 on the vehicle 212 to the support plane 201 in the camera coordinate system. The direction of the normal vector of the supporting plane 201 in the camera coordinate system. and height With normal vector direction covariance Σ g .

[0072] It should be noted that the direction of the normal vector It can be obtained through angle values To uniquely represent, Normal vector direction The coordinates in polar coordinates, where T is the transpose matrix operator.

[0073] Wherein, the direction of the normal vector and The transformation relationship can be expressed by equation (1):

[0074] According to the technical scheme provided in the embodiments of the present disclosure, by limiting the parameters of the support plane to include the height of the image acquisition device to the support plane, the direction of the normal vector of the support plane in the image acquisition device coordinate system, and the covariance of the height and the direction, the composition of the parameters of the support plane can be simplified, the difficulty of obtaining the parameters of the support plane can be reduced, and the calculation complexity can be reduced.

[0075] In an embodiment of the present disclosure, in step S104, the visual odometry constraint term is generated according to the image feature points, the image feature point matching relationship, and the parameters of the support plane, which can be achieved by the following steps:

[0076] According to the parameters of the support plane, the image feature points, and the intrinsic parameters of the image acquisition device, the observed depth of the image feature points is obtained.

[0077] Based on the image feature points, the image feature point matching relationship, and the observed depth of the image feature points, the depth estimation error constraint of the image feature points is established.

[0078] The visual odometry constraint term including at least the depth estimation error constraint is generated.

[0079] In an embodiment of the present disclosure, according to the parameters of the support plane, the image feature points, and the intrinsic parameters of the image acquisition device, the observed depth of the image feature points can be understood as being calculated by substituting the parameters of the support plane, the coordinates of the image feature points, and the intrinsic parameters of the image acquisition device into a pre-obtained algorithm to obtain the observed depth of the image feature points; or it can also be understood as obtaining a pre-trained observed depth model, and inputting the parameters of the support plane, the coordinates of the image feature points, and the intrinsic parameters of the image acquisition device as inputs into the observed depth model to obtain the observed depth of the image feature points.

[0080] In an embodiment of the present disclosure, the depth estimation error constraint of the image feature points belonging to the support plane is established based on the image coordinates, the matching relationship, and the observed depth of the image feature points belonging to the support plane, which can be understood as obtaining the error between the estimated depth of the image feature points and the observed depth of the image feature points, and effectively constraining the scale implementation in the visual odometry process based on the error.

[0081] For example, the observed depth of the jth image feature point in the ith frame image is obtained.

[0082] The observed depth of the jth image feature point in the ith frame image which can be represented by formula (2):

[0083] wherein K is the intrinsic parameter of the image acquisition device, can be used to indicate the coordinate of the jth image feature point in the ith image in a two-dimensional image coordinate system, can be uniquely represented by (u, v) T , where (u, v) is the coordinate of the jth image feature point in the ith image in a two-dimensional image coordinate system.

[0084] Based on the observation depth The depth estimation error constraint as shown in equation (3) can be obtained:

[0085] wherein, is the estimated depth of the jth image feature point in the ith image, wherein can be represented by equation (4):

[0086] wherein, is the pose of the image acquisition device when acquiring the ith image, is the spatial position of the jth image feature point in the image acquisition device coordinate system.

[0087] Σ dep is the confidence of the observation depth of the jth image feature point in the ith image, Σ dep can be represented by equation (5):

[0088] J g is the partial derivative of the observation depth of the jth image feature point in the ith image with respect to , J p is the partial derivative of the observation depth of the jth image feature point in the ith image with respect to , Σ p is the covariance of the horizontal coordinate and the vertical coordinate of the jth image feature point in the ith image in a two-dimensional image coordinate system;

[0089] J g can be represented by equation (6):

[0090] δ is the differential operator symbol;

[0091] J p can be represented by equation (7):

[0092] can be represented by equation (8):

[0093] is the set of feature points belonging to the support plane among the j image feature points of the ith image.

[0094] Based on the depth estimation error constraint as shown in formula (3), a visual odometry constraint term including at least the depth estimation error constraint can be obtained, which is shown in formula (9):

[0095] In the visual odometry constraint term, the more accurate the estimated and are, the closer and are, and the smaller is, and the smaller the visual odometry constraint term is.

[0096] According to the technical scheme provided by the embodiments of the present disclosure, the observation depth of the image feature point is obtained according to the parameters of the support plane, the image feature point and the internal parameters of the image acquisition device; the depth estimation error constraint of the image feature point is established based on the image feature point, the matching relationship of the image feature point and the observation depth of the image feature point; and the visual odometry constraint term including at least the depth estimation error constraint is generated. By constraining the visual odometry model through the visual odometry constraint term including at least the depth estimation error constraint, the scale can be constrained by means of the error of the observation depth of the image feature point, which helps to further ensure the stability of the scale in the visual odometry process.

[0097] In an embodiment of the present disclosure, the method further comprises:

[0098] obtaining an initial pose of the image acquisition device when capturing the first image in the two or more images and an initial spatial position of the image feature point of the first image;

[0099] establishing the depth estimation error constraint of the image feature point based on the image feature point, the matching relationship of the image feature point and the observation depth of the image feature point, comprising:

[0100] establishing the depth estimation error constraint of the image feature point based on the initial pose of the image acquisition device, the initial spatial position of the image feature point of the first image, the image feature point, the matching relationship of the image feature point and the observation depth of the image feature point.

[0101] In an embodiment of the present disclosure, obtaining the initial pose of the image acquisition device when capturing the first image in the two or more images can be understood as obtaining the initial pose through a positioning device such as an Inertial Measurement Unit (IMU) or a Global Positioning System (GPS) matched with the image acquisition device, or can be understood as receiving the initial pose sent by other devices or systems.

[0102] In an embodiment of the present disclosure, the initial spatial position of the image feature point of the first frame image when the image acquisition device acquires more than two frames of images can be understood as being acquired by a spatial position measuring device matched with the image acquisition device, such as a laser radar, or can be understood as receiving an initial spatial position sent by other devices or systems.

[0103] According to the technical scheme provided by the embodiment of the present disclosure, the initial pose when the image acquisition device acquires more than two frames of images and the initial spatial position of the image feature point of the first frame image are acquired, and based on the initial pose of the image acquisition device, the initial spatial position of the image feature point of the first frame image, the image feature point, the matching relationship of the image feature point, and the observation depth of the image feature point, the depth estimation error constraint of the image feature point is established, the initial value of the depth estimation error constraint is provided, the computational complexity of the visual odometry based on the visual odometry model is effectively reduced, and the efficiency is improved.

[0104] In an embodiment of the present disclosure, the method further comprises:

[0105] The pose of the image acquisition device when acquiring more than two frames of images and the spatial position of the image feature point in the more than two frames of images are converted to a two-dimensional image coordinate system by a projection equation;

[0106] Based on the image feature point, the pose of the image acquisition device when acquiring more than two frames of images in the two-dimensional image coordinate system, and the spatial position of the image feature point in the more than two frames of images in the two-dimensional image coordinate system, an image coordinate estimation error constraint is obtained;

[0107] The visual odometry constraint term including at least the depth estimation error constraint is obtained, including:

[0108] The visual odometry constraint term including at least the image coordinate estimation error constraint and the depth estimation error constraint is generated.

[0109] In an embodiment of the present disclosure, before the pose of the image acquisition device when acquiring more than two frames of images and the spatial position of the image feature point in the more than two frames of images are converted to a two-dimensional image coordinate system by a projection equation, the projection equation can be established according to the initial pose of the image acquisition device, the initial spatial position of the image feature point, and the intrinsic parameter of the image acquisition device. The projection equation can map a three-dimensional spatial position to a two-dimensional image space. This embodiment can reduce the difficulty of obtaining the projection equation.

[0110] For example, the pose of the image acquisition device when acquiring the i-th frame of image in more than two frames of images and the spatial position of the image feature point in the more than two frames of images In the two-dimensional image coordinate system, the following can be obtained Pi can be understood as a projection symbol, where, is used to indicate the pose of the image acquisition device when acquiring the i-th image in the two-dimensional image coordinate system and the spatial position of the j-th image feature point in the image acquisition device coordinate system.

[0111] Based on and The image coordinate estimation error constraint as shown in equation (10) can be obtained:

[0112] where, when The more accurate, and The closer, the smaller the image coordinate estimation error constraint.

[0113] Based on the image coordinate estimation error constraint as shown in equation (10) and the depth estimation error constraint as shown in equation (3), the visual odometry estimation constraint term as shown in equation (11) can be generated:

[0114] Wherein, the visual odometry estimation constraint term as shown in equation (11) can be used as a constraint on the pose of the image acquisition device when acquiring the corresponding image and the spatial position of the image feature point in the corresponding image based on the Bundle Adjustment (BA) problem construction, to obtain a visual odometry model, which can be as shown in equation (12):

[0115] wherein, represents a set of poses of the image acquisition device when acquiring i-th images, represents a set of maximum likelihood estimates of the pose of the image acquisition device when acquiring i-th images, represents a set of spatial positions of j image feature points in the image acquisition device coordinate system, represents a set of maximum likelihood estimates of the spatial positions of j image feature points in the image acquisition device coordinate system, represents other factors used in the bundle adjustment factor graph of the visual odometry.

[0116] By solving the unconstrained nonlinear optimization problem of the visual odometry estimation model as shown in equation (14) (for example, the Levenberg-Marquardt (L-M) method can be used for solving), the pose of the image acquisition device can be obtained.

[0117] According to the technical scheme provided by the embodiment of the present disclosure, the spatial positions of the image feature points in the two or more images and the pose of the image acquisition device when collecting the two or more images are converted to the two-dimensional image coordinate system through the projection equation; the image coordinate estimation error constraint is obtained based on the image feature points, the pose of the image acquisition device when collecting the two or more images in the two-dimensional image coordinate system, and the spatial positions of the image feature points in the two or more images in the two-dimensional image coordinate system; the visual odometry estimation constraint term including at least the image coordinate estimation error constraint and the depth estimation error constraint is generated; and then the visual odometry estimation constraint term can be used as the constraint on the spatial positions of the image feature points in the corresponding images and the pose of the image acquisition device when collecting the corresponding images based on the bundle adjustment (BA) problem to obtain the visual odometry estimation model. In the above scheme, the image coordinate estimation error constraint is introduced to further constrain the scale in the visual odometry estimation process, the dimension for constraining the scale is increased, and the stability of the scale is further ensured, which helps to improve the accuracy of collecting map data based on the visual odometry estimation result.

[0118] FIG. 3 shows a structural block diagram of a constraint construction device of a visual odometer according to an embodiment of the present disclosure. The device can be realized by software, hardware, or a combination of both as part or all of an electronic device.

[0119] As shown in FIG. 3, the constraint construction device 300 of the visual odometer includes:

[0120] The parameter acquisition module 301 is configured to acquire parameters of a support plane, the support plane being used to carry an image acquisition device;

[0121] The feature point acquisition module 302 is configured to acquire image feature points of two or more images, the image feature points belonging to the support plane, and the images being collected by the image acquisition device;

[0122] The relationship establishment module 303 is configured to establish an image feature point matching relationship between the image feature points of different images;

[0123] The constraint generation module 304 is configured to generate a visual odometry estimation constraint term according to the image feature points, the image feature point matching relationship, and the parameters of the support plane, the visual odometry estimation constraint term being used to constrain the scale in visual odometry estimation.

[0124] The above various modules execute specific embodiments of corresponding processes, which are described in the foregoing method embodiment related parts, and will not be expanded here.

[0125] The present disclosure also discloses an electronic device, and FIG. 4 shows a structural block diagram of an electronic device according to an embodiment of the present disclosure.

[0126] As shown in FIG. 4, the electronic device includes a memory and a processor, wherein the memory is configured to store one or more computer instructions, and the one or more computer instructions are executed by the processor to implement the method for constraint construction of visual odometry provided by the foregoing embodiments of the present disclosure. Please refer to the foregoing embodiments for details, which will not be repeated here.

[0127] FIG. 5 shows a structural schematic diagram of a computer system suitable for implementing the method according to the embodiments of the present disclosure.

[0128] As shown in FIG. 5, the computer system includes a processing unit, which can execute various methods in the above embodiments according to programs stored in a Read-Only Memory (ROM) or loaded from a storage part into a Random Access Memory (RAM). Various programs and data required for the operation of the computer system are also stored in the RAM. The processing unit, the ROM, and the RAM are connected to each other through a bus. An Input / Output (I / O) interface is also connected to the bus.

[0129] The following components are connected to the I / O interface: an input part including a keyboard, a mouse, etc.; an output part including a Cathode Ray Tube (CRT), a Liquid Crystal Display (LCD), etc., and a speaker, etc.; a storage part including a hard disk, etc.; and a communication part including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication part performs a communication process via a network such as the Internet. A drive is also connected to the I / O interface as needed. A removable medium such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive as needed, so that a computer program read therefrom is installed in the storage part as needed. The processing unit can be implemented as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a TPU (Tensor Processing Unit), a FPGA (Field Programmable Gate Array), a NPU (Neural network Processing Unit), etc.

[0130] In particular, the method described above can be implemented as a computer software program according to embodiments of the present disclosure. For example, embodiments of the present disclosure include a computer program product comprising computer instructions tangibly embodied on a machine-readable medium, the computer instructions comprising program code for executing the method described above. In such embodiments, the computer program can be downloaded and installed from a network via a communication part, and / or installed from a detachable medium.

[0131] The flow and block diagrams in the drawings show the architectural, functional and operational aspects of possible implementations of systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow and block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations thereof, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0132] The units or modules described in the embodiments of the present disclosure can be implemented by means of software, or by means of programmable hardware. The described units or modules can also be provided in a processor, and the names of these units or modules do not constitute a limitation on the units or modules themselves in some cases.

[0133] As another aspect, the present disclosure also provides a computer readable storage medium, which can be the computer readable storage medium contained in the electronic device or computer system in the above embodiments; or can exist separately, and is not assembled into the device. The computer readable storage medium stores one or more programs, which are used by one or more processors to execute the method described in the present disclosure.

[0134] The above description is merely that of the preferred embodiments of the present disclosure and a description of the technical principles of the present disclosure. It should be understood by those skilled in the art that the inventive scope involved in the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by the combinations of the above technical features or equivalent features without departing from the inventive concept. For example, the technical solutions formed by the mutual replacement of the above features and the technical features with similar functions disclosed in the present disclosure (but not limited to) without departing from the inventive concept.

Claims

1. A constraint construction method for visual odometry, wherein, include: Obtain the parameters of the supporting plane, which is used to support the image acquisition device; Acquire image feature points from two or more frames of images, wherein the image feature points belong to the supporting plane, and the images are acquired by the image acquisition device; Establish image feature point matching relationships between image feature points of different frames; Based on the image feature points, the image feature point matching relationships, and the parameters of the supporting plane, a visual odometry estimation constraint term is generated, which is used to constrain the scale in visual odometry estimation.

2. The method according to claim 1, wherein, The parameters of the supporting plane include the height of the image acquisition device from the supporting plane, the direction of the normal vector of the supporting plane in the coordinate system of the image acquisition device, and the covariance between the height and the direction.

3. The method according to claim 1 or 2, wherein, The step of generating visual odometry estimation constraints based on the image feature points, the image feature point matching relationships, and the parameters of the supporting plane includes: The observation depth of the image feature points is obtained based on the parameters of the supporting plane, the image feature points, and the intrinsic parameters of the image acquisition device. Based on the image feature points, the image feature point matching relationship, and the observation depth of the image feature points, a depth estimation error constraint for the image feature points is established. Generate visual odometry estimation constraints that include at least the depth estimation error constraints.

4. The method according to claim 3, wherein, The method further includes: The initial pose of the first frame of the image acquisition device when it acquires the two or more frames of images, and the initial spatial position of the image feature points of the first frame image are obtained. The step of establishing depth estimation error constraints for image feature points based on the image feature points, the matching relationships between the image feature points, and the observation depth of the image feature points includes: Based on the initial pose of the image acquisition device, the initial spatial position of the image feature points of the first frame image, the image feature points, the matching relationship of the image feature points, and the observation depth of the image feature points, a depth estimation error constraint for the image feature points is established.

5. The method according to claim 4, wherein, The method further includes: The pose of the image acquisition device when acquiring two or more frames of images and the spatial position of image feature points in the two or more frames of images are transformed into a two-dimensional image coordinate system by using the projection equation. Based on the image feature points, the pose of the image acquisition device when acquiring two or more frames of images in the two-dimensional image coordinate system, and the spatial position of the image feature points in the two or more frames of images in the two-dimensional image coordinate system, the image coordinate estimation error constraint is obtained. The generation of visual odometry estimation constraints, which includes at least the depth estimation error constraints, includes: Generate a visual odometry estimation constraint term that includes at least the image coordinate estimation error constraint and the depth estimation error constraint.

6. The method according to claim 5, wherein, Before transforming the pose of the image acquisition device when acquiring two or more frames of images and the spatial positions of image feature points in the two or more frames of images to a two-dimensional image coordinate system using the projection equation, the method further includes: The projection equation is established based on the initial pose of the image acquisition device, the initial spatial position of the image feature points, and the intrinsic parameters of the image acquisition device.

7. The method according to any one of claims 1-6, wherein, The method further includes: The visual odometry estimation constraint term is used as a constraint on the pose of the image acquisition device when acquiring the corresponding image and the spatial position of the image feature points in the corresponding image, which is constructed based on the bundle adjustment (BA) problem, to obtain the visual odometry estimation model.

8. A constraint construction device for a visual odometry, wherein, include: The parameter acquisition module is configured to acquire parameters of the support plane, which is used to support the image acquisition device; The feature point acquisition module is configured to acquire image feature points from two or more frames of images, wherein the image feature points belong to the supporting plane, and the images are acquired by the image acquisition device. The relationship establishment module is configured to establish image feature point matching relationships between image feature points of different frames; The constraint generation module is configured to generate visual odometry estimation constraint terms based on the image feature points, the image feature point matching relationships, and the parameters of the support plane. The visual odometry estimation constraint terms are used to constrain the scale in visual odometry estimation.

9. An electronic device, wherein, It includes a memory and a processor; the memory is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the method of any one of claims 1-7.

10. A computer-readable storage medium having stored thereon computer instructions, wherein, When executed by a processor, the computer instructions implement the method of any one of claims 1-7.

11. A computer program product, wherein, Includes computer instructions that, when executed by a processor, implement the method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Real scale obtaining method of monocular vision odometer

    CN105976402A

  • Method and device for establishing beacon map based on visual beacons

    CN112183171A

  • Monocular visual odometer scale recovery method based on optimized angular point screening road surface points

    CN117557592A

  • Visual odometer method and device fused with depth vision constraint, terminal and medium

    CN118052872A