A method for constructing a passive reference based on scene features

By mounting a monocular camera on the motion platform and constructing a reference coordinate system using scene features, image processing and optimization are performed, solving the problem of unified reference for the motion platform. This enables stable and accurate self-construction of reference in indoor and outdoor environments, and is applicable to motion platforms such as drones and mobile robots.

CN116385563BActive Publication Date: 2026-01-02SUN YAT SEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310510553.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-08
Publication Date
2026-01-02
Estimated Expiration
2043-05-08

AI Technical Summary

Technical Problem

Existing technologies make it difficult to build a stable and accurate unified benchmark on a moving platform. Traditional positioning technologies such as GPS are limited in indoor use. Monocular cameras lack depth information, and the accuracy of binocular or depth camera calculations is affected by the observation distance. Methods that combine visual and inertial navigation information have low calculation accuracy.

Method used

The passive benchmark self-construction method based on scene features on a moving platform acquires image data by mounting a monocular camera on the moving platform, establishes a benchmark coordinate system using cooperative markers, performs preliminary camera extrinsic parameter estimation, feature extraction and matching, triangulation and reprojection error optimization, generates scene point clouds, and finally achieves benchmark alignment.

Benefits of technology

It achieves stable and accurate self-construction of benchmarks applicable to both indoor and outdoor environments. The method is flexible, convenient, and low-cost, and is suitable for motion platforms such as drones and mobile robots. It can be applied to indoor operations of industrial robots and outdoor inspection and surveying of drones.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116385563B_ABST
    Figure CN116385563B_ABST
Patent Text Reader

Abstract

The application provides a moving platform passive reference self-construction method based on scene features, which comprises the following steps: based on the two-dimensional image coordinates of each cooperative mark in two images and the three-dimensional world coordinates of each cooperative mark in a reference coordinate system, completing the preliminary camera external parameter estimation corresponding to the two images; combining the preliminary camera external parameter estimation results of the two images, performing triangulation on all homonymous feature point pairs in the two images, estimating the preliminary three-dimensional coordinates of the three-dimensional space feature points corresponding to each homonymous feature point pair, and generating an initial three-dimensional point cloud of the scene; constructing a re-projection error objective function based on the preliminary three-dimensional coordinates of the three-dimensional space feature points and the image coordinates of the three-dimensional space feature points in the image planes of the two images and solving the re-projection error objective function, obtaining the optimized camera external parameter estimation results and the scene point cloud; and performing pose estimation based on the optimized scene point cloud, so as to realize the reference alignment of the moving platform.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application mainly relates to the field of radar signal processing, and particularly relates to a passive reference self-construction method for a moving platform based on scene features. BACKGROUND

[0002] In the field of outdoor inspection and survey, a moving platform such as a UAV or an unmanned vehicle is required to collect image information of a target to be measured within a specified range, and an observation result is obtained through data processing. In order to ensure the observation accuracy, the moving platform needs to approach the target to collect local information of the target, and then the overall information of the target is obtained through data fusion. Since the moving platform itself is in a moving state, the spatial pose changes at any time, and there is a problem of reference alignment in data fusion, and a unified observation reference needs to be constructed. On the other hand, for a moving platform such as an industrial robot or a mechanical arm used in indoor operation, the position and attitude information of the moving platform needs to be obtained when performing a task, and a relative relationship with a scene or a target is established, and motion planning is performed under a unified reference.

[0003] Reference unification is the basis for most moving platforms to realize information fusion and task planning, and its essence is pose estimation of the moving platform. First, a reference coordinate system is constructed in a scene, and pose information of the moving platform relative to the reference at each time is obtained through a pose estimation method. Traditional positioning technologies such as GPS are difficult to use indoors due to the limitation of wireless signals, and GPS can only output the position information of the moving platform, lacking attitude information. The pose estimation method based on epipolar geometry constraint of a monocular camera lacks depth information, and the solved pose lacks scale information. The pose estimation method based on binocular cameras or depth cameras has poor stability due to the influence of the observation distance on the solution accuracy.

[0004] The visual odometry method combining vision and inertial navigation information combines a camera and an inertial sensor, and can realize stable pose estimation, but has low solution accuracy and is not suitable for scenes with high accuracy requirements.

[0005] In summary, reference unification is a common requirement for most moving platforms, and existing technologies are difficult to construct a stable and accurate unified reference. Therefore, it is necessary to study a moving platform reference self-construction method which is suitable for a wide range of scenes, has strong stability and high solution accuracy. SUMMARY

[0006] In view of the technical problems existing in the prior art, the present application provides a passive reference self-construction method for a moving platform based on scene features.

[0007] To achieve the above purpose, the technical scheme adopted by the present application is as follows:

[0008] On the one hand, the present application provides a passive reference self-construction method for a moving platform based on scene features, comprising:

[0009] A reference coordinate system is established according to the cooperative markers with known world coordinates in the scene, and three-dimensional world coordinates of each cooperative marker in the reference coordinate system are obtained;

[0010] Two images containing the cooperative markers collected by the moving platform are acquired, and two-dimensional image coordinates of each cooperative marker in the two images are obtained;

[0011] Based on the two-dimensional image coordinates of each cooperative marker in the two images and the three-dimensional world coordinates of each cooperative marker in the reference coordinate system, a camera extrinsic parameter in the reference coordinate system that minimizes the re-projection error is solved, and preliminary camera extrinsic parameter estimation corresponding to the two images is completed;

[0012] Feature extraction and matching are performed on the two images, and all corresponding feature point pairs in the two images are obtained;

[0013] Based on the preliminary camera extrinsic parameter estimation results corresponding to the two images, all corresponding feature point pairs in the two images are triangulated, and preliminary estimated three-dimensional coordinates of each corresponding feature point pair in the two images are estimated, and an initial three-dimensional point cloud of the scene is generated;

[0014] Based on the preliminary estimated three-dimensional coordinates of the three-dimensional space feature points and the image coordinates of the three-dimensional space feature points in the image plane of the two images, a re-projection error objective function is constructed, and the camera extrinsic parameter and the three-dimensional coordinates of the three-dimensional space feature points that minimize the re-projection error are solved, to obtain the optimized camera extrinsic parameter estimation result and the scene point cloud;

[0015] Based on the optimized scene point cloud, pose estimation is performed to realize reference alignment of the moving platform.

[0016] Further, as a preferred embodiment, the preliminary camera extrinsic parameter estimation result of the image is solved by the following steps, specifically including:

[0017] For any image I, the image I is any one of the two images, and the spatial three-dimensional coordinates of the i-th cooperative marker are i=1,2,...,n, and the two-dimensional image coordinates of the i-th cooperative marker in the image I are p i (x i ,y i ), according to the imaging relationship, we have:

[0018]

[0019] Where s i is a scale factor, K is a camera intrinsic parameter, which is obtained by camera calibration, and T is a camera extrinsic parameter matrix;

[0020] A first objective function T* is constructed as follows:

[0021]

[0022] By adjusting T, T is continuously reduced. * When T * When the minimum value is reached, the current T is output as the preliminary camera extrinsic parameter estimation result.

[0023] In this invention, for the two images, the corresponding camera optical centers are O and O, respectively. A and O B For the j-th pair of corresponding feature points in two images Where j = 1, 2, ..., m, m is the total number of pairs of feature points with the same name in the two images; the straight line and They will intersect at a point M in the scene. j Point M j That is, pairs of feature points with the same name in two images. The corresponding three-dimensional spatial feature points.

[0024] Further, as a preferred embodiment, triangulation is performed on all pairs of corresponding feature points in the two images to estimate the preliminary estimated three-dimensional coordinates of the three-dimensional spatial feature points corresponding to each pair of corresponding feature points in the two images, including:

[0025] set up For the pair of feature points with the same name Normalized coordinates are used to calculate the depth values ​​corresponding to two feature points in a pair of identical feature points, based on the following formula:

[0026]

[0027] Where R and t are the camera rotation matrix and camera translation matrix in the preliminary camera extrinsic parameter estimation results, respectively, and s A ,s B Let be the depth values ​​corresponding to the two feature points in the pair of feature points with the same name to be solved;

[0028] Based on the depth values ​​corresponding to two feature points in a pair of identical feature points, the preliminary estimated three-dimensional coordinates of the three-dimensional spatial feature points corresponding to the pair of identical feature points are calculated according to the imaging equation.

[0029] Furthermore, as a preferred implementation, the optimized camera extrinsic parameter estimation results and scene point cloud are obtained through the following steps:

[0030] Based on the camera imaging model, each 3D spatial feature point M is... j Projecting the two images onto the image plane, we obtain three-dimensional spatial feature points M respectively. j Image coordinates in the image plane of the two images and

[0031] A re-projection error objective function is constructed as follows:

[0032]

[0033] wherein M j is a spatial coordinate of a jth three-dimensional spatial feature point to be solved in optimization, [R|t] represents a camera extrinsic parameter to be solved in optimization, including a camera rotation matrix and a camera translation matrix R, t; [R A |t A ] and [R B |t B ] are preliminary camera extrinsic parameter estimation results corresponding to the two images respectively; K is a camera intrinsic parameter, K A , K B are camera intrinsic parameters corresponding to the two images respectively, and are obtained through camera calibration; s i is a scale factor; is a preliminary estimation three-dimensional coordinate of the jth three-dimensional spatial feature point obtained through triangulation;

[0034] The re-projection error objective function is solved in optimization to obtain a camera extrinsic parameter and a spatial coordinate of a three-dimensional spatial feature point when a re-projection error ε(M j , K, [R|t]) reaches a minimum value, and then an optimized camera extrinsic parameter estimation result and a scene point cloud are obtained.

[0035] In another aspect, the application provides a passive reference self-construction device based on scene features, comprising:

[0036] A preliminary camera extrinsic parameter estimation module comprises a first module, a second module and a third module, wherein the first module is used to establish a reference coordinate system according to known world coordinates of cooperative markers in a scene, and obtain three-dimensional world coordinates of each cooperative marker in the reference coordinate system; the second module obtains two images containing cooperative markers collected by a moving platform, and obtains two-dimensional image coordinates of each cooperative marker in the two images; and the third module solves a camera extrinsic parameter corresponding to the two images in the reference coordinate system based on the two-dimensional image coordinates of each cooperative marker in the two images and the three-dimensional world coordinates of each cooperative marker in the reference coordinate system, so as to complete preliminary camera extrinsic parameter estimation of the two images;

[0037] The initial three-dimensional point cloud generation module comprises a fourth module and a fifth module, wherein the fourth module is configured to perform feature extraction and matching on the two images to obtain all homonymous feature point pairs in the two images; and the fifth module is configured to perform triangulation on all homonymous feature point pairs in the two images in combination with the preliminary camera extrinsic parameter estimation results corresponding to the two images to estimate preliminary estimation three-dimensional coordinates of three-dimensional space feature points corresponding to each homonymous feature point pair in the two images, and generate an initial three-dimensional point cloud of a scene;

[0038] The optimization module is configured to construct a re-projection error objective function based on the preliminary estimation three-dimensional coordinates of the three-dimensional space feature points and image coordinates of the three-dimensional space feature points in the image planes of the two images, solve camera extrinsic parameters and three-dimensional coordinates of the three-dimensional space feature points that minimize the re-projection error, and obtain an optimized camera extrinsic parameter estimation result and a scene point cloud.

[0039] The moving platform reference alignment module is configured to perform pose estimation based on the optimized scene point cloud to realize moving platform reference alignment.

[0040] In another aspect, the present application provides a computer device comprising a memory and a processor, the memory storing a computer program, and the processor realizing the following steps when executing the computer program:

[0041] A reference coordinate system is established according to the known world coordinates of the cooperative markers in the scene, and three-dimensional world coordinates of each cooperative marker in the reference coordinate system are obtained;

[0042] Two images containing the cooperative markers collected by the moving platform are obtained, and two-dimensional image coordinates corresponding to each cooperative marker in the two images are obtained;

[0043] Based on the two-dimensional image coordinates corresponding to each cooperative marker in the two images and the three-dimensional world coordinates of each cooperative marker in the reference coordinate system, camera extrinsic parameters in the reference coordinate system that minimize the re-projection error are solved, and preliminary camera extrinsic parameter estimation corresponding to the two images is completed.

[0044] Feature extraction and matching are performed on the two images to obtain all homonymous feature point pairs in the two images;

[0045] Triangulation is performed on all homonymous feature point pairs in the two images in combination with the preliminary camera extrinsic parameter estimation results corresponding to the two images to estimate preliminary estimation three-dimensional coordinates of three-dimensional space feature points corresponding to each homonymous feature point pair in the two images, and an initial three-dimensional point cloud of a scene is generated.

[0046] Based on the preliminary estimation three-dimensional coordinates of three-dimensional space feature points and the image coordinates of three-dimensional space feature points in the image plane of the two images, a re-projection error objective function is constructed, the camera extrinsic parameters and the three-dimensional coordinates of three-dimensional space feature points that make the re-projection error minimum are solved, and the optimized camera extrinsic parameter estimation result and scene point cloud are obtained;

[0047] Based on the optimized scene point cloud, pose estimation is carried out to realize the alignment of the dynamic platform reference.

[0048] In another aspect, the application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to realize the following steps:

[0049] A reference coordinate system is established according to the cooperative signs with known world coordinates in the scene, and the three-dimensional world coordinates of each cooperative sign in the reference coordinate system are obtained;

[0050] Two images containing cooperative signs collected by a dynamic platform are obtained, and the two-dimensional image coordinates of each cooperative sign in the two images are obtained;

[0051] Based on the two-dimensional image coordinates of each cooperative sign in the two images and the three-dimensional world coordinates of each cooperative sign in the reference coordinate system, the camera extrinsic parameters in the reference coordinate system that make the re-projection error minimum are solved, and the preliminary camera extrinsic parameter estimation corresponding to the two images is completed;

[0052] Feature extraction and matching are performed on the two images, and all homonymous feature point pairs in the two images are obtained;

[0053] Combined with the preliminary camera extrinsic parameter estimation result corresponding to the two images, the three-dimensional space feature points corresponding to each homonymous feature point pair in the two images are estimated, the preliminary estimation three-dimensional coordinates of the three-dimensional space feature points are obtained, and the initial three-dimensional point cloud of the scene is generated;

[0054] Based on the preliminary estimation three-dimensional coordinates of three-dimensional space feature points and the image coordinates of three-dimensional space feature points in the image plane of the two images, a re-projection error objective function is constructed, the camera extrinsic parameters and the three-dimensional coordinates of three-dimensional space feature points that make the re-projection error minimum are solved, and the optimized camera extrinsic parameter estimation result and scene point cloud are obtained;

[0055] Based on the optimized scene point cloud, pose estimation is carried out to realize the alignment of the dynamic platform reference.

[0056] Compared with the prior art, the technical effects of the application are:

[0057] The technical problem solved by the present application is to unify the reference of the moving platform, and the present application is characterized in that a passive sensor such as a monocular camera is mounted on a moving platform such as a UAV or a mobile robot to collect image data of a scene, a scene point cloud is constructed according to cooperative information, and pose estimation is performed through a PnP method to realize self-construction of the reference of the moving platform. BRIEF DESCRIPTION OF DRAWINGS

[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description only show some embodiments of the present application, and for those skilled in the art, other drawings can be obtained from the structures shown in the drawings without creative labor.

[0059] Figure 1 is a flow chart of an embodiment of the present application;

[0060] Figure 2 is a scene schematic diagram of an embodiment of the present application;

[0061] Figure 3 is a schematic diagram of the principle of triangulation;

[0062] Figure 4 is a scene point cloud optimization principle diagram. DETAILED DESCRIPTION

[0063] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0064] Referring to Figure 1 An embodiment provides a passive reference self-construction method of a moving platform based on scene features, comprising:

[0065] S1. Reference construction based on cooperative signs;

[0066] A reference coordinate system is established according to the cooperative signs with known world coordinates in the scene, and three-dimensional world coordinates of each cooperative sign in the reference coordinate system are obtained.

[0067] A reference coordinate system of the scene is constructed by using the cooperative markers with known world coordinates, wherein the cooperative markers can be artificially laid out.

[0068] S2. Obtain the two-dimensional image coordinates of each cooperative marker in the two images;

[0069] Two images containing cooperative markers and having good intersection conditions are selected from the image data collected by the moving platform, and the two-dimensional image coordinates of each cooperative marker in the two images are obtained through feature extraction and recognition.

[0070] S3. Complete the preliminary camera extrinsic parameter estimation of the two images;

[0071] Based on the two-dimensional image coordinates of each cooperative marker in the two images and the three-dimensional world coordinates of each cooperative marker in the reference coordinate system, the camera extrinsic parameter in the reference coordinate system that minimizes the re-projection error is solved, and the preliminary camera extrinsic parameter estimation of the two images is completed.

[0072] S4. Obtain all corresponding feature point pairs in the two images;

[0073] All corresponding feature point pairs in the two images are obtained through feature extraction and matching of the two images.

[0074] S5. Triangulate all corresponding feature point pairs in the two images to generate an initial three-dimensional point cloud of the scene;

[0075] All corresponding feature point pairs in the two images are triangulated to estimate the preliminary three-dimensional coordinates of the three-dimensional spatial feature points corresponding to each corresponding feature point pair in the two images, and an initial three-dimensional point cloud of the scene is generated, in combination with the preliminary camera extrinsic parameter estimation results of the two images.

[0076] S6. Optimize the camera extrinsic parameter estimation results and the scene point cloud;

[0077] A re-projection error objective function is constructed based on the preliminary three-dimensional coordinates of the three-dimensional spatial feature points and the image coordinates of the three-dimensional spatial feature points in the image planes of the two images, and the three-dimensional coordinates of the three-dimensional spatial feature points and the camera extrinsic parameter that minimize the re-projection error are solved, to obtain the optimized camera extrinsic parameter estimation results and the scene point cloud.

[0078] S7. Reference alignment based on the scene point cloud;

[0079] Pose estimation is performed based on the optimized scene point cloud to realize reference alignment of the moving platform.

[0080] In an embodiment, a drone is taken as an example, i.e., the moving platform to be targeted is a drone, and a scene diagram of passive reference self-construction of the moving platform based on scene features is as followsFigure 1 As shown, after the unmanned aerial vehicle completes image data acquisition, a method as shown in the following figure is used to realize dynamic platform reference alignment. Figure 1 As shown in the figure, the reference coordinate system of the scene is constructed by using the known world coordinates of the cooperative markers, and the image data of the dynamic platform camera and the camera intrinsic parameters are obtained. The camera intrinsic parameters can be obtained by camera calibration. The image data of the dynamic platform is processed, and the initial pose of the dynamic platform is obtained by combining the world coordinates of the known marker points, feature extraction, and PnP pose estimation. The initial point cloud Map of the scene is obtained by triangulation, and BA optimization is performed. In the subsequent solving stage, the relationship between the two-dimensional image coordinates of the features and the three-dimensional space coordinates of the features in the scene point cloud Map is obtained by feature tracking, and the pose is solved by the PnP method, and the scene point cloud Map is updated by triangulation and BA optimization, to obtain the pose information of the dynamic platform in the reference coordinate system, and complete the reference unification.

[0081] In an embodiment, the reference construction based on cooperative markers: a reference coordinate system W is established according to n (n is greater than or equal to 4) cooperative markers with known world coordinates, and the three-dimensional space coordinates of each cooperative marker in the reference coordinate system are obtained

[0082] In the image data set of the dynamic platform, two images I A ,I B are selected, which contain all the cooperative markers and have good intersection conditions.

[0083] The image data of the dynamic platform is processed, and the initial pose of the dynamic platform is obtained by combining the world coordinates of the known marker points, feature extraction, and PnP pose estimation. A ,I B A ,I B and

[0084]

[0085] Specifically, if the cooperative marker is a scene feature with known three-dimensional coordinates, the cooperative marker in the scene is manually selected in the image I A by using the feature tracking method, and the two-dimensional image coordinates of the cooperative marker in the image I A in the image I B are obtained by image matching.

[0086] If the cooperative marker is artificially laid out, the two-dimensional image coordinates of the cooperative marker in the two images are obtained by feature extraction and recognition.

[0087] The three-dimensional space coordinates of each cooperative marker and the two-dimensional image coordinates of each cooperative marker in the image I A , I​​B The corresponding two-dimensional image coordinates and Then, by constructing the PnP problem, the preliminary camera extrinsic parameter estimates corresponding to the two images in the reference coordinate system are obtained.

[0088] For any image I, image I is either of the two images (image I...). A or I B The three-dimensional coordinates of the i-th cooperation marker are... i = 1, 2, ..., n, where the i-th cooperation flag has two-dimensional image coordinates p in image I. i (x i ,y i Based on the imaging relationship, we have:

[0089]

[0090] Where s i Let K be the scaling factor, K be the camera intrinsic parameter obtained through camera calibration, and T be the camera extrinsic parameter matrix. The above equation can be written in matrix form as follows:

[0091] The above equation implicitly involves a transformation from homogeneous to non-homogeneous coordinates. Due to the unknown camera pose and noise at the observation points, the equation contains an error. Summing this error, we construct a least-squares problem and then search for the optimal camera extrinsic parameters to minimize the error. Based on this, we construct the first objective function T*, as follows:

[0092]

[0093] By adjusting T, T is continuously reduced. * When T * When the minimum value is reached, the current T is output as the preliminary camera extrinsic parameter estimation result.

[0094] For image I A ,I B Simultaneously performing the above PnP pose estimation yields two images I. A ,I B Preliminary camera extrinsic parameter estimation in their respective reference coordinate systems, i.e., the two images I A ,I B The corresponding camera rotation matrix and camera translation matrix R A ,t A and R B ,t B .

[0095] This invention obtains all pairs of identical feature points in the two images by performing feature extraction and matching. It is understood that those skilled in the art can obtain all pairs of identical feature points in two images using any existing feature extraction and matching method, including but not limited to ORB, SIFT, SURF, and other feature matching methods.

[0096] The triangulation described in this invention refers to observing the same spatial point from different locations and inferring the distance of the spatial point from the observed locations. Figure 3 This is a schematic diagram of triangulation, for the initial image I. A ,I B , with I A For reference, I B The transformation matrix is ​​T, and the optical centers of the two cameras are O and O. A and O B For the j-th pair of corresponding feature points in two images In an ideal situation, a straight line and They will intersect at a point M in the scene. j Point M j That is, pairs of feature points with the same name in two images. The corresponding three-dimensional spatial feature points. However, due to the influence of noise, these two lines often cannot intersect directly, which can be solved using the least squares method.

[0097] In one embodiment of the present invention, triangulation is performed on all pairs of corresponding feature points in the two images to estimate the preliminary estimated three-dimensional coordinates of the three-dimensional spatial feature points corresponding to each pair of corresponding feature points in the two images, including:

[0098] set up For the pair of feature points with the same name Normalized coordinates are used to calculate the depth values ​​corresponding to two feature points in a pair of identical feature points, based on the following formula:

[0099]

[0100] Where R and t are the camera rotation matrix and camera translation matrix in the preliminary camera extrinsic parameter estimation results, respectively, and s A ,s B Let be the depth values ​​corresponding to the two feature points in the pair of feature points with the same name to be solved;

[0101] Based on the depth values ​​corresponding to two feature points in a pair of identical feature points, the preliminary estimated 3D coordinates of the corresponding 3D spatial feature points are calculated according to the imaging equation. Geometrically, this can be achieved through ray... The above uses an optimization search to find pairs of feature points with the same name. The corresponding three-dimensional spatial feature points, whose projected positions are closest to The three-dimensional spatial feature points can be considered as pairs of feature points with the same name. Preliminary estimation of the three-dimensional coordinates of the corresponding three-dimensional spatial feature points

[0102] By performing the above triangulation on all pairs of corresponding feature points obtained from image feature matching, a preliminary estimate of the three-dimensional coordinates of the corresponding three-dimensional feature points in the reference coordinate system can be obtained, and the initial point cloud of the scene can be constructed accordingly.

[0103] In one embodiment, the optimized camera extrinsic parameter estimation results and scene point cloud are obtained through the following steps:

[0104] Based on the camera imaging model, each 3D spatial feature point M is... j Projecting the two images onto the image plane, we obtain three-dimensional spatial feature points M respectively. j Image coordinates in the image plane of the two images and like Figure 4 As shown, the three-dimensional spatial feature point M in the reference coordinate system j In image I A ,I B The image coordinates on are respectively and The camera intrinsic parameter K is obtained through camera calibration. A and K B .

[0105] Due to errors in pose estimation and triangulation, there is a deviation between the projected image coordinates of spatial feature points and the actual image coordinates of the features; this deviation is known as reprojection error. BA optimization, also called bundle adjustment, uses optimization algorithms to optimize camera extrinsic parameters and the spatial coordinates of 3D feature points to minimize the reprojection error, thereby obtaining accurate camera extrinsic parameters and 3D coordinates of the feature points.

[0106] Construct the following objective function for reprojection error:

[0107]

[0108] Where M j Let [R|t] represent the spatial coordinates of the j-th 3D feature point to be optimized, and let [R|t] represent the camera extrinsic parameters to be optimized, including the camera rotation matrix and the camera translation matrix R,t; [R A |t A ] and [R B |t B [ ] represents the preliminary camera extrinsic parameter estimation results corresponding to the two images; K represents the camera intrinsic parameter, K A, K B are the camera intrinsic parameters of the two images respectively, which are obtained by camera calibration; i is a scale factor; is the preliminary estimation of the three-dimensional coordinates of the jth three-dimensional feature point obtained by triangulation;

[0109] The re-projection error objective function is optimized and solved to obtain the camera extrinsic parameters and the spatial coordinates of the three-dimensional feature points when the re-projection error ε(M j , K, [R|t]) reaches a minimum value, and then the optimized camera extrinsic parameter estimation result and the scene point cloud are obtained.

[0110] Due to the continuity of the motion of the moving platform, it can be considered that the two-dimensional coordinates of the same spatial feature point in adjacent two images will not change dramatically. Feature tracking is performed on adjacent two images by using an optical flow method to obtain the two-dimensional image coordinates of a part of the spatial feature points in the scene point cloud in the current frame image. By using the correspondence between the three-dimensional coordinates of the feature points in the scene point cloud and the two-dimensional coordinates of the image, the camera pose of the current frame image in the reference coordinate system is obtained by using a PnP pose estimation method. At the same time, the scene point cloud and the camera pose are optimized by BA to further optimize the accuracy of the point cloud and the pose.

[0111] At this point, the pose parameters R and t of the moving platform at any time in the reference coordinate system W have been obtained, and the reference construction and reference alignment are completed.

[0112] The application carries a monocular camera on the moving platform to collect image data of the environment. In the initialization stage, a reference coordinate system is constructed according to the scene cooperation information or artificial markers. The scene point cloud is obtained by pose estimation, triangulation and BA optimization. In the subsequent stage, the scene point cloud is used for pose estimation to complete the reference alignment.

[0113] The method proposed in the application is not limited by wireless signals. By carrying a monocular camera on the moving platform, the environmental image data is passively obtained. The cooperation information and natural features in the scene are combined to realize the self-construction of the reference of the moving platform.

[0114] A passive reference self-construction device for a moving platform based on scene features, comprising:

[0115] The preliminary camera extrinsic parameter estimation module includes a first module, a second module and a third module, wherein the first module is configured to establish a reference coordinate system according to the cooperative markers with known world coordinates in the scene, and obtain the three-dimensional world coordinates of each cooperative marker in the reference coordinate system; the second module obtains two images containing the cooperative markers collected by the moving platform, and obtains the two-dimensional image coordinates of each cooperative marker in the two images; and the third module solves the camera extrinsic parameter in the reference coordinate system corresponding to the two images which minimizes the re-projection error based on the two-dimensional image coordinates of each cooperative marker in the two images and the three-dimensional world coordinates of each cooperative marker in the reference coordinate system, and completes the preliminary camera extrinsic parameter estimation corresponding to the two images.

[0116] The initial three-dimensional point cloud generation module includes a fourth module and a fifth module, wherein the fourth module is configured to perform feature extraction and matching on the two images to obtain all homonymous feature point pairs in the two images; and the fifth module is configured to combine the preliminary camera extrinsic parameter estimation results corresponding to the two images, perform triangulation on all homonymous feature point pairs in the two images, estimate the preliminary three-dimensional coordinates of the three-dimensional space feature points corresponding to each homonymous feature point pair in the two images, and generate the initial three-dimensional point cloud of the scene.

[0117] The optimization module is configured to construct a re-projection error objective function based on the preliminary three-dimensional coordinates of the three-dimensional space feature points and the image coordinates of the three-dimensional space feature points in the image planes of the two images, solve the camera extrinsic parameter and the three-dimensional coordinates of the three-dimensional space feature points that minimize the re-projection error, and obtain the optimized camera extrinsic parameter estimation result and the scene point cloud.

[0118] The moving platform reference alignment module is configured to perform pose estimation based on the optimized scene point cloud to realize the alignment of the moving platform reference.

[0119] The implementation methods of the above modules and the construction of the model can adopt the methods described in any of the preceding embodiments, and will not be described here.

[0120] In another aspect, the present application provides a computer device comprising a memory and a processor, the memory storing a computer program, the processor implementing the steps of the method for self-construction of a passive reference based on scene features of a moving platform according to any one of the above embodiments when executing the computer program. The computer device can be a server. The computer device comprises a processor, a memory, a network interface and a database connected by a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device comprises a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The database of the computer device is configured to store sample data. The network interface of the computer device is configured to communicate with an external terminal through a network connection.

[0121] In another aspect, the present application provides a computer readable storage medium storing a computer program, the computer program being executed by a processor to implement the steps of the method for self-construction of a passive reference based on scene features of a moving platform according to any one of the above embodiments.

[0122] It is understood by those skilled in the art that all or part of the processes of the above embodiments can be completed by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer readable storage medium and executed to include the processes of the above embodiments. Any reference to memory, storage, database or other medium in the embodiments provided by the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM) and memory bus dynamic RAM (RDRAM), etc.

[0123] The details of the present application are as described above.

[0124] Any technical features in the above-described embodiments can be combined in any manner, and for the sake of brevity, not all possible combinations are described, however, any combination of the technical features is considered to be within the scope of the present specification.

[0125] The above-described embodiments are merely illustrative for the present application and are not used to limit the present application. It should be pointed out that, for those skilled in the art, some modifications and improvements can be made without departing from the concept of the present application, and these should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

[0126] The above-described embodiments are merely illustrative for the present application and are not used to limit the present application. It should be pointed out that, for those skilled in the art, some modifications and improvements can be made without departing from the concept of the present application, and these should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims. The above-described embodiments are merely illustrative for the present application and are not used to limit the present application. It should be pointed out that, for those skilled in the art, some modifications and improvements can be made without departing from the concept of the present application, and these should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A method for constructing a passive reference frame from a moving platform based on scene features, characterized in that, The method comprises the following steps: establishing a reference coordinate system according to cooperative markers with known world coordinates in a scene, and obtaining three-dimensional world coordinates of each cooperative marker in the reference coordinate system; obtaining two images containing cooperative markers collected by a moving platform, and obtaining two-dimensional image coordinates of each cooperative marker in the two images; based on the two-dimensional image coordinates of each cooperative marker in the two images and the three-dimensional world coordinates of each cooperative marker in the reference coordinate system, solving camera extrinsic parameters in the reference coordinate system corresponding to the two images that minimize the re-projection error, and completing preliminary camera extrinsic parameter estimation corresponding to the two images; performing feature extraction and matching on the two images to obtain all homonymous feature point pairs in the two images; combining the preliminary camera extrinsic parameter estimation results corresponding to the two images, performing triangulation on all homonymous feature point pairs in the two images, estimating preliminary three-dimensional coordinates of three-dimensional space feature points corresponding to each homonymous feature point pair in the two images, and generating an initial three-dimensional point cloud of the scene; based on the preliminary three-dimensional coordinates of the three-dimensional space feature points and the image coordinates of the three-dimensional space feature points in the image planes of the two images, constructing a re-projection error objective function, solving camera extrinsic parameters and three-dimensional coordinates of the three-dimensional space feature points that minimize the re-projection error, and obtaining optimized camera extrinsic parameter estimation results and scene point cloud; based on the optimized scene point cloud, performing pose estimation to realize reference alignment of the moving platform.

2. The method of claim 1, wherein, Solving the preliminary camera extrinsic parameter estimation results of the images comprises: For any image I, image I being either of the two images, the i-th cooperating marker has its spatial three-dimensional coordinates The i-th cooperating marker has its corresponding two-dimensional image coordinates in image I, p i (x i ,y i ), according to the imaging relationship, there is: where s i is a scale factor, K is the camera intrinsic parameter, which is obtained by camera calibration, and T is the camera extrinsic matrix. constructing a first objective function T* as follows: By adjusting T, T is constantly narrowed * When T * reaches the minimum value, the output current T is the preliminary camera extrinsic parameter estimation result.

3. The method according to claim 1 or 2, wherein, For the two images, the camera optical centers corresponding to the two images are O A and O B For the jth pair of corresponding feature points in the two images where j = 1, 2, …, m, m is the total number of pairs of corresponding feature points in the two images; the straight lines and intersect at a point M j in the scene j , i.e., the three-dimensional feature points corresponding to the pairs of corresponding feature points in the two images .

4. The method of claim 3, wherein, performing triangulation on all homonymous feature point pairs in the two images to estimate preliminary three-dimensional coordinates of three-dimensional space feature points corresponding to each homonymous feature point pair in the two images, comprising: Let be the normalized coordinates of the corresponding feature point pair , the depth values corresponding to the two feature points in the corresponding feature point pair are solved based on the following formula: where R and t are camera rotation matrix and camera translation matrix in the preliminary camera extrinsic parameter estimation result, respectively, s A B are the depth values corresponding to the two feature points in the same-named feature point pair to be solved.​ based on the depth values corresponding to the two feature points in the homonymous feature point pair, calculating the preliminary three-dimensional coordinates of the three-dimensional space feature points corresponding to the homonymous feature point pair according to the imaging equation.

5. The method of claim 3, wherein, Obtaining the optimized camera extrinsic parameter estimation results and the scene point cloud comprises: According to the camera imaging model, each three-dimensional spatial feature point M j is projected onto the two image image planes to obtain three-dimensional spatial feature point M j image coordinates in the two image image planes and constructing a re-projection error objective function as follows: wherein M j is the spatial coordinate of the jth three-dimensional spatial feature point to be solved for optimization, and [R|t] represents the camera extrinsic parameter to be solved for optimization, including a camera rotation matrix and a camera translation matrix R, t; [R A |t A ] and [R B |t B ] are respectively the preliminary camera extrinsic parameter estimation results corresponding to the two images; K is a camera intrinsic parameter, K A and K B are respectively the camera intrinsic parameters corresponding to the two images, obtained through camera calibration; s i is a scale factor; is a preliminary estimated three-dimensional coordinate of the jth three-dimensional spatial feature point obtained through triangulation. Optimizing the re-projection error objective function to obtain camera extrinsic parameters and spatial coordinates of three-dimensional feature points when the re-projection error ε(M j , K, [R | t]) reaches a minimum value, and further obtain an optimized camera extrinsic parameter estimation result and a scene point cloud.

6. A scene feature based passive reference self-constructing device for a moving platform, characterized in that, comprising: a preliminary camera extrinsic parameter estimation module comprising a first module, a second module and a third module, wherein the first module is used to establish a reference coordinate system according to cooperative markers with known world coordinates in a scene, and obtain three-dimensional world coordinates of each cooperative marker in the reference coordinate system; the second module obtains two images containing cooperative markers collected by a moving platform, and obtains two-dimensional image coordinates of each cooperative marker in the two images; and the third module is based on the two-dimensional image coordinates of each cooperative marker in the two images and the three-dimensional world coordinates of each cooperative marker in the reference coordinate system, solves camera extrinsic parameters in the reference coordinate system corresponding to the two images that minimize the re-projection error, and completes preliminary camera extrinsic parameter estimation corresponding to the two images; The initial three-dimensional point cloud generation module comprises a fourth module and a fifth module. The fourth module is configured to perform feature extraction and matching on the two images to obtain all homonymous feature point pairs in the two images. The fifth module is configured to perform triangulation on all homonymous feature point pairs in the two images in combination with the preliminary camera extrinsic parameter estimation results corresponding to the two images to estimate preliminary estimated three-dimensional coordinates of three-dimensional space feature points corresponding to each homonymous feature point pair in the two images, and generate an initial three-dimensional point cloud of a scene. The optimization module is configured to construct a re-projection error objective function based on the preliminary estimated three-dimensional coordinates of the three-dimensional space feature points and image coordinates of the three-dimensional space feature points in the image planes of the two images, solve camera extrinsic parameters and three-dimensional coordinates of the three-dimensional space feature points that minimize the re-projection error, and obtain an optimized camera extrinsic parameter estimation result and a scene point cloud. The moving platform reference alignment module is configured to perform pose estimation based on the optimized scene point cloud to achieve moving platform reference alignment.

7. The device according to claim 6, wherein, In the third module, the preliminary camera extrinsic parameter estimation result of the image is solved by the following method, comprising: For any image I, image I being either of the two images, the i-th cooperating marker has its spatial three-dimensional coordinates The i-th cooperating marker has its corresponding two-dimensional image coordinates in image I, p i (x i ,y i ), according to the imaging relationship, there is: where s i is a scale factor, K is the camera intrinsic parameter, which is obtained by camera calibration, and T is the camera extrinsic matrix. A first objective function T* is constructed as follows: By adjusting T, T is constantly narrowed * When T * reaches the minimum value, the output of the current T is the preliminary camera extrinsic parameter estimation result.

8. The device according to claim 6, wherein, In the optimization module, the optimized camera extrinsic parameter estimation result and the scene point cloud are obtained by the following steps, comprising: According to the camera imaging model, each three-dimensional spatial feature point M j is projected onto the two image image planes to obtain three-dimensional spatial feature point M j image coordinates in the two image image planes and A re-projection error objective function is constructed as follows: wherein M j is the spatial coordinate of the jth three-dimensional spatial feature point to be solved for optimization, and [R|t] represents the camera extrinsic parameter to be solved for optimization, including a camera rotation matrix and a camera translation matrix R, t; [R A |t A ] and [R B |t B ] are respectively the preliminary camera extrinsic parameter estimation results corresponding to the two images; K is a camera intrinsic parameter, K A and K B are respectively the camera intrinsic parameters corresponding to the two images, obtained through camera calibration; s i is a scale factor; is a preliminary estimated three-dimensional coordinate of the jth three-dimensional spatial feature point obtained through triangulation. Optimizing the re-projection error objective function to obtain camera extrinsic parameters and spatial coordinates of three-dimensional feature points when the re-projection error ε(M j , K, [R | t]) reaches a minimum value, and further obtain an optimized camera extrinsic parameter estimation result and a scene point cloud.

9. Computer device, characterized in that The computer program is executed by the processor to implement the moving platform passive reference self-construction method based on scene features as claimed in claim 1 or 2 or 4 or 5.

10. A computer readable storage medium having stored thereon a computer program, characterized in that: The computer program is executed by the processor to implement the moving platform passive reference self-construction method based on scene features as claimed in claim 1 or 2 or 4 or 5.

Citation Information

Patent Citations

  • Method and device for calibrating relative parameters of collector, apparatus and medium

    CN109242913A

  • Three-dimensional information restoration device, three-dimensional information restoration system, and three-dimensional information restoration method

    WO2016103621A1