Image matching method

By using the matching method of image feature set and three-dimensional coordinates in the image RTK measurement technology, the problems of low efficiency and low accuracy of multi-frame common viewpoint matching are solved, and high-precision automatic matching is achieved, which improves the usability and efficiency of measurement.

WO2025111822A1PCT designated stage expired Publication Date: 2025-06-05FEYMAN BEIJING TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2023/134862
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-28
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

In the existing image RTK measurement technology, the multi-frame common viewpoint matching efficiency is low and the accuracy is not high, especially in scenarios where there are no obvious corner points, it is difficult to achieve automatic matching.

Method used

By receiving photos taken by the surveying and mapping camera, selecting basic frames and reference frame groups, using image feature sets and three-dimensional coordinate back projection, gradually matching pixels around the point to be measured, and combining image feature sets to achieve multi-frame common-view matching.

Benefits of technology

High-precision multi-frame common viewpoint matching is achieved, the matching accuracy can reach the sub-pixel level, and it can automatically match user-specified points in a feature-free environment, improving the usability and efficiency of measurement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023134862_05062025_PF_FP_ABST
    Figure CN2023134862_05062025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of image matching, and aims to provide an image matching method. The present invention comprises: receiving captured photographs of an object to be measured captured at different positions by a surveying and mapping camera; using any of the captured photographs as a basic frame, and selecting on the basic frame a point A1 to be measured; then selecting a reference frame group; using a specified two-dimensional square area located around said point A1 in the basic frame as a processing area R1, and on the basis of the reference frame group, obtaining an image feature set M1 corresponding to the processing area R1 and an initial value of the three-dimensional coordinate X of said point A1 in the geodetic coordinate system; back-projecting the initial value of the three-dimensional coordinate X to any adjacent photograph in the reference frame group, to obtain a point A2 to be measured corresponding to said point A1; and finally obtaining a co-visibility matching result of all pixels around said point A1. Thus, co-visibility matching of multiple frames of photographs is realized. The present invention has high matching precision and high availability.
Need to check novelty before this filing date? Find Prior Art

Description

An image matching method Technical Field

[0001] The present invention belongs to the technical field of image matching, and in particular relates to an image matching method. Background Art

[0002] An imaging RTK (Real-time kinematic) camera is a non-contact measurement device that combines an RTK positioning system, an IMU (inertial measurement unit) navigation system (consisting of three single-axis accelerometers and three single-axis gyroscopes), and a monocular vision system. After capturing an image, an imaging RTK camera can select a target point in the image and perform three-dimensional coordinate measurement. The measurement results are referenced to the Earth's coordinate system and provide absolute position information. Imaging RTK measurement technology can measure target points ranging from a few meters to tens of meters, while avoiding interference from the surrounding environment during RTK positioning. This significantly improves measurement efficiency and usability.

[0003] A core technical point in image RTK measurement is how to achieve multi-frame common viewpoint matching at user-selected points. In existing technologies, the following two solutions are usually used:

[0004] a. Multi-frame common viewpoint matching is achieved by manually selecting matching points. However, this method is mainly performed manually, which is cumbersome and inefficient. The matching accuracy can only be achieved to 1 pixel, which is poor.

[0005] b. Automatically complete multi-frame common viewpoint matching through software. This method is relatively simple to operate and can achieve sub-pixel matching accuracy. However, conventional automatic matching cannot usually perform multi-frame matching based on user input points. Instead, it can only complete multi-frame matching by searching for corner points or feature points near the user input points. Although this method achieves automatic multi-frame common viewpoint matching to a certain extent, the matching rate and accuracy are low, especially in scenes without obvious corner points, where automatic matching is basically impossible.

[0006] Summary of the Invention

[0007] The present invention aims to solve the above technical problems at least to a certain extent, and provides an image matching method.

[0008] In order to achieve the above object, the present invention adopts the following technical solutions:

[0009] In a first aspect, the present invention provides an image matching method, comprising:

[0010] Receiving photos of the object to be measured taken by a surveying camera at different positions;

[0011] Any one of the photographs is used as a base frame, and a point A1 to be measured is selected on the base frame;

[0012] Selecting a group of adjacent photos that have a co-viewing relationship with the base frame from the captured photos, and using the group of adjacent photos as a reference frame group;

[0013] A designated two-dimensional square area around the test point A1 in the base frame is used as a processing area R1, and an image feature set M1 corresponding to the processing area R1 and an initial value of the three-dimensional coordinate X of the test point A1 in the geodetic coordinate system are obtained based on the reference frame group; wherein the image feature set M1 includes the final effective depth and normal vector information of all pixels in the processing area R1;

[0014] Back-projecting the initial value of the three-dimensional coordinate X onto any adjacent photo in the reference frame group to obtain the test point A2 corresponding to the test point A1, setting this adjacent photo as a new base frame, and then reselecting a set of adjacent photos in the reference frame group that have a co-viewing relationship with the new base frame as a new reference frame group, until obtaining image feature sets M1, M2, ..., and Mn corresponding to all processing areas in the reference frame group; where n is the number of adjacent photos in the initial reference frame group;

[0015] The image feature set M1, the image feature set M2, ... and the image feature set Mn are combined to obtain the common view matching result of all pixels around the point A1 to be measured.

[0016] The beneficial effects of the present invention are as follows: high matching accuracy, which can reach the sub-pixel level; high usability, which can realize automatic matching of multi-frame common viewpoints of user-selected points in an environment without feature points. Specifically, during the implementation of the present invention, by receiving photos of the object to be measured taken by a surveying camera at different positions, any one of the photos is used as a basic frame, and a point to be measured A1 is selected on the basic frame, and then a reference frame is selected, and the specified two-dimensional square area around the point to be measured A1 in the basic frame is used as the processing area R1, and the image feature set M1 corresponding to the processing area R1 and the initial value of the three-dimensional coordinate X of the point to be measured A1 in the geodetic coordinate system are obtained according to the reference frame group, and then the initial value of the three-dimensional coordinate X is reversely projected to any adjacent photo in the reference frame group to obtain the point to be measured A2 corresponding to the point to be measured A1. , and set this adjacent photo as a new base frame, and then reselect a group of adjacent photos in the reference frame group that have a common view relationship with the new base frame as a new reference frame group, until the image feature set M1, image feature set M2, ... and image feature set Mn corresponding to all processing areas in the reference frame group are obtained. Finally, by merging the image feature set M1, image feature set M2, ... and image feature set Mn, the common view matching result of all pixels around the test point A1 is obtained, thereby realizing the common view matching of multiple frames of photos. In this process, the matching relationship can reach the sub-pixel level, and even in areas without corner points or feature points, automatic matching of user-specified points can be completed, which has higher usability.

[0017] In one possible design, obtaining the image feature set M1 corresponding to the processing area R1 and the initial value of the three-dimensional coordinate X of the measured point A1 in the geodetic coordinate system according to the reference frame group includes:

[0018] Select any pixel P in the processing area R1, and determine that the pixel P and a specified number of adjacent pixels adjacent to the pixel P are in the same physical plane of the object to be measured, and the specified number of adjacent pixels can be referred to as an adjacent sub-block of the pixel P;

[0019] Obtaining final effective depth and normal vector information of all pixels in the processing area R1 according to the reference frame group;

[0020] Reselecting a pixel in the processing area R1 until final effective depth and normal vector information of all pixels in the processing area R1 is obtained, and storing the final effective depth and normal vector information of all pixels in the processing area R1 into the image feature set M1 to obtain the image feature set M1 corresponding to the processing area R1;

[0021] The average value of all final effective depths and normal vector information in the image feature set M1 is taken to obtain the initial value of the three-dimensional coordinate X of the point A1 to be measured in the geodetic coordinate system.

[0022] In one possible design, obtaining final effective depth and normal vector information of all pixels in the processing area R1 according to the reference frame group includes:

[0023] Randomly initialize all pixels in the processing area R1 to obtain the initial depth and normal vector of the physical plane where each pixel in the processing area R1 is located;

[0024] Obtaining the similarity between all pixels in the processing region R1 and the reference frame group, and when the similarity is greater than a preset similarity threshold, using the initial depth and normal vector of the physical plane where the corresponding pixel is located as the effective depth and normal vector of the pixel;

[0025] Depth and normal vector propagation optimization and depth and normal vector random optimization are performed on all pixels in the processing area R1 in sequence, so as to obtain final effective depth and normal vector information of all pixels based on the effective depth and normal vector of all pixels in the processing area R1.

[0026] In one possible design, obtaining the similarity between a pixel P in the processing region R1 and the reference frame group, and when the similarity is greater than a preset similarity threshold, using the initial depth and normal vector of the physical plane where the pixel P is located as the effective depth and normal vector of the pixel, includes:

[0027] A homographic transformation is performed on the coordinates of the adjacent sub-image blocks of the pixel P in the processing area R1 to obtain the transformed pixel coordinates corresponding to the adjacent sub-image blocks in all adjacent photos in the reference frame group, and a reference frame pixel information group of the area corresponding to the transformed pixel coordinates in all adjacent photos in the reference frame group is obtained, so as to calculate the similarity between each reference frame pixel information group and the current adjacent sub-image block, and obtain similarity s1, similarity s2, ... and similarity sn, where n is the number of adjacent photos in the initial reference frame group; the average of the largest three similarities among the similarities s1, similarity s2, ... and similarity sn is used as the similarity S between the pixel P and the reference frame group, and when the similarity S is greater than a preset similarity threshold, the initial depth and normal vector of the physical plane where the pixel P is located are used as the effective depth and normal vector E1 of the pixel P.

[0028] In one possible design, while depth and normal vector propagation optimization and depth and normal vector random optimization are sequentially performed on all pixels in the processing area R1, depth and normal vector propagation optimization is performed on a pixel P in the processing area R1, including: if the pixel P has a valid depth and normal vector E1, assigning the valid depth and normal vector E1 to a neighboring pixel of the pixel P, and using the valid depth and normal vector E1 to calculate the similarity between the current neighboring pixel and the reference frame group; if the similarity is greater than the current similarity of the current neighboring pixel, accepting the assignment and updating the similarity of the current neighboring pixel, and then proceeding to the next step; otherwise, rejecting the assignment and directly performing depth and normal vector random optimization on the neighboring pixels of the pixel P;

[0029] The depth and normal vector of the pixel P in the processing area R1 are randomly optimized, including: setting disturbances on the estimated values ​​of the current effective depth and normal vector of the pixel P, and obtaining the similarity between the pixel P and the reference frame group after setting the disturbance, and then retaining the estimated value with higher similarity as the final effective depth and normal vector information and assigning it to the pixel P.

[0030] In a possible design, the following formula is used to perform homography transformation on any coordinate in the adjacent sub-image block of the pixel P in the processing area R1:

[0031] ;

[0032] Where, is the current coordinate in the adjacent sub-image block; is the transformed pixel coordinates corresponding to the current coordinates in the adjacent sub-image block of any adjacent photo in the reference frame group; Obtained by the following formula:

[0033] ;

[0034] Where, and is the internal reference information of the surveying and mapping camera; and is the rotation matrix of the mapping camera; and is the translation vector of the mapping camera; is the normal vector of the preset physical plane; The depth vector of the preset physical plane in the camera coordinate system.

[0035] In one possible design, a weighted NCC matching algorithm is used to calculate the similarity between each reference frame pixel information group and the current adjacent sub-image block. Accordingly, the similarity between any reference frame pixel information group and the current adjacent sub-image block is:

[0036] ;

[0037] [Corrected 29.12.2023 according to Rule 91] where w(i,j) is the preset NCC weight; I(i,j) is the grayscale value of the pixel in the current adjacent sub-image block, p(i,j) is the coordinate of the pixel in the current adjacent sub-image block, I c is the gray value of the pixel at the center of the current adjacent sub-image block, p c is the coordinate of the pixel at the center of the current adjacent sub-image block, σ I is the preset grayscale weight parameter, σ p is the preset coordinate weight parameter.

[0038] In a possible design, the initial value of the three-dimensional coordinate X is reversely projected onto any adjacent photo in the reference frame group using the following formula to obtain the measured point A2 corresponding to the measured point A1:

[0039] ;

[0040] Where, is the coordinate of the point A2 to be measured, is the internal reference information of the surveying and mapping camera, is the initial value of the three-dimensional coordinate X, is the three-dimensional coordinate of the surveying camera, is the transpose of the rotation matrix of the camera coordinate system of the surveying camera relative to the geodetic reference system, and Z is the depth value of the three-dimensional coordinate X in the camera coordinate system of the surveying camera.

[0041] In one possible design, the initial value of the three-dimensional coordinate X is back-projected onto any adjacent photo in the reference frame group to obtain the measured point A2 corresponding to the measured point A1, and this adjacent photo is set as a new base frame. Then, a group of adjacent photos in the reference frame group that have a co-viewing relationship with the new base frame are reselected as a new reference frame group, and multi-core CPUs are used for parallel execution.

[0042] In one possible design, the image feature set M1, the image feature set M2, ..., and the image feature set Mn are combined to obtain the common view matching result of all pixels around the test point A1, including:

[0043] Randomly select an image feature set M from the image feature set M1, the image feature set M2, ..., and the image feature set Mn, and obtain any final effective depth and normal vector information E3 in the image feature set M, where the final effective depth and normal vector information E3 corresponds to the pixel P1;

[0044] Projecting the final effective depth and normal vector information E3 to any reference frame in the reference frame group to obtain the corresponding pixel P2 in the reference frame;

[0045] Using the coordinates of pixel P2 as an index, the final effective depth and normal vector information E4 corresponding to pixel P2 is searched in the image feature set Mn, and the depth difference and normal vector angle between the final effective depth and normal vector information E3 and the final effective depth and normal vector information E4 are compared to see whether they are less than a preset threshold. If so, it is determined that pixel P1 and pixel P2 are common view matching points, and then the next step is entered;

[0046] Deleting the final effective depth and normal vector information E3 and the final effective depth and normal vector information E4 from the image feature set M and the image feature set Mn respectively;

[0047] Any one of the image feature sets M1, M2, ... and Mn is selected again until a common view matching result of all pixels around the point A1 to be measured is obtained. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] FIG1 is a flow chart of an image matching method according to an embodiment;

[0049] FIG2 is a module block diagram of an electronic device according to an embodiment.

[0050] DETAILED DESCRIPTION

[0051] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the present invention will be briefly introduced below in conjunction with the drawings and the description of the embodiments or the prior art. Obviously, the following description of the structure of the drawings is only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work. It should be noted that the description of these embodiments is used to help understand the present invention, but does not constitute a limitation of the present invention.

[0052] Example 1:

[0053] This embodiment discloses an image matching method, which can be executed by, but is not limited to, a computer device or a virtual machine with certain computing resources, such as a personal computer, a smart phone, a personal digital assistant, or a wearable device, or by a virtual machine.

[0054] As shown in FIG1 , an image matching method may include, but is not limited to, the following steps:

[0055] S1. Receive photos of the object to be measured taken at different locations by a surveying camera, and obtain the absolute position of the surveying camera in the geodetic coordinate system based on the photos for subsequent point selection and measurement. In this embodiment, the surveying camera is an imaging RTK camera. The object to be measured may be a designated building, etc., without limitation.

[0056] S2. Take any one of the photographs as a base frame, and select a point A1 to be measured on the base frame; in this embodiment, the coordinates of the point A1 to be measured are recorded as A1.

[0057] S3. Select a group of adjacent photos that have a common viewing relationship with the basic frame from the captured photos, and use the group of adjacent photos as a reference frame group. It should be noted that the number of adjacent photos depends on the common viewing relationship of the captured photos. For example, if all five captured photos can see the object to be measured, then there is a common viewing relationship between these five photos, and they are reference frames of each other. These five photos constitute a reference frame group.

[0058] S4. A designated two-dimensional square region surrounding the test point A1 in the base frame is defined as a processing region R1. Based on the reference frame group, an image feature set M1 corresponding to the processing region R1 and the initial value of the three-dimensional coordinate X of the test point A1 in the geodetic coordinate system are obtained. The image feature set M1 includes the final effective depth and normal vector information for all pixels within the processing region R1. Specifically, each pixel within the processing region R1 is processed one by one, and the coordinates of the pixels within the processing region R1 are defined as P.

[0059] Specifically, in this embodiment, obtaining the image feature set M1 corresponding to the processing area R1 and the initial value of the three-dimensional coordinate X of the measured point A1 in the geodetic coordinate system according to the reference frame group includes:

[0060] S401. Select any pixel P in the processing area R1. In this embodiment, the coordinates of the pixel P are recorded as P; and it is determined that the pixel P and a specified number of adjacent pixels adjacent to the pixel P are in the same physical plane of the object to be measured; it should be noted that, if the specified number is 5, it means that the 5 adjacent pixels adjacent to the pixel P and the pixel P are considered to be in the same physical plane of the object to be measured. In this embodiment, there is no limit on the size of the specified number.

[0061] S402. Obtain final effective depth and normal vector information of all pixels in the processing area R1 according to the reference frame group.

[0062] Specifically, in this embodiment, obtaining the final effective depth and normal vector information of all pixels in the processing area R1 according to the reference frame group includes:

[0063] S4021. Randomly initialize all pixels in the processing area R1 to obtain the initial depth and normal vector of the physical plane where each pixel in the processing area R1 is located.

[0064] S4022. Obtain the similarity between all pixels in the processing area R1 and the reference frame group, and when the similarity is greater than a preset similarity threshold, use the initial depth and normal vector of the physical plane where the corresponding pixel is located as the effective depth and normal vector of the pixel;

[0065] In step S4022, the similarity between the pixel P in the processing area R1 and the reference frame group is obtained, and when the similarity is greater than a preset similarity threshold, the initial depth and normal vector of the physical plane where the pixel P is located are used as the effective depth and normal vector of the pixel, including:

[0066] A homographic transformation is performed on the coordinates of the adjacent sub-image blocks of the pixel P in the processing region R1 to obtain the transformed pixel coordinates corresponding to the adjacent sub-image blocks in all adjacent photos in the reference frame group, and reference frame pixel information groups corresponding to the transformed pixel coordinates in all adjacent photos in the reference frame group are obtained to calculate the similarity between each reference frame pixel information group and the current adjacent sub-image block, thereby obtaining similarities s1, s2, ..., and sn, where n is the number of adjacent photos in the initial reference frame group; the average of the largest three similarities among s1, s2, ..., and sn is used as the similarity S between the pixel P and the reference frame group, and when the similarity S is greater than a preset similarity threshold, the initial depth and normal vector of the physical plane where the pixel P is located are used as the effective depth and normal vector E1 of the pixel P. It should be noted that the similarity S reflects the credibility of the initial depth and normal vector. In this embodiment, the similarity S may be used as an initial similarity value, which is the average of the three maximum values ​​of the similarities s1, s2, ... and sn.

[0067] In step S4022, a homography transformation is performed on any coordinate in the adjacent sub-image block of the pixel P in the processing area R1 using the following formula:

[0068] ;

[0069] Where, is the current coordinate in the adjacent sub-image block; is the transformed pixel coordinates corresponding to the current coordinates in the adjacent sub-image block of any adjacent photo in the reference frame group; Obtained by the following formula:

[0070] ;

[0071] Where, and is the internal reference information of the surveying and mapping camera; and is the rotation matrix of the mapping camera; and is the translation vector of the mapping camera; is the normal vector of the preset physical plane; The depth vector of the preset physical plane in the camera coordinate system.

[0072] In step S4022, a weighted NCC (Normalized Cross Correlation) matching algorithm is used to calculate the similarity between each reference frame pixel information group and the current adjacent sub-image block. Accordingly, the similarity between any reference frame pixel information group and the current adjacent sub-image block is:

[0073] ;

[0074] [Corrected 29.12.2023 according to Rule 91] where w(i,j) is the preset NCC weight; I(i,j) is the grayscale value of the pixel in the current adjacent sub-image block, p(i,j) is the coordinate of the pixel in the current adjacent sub-image block, I c is the gray value of the pixel at the center of the current adjacent sub-image block, p c is the coordinate of the pixel at the center of the current adjacent sub-image block, σ I is the preset grayscale weight parameter, σ p is the preset coordinate weight parameter.

[0075] S4023. Perform depth and normal vector propagation optimization and depth and normal vector random optimization on all pixels in the processing area R1 in sequence, so as to obtain the final effective depth and normal vector information of all pixels based on the effective depth and normal vector of all pixels in the processing area R1.

[0076] During the depth and normal vector propagation optimization and the depth and normal vector random optimization are sequentially performed on all pixels in the processing area R1, the depth and normal vector propagation optimization is performed on the pixel P in the processing area R1, including: if the pixel P has a valid depth and normal vector E1, the valid depth and normal vector E1 are assigned to the adjacent pixels of the pixel P, and the similarity between the current adjacent pixel and the reference frame group is calculated using the valid depth and normal vector E1, if the similarity is greater than the current similarity of the current adjacent pixel, the assignment is accepted and the similarity of the current adjacent pixel is updated, and then the next step is entered; otherwise, the assignment is rejected, and the depth and normal vector random optimization is directly performed on the adjacent pixels of the pixel P;

[0077] performing random depth and normal optimization on a pixel P in the processing region R1, including: setting a perturbation on an estimated value of a current effective depth and a normal vector of the pixel P, obtaining a similarity between the pixel P and the reference frame group after setting the perturbation, retaining an estimated value with a higher similarity as final effective depth and normal vector information and assigning it to the pixel P, and recording the similarity;

[0078] By performing multiple iterations of propagation optimization and random optimization on all pixels in the region R1 , final effective depth and normal vector information of all pixels in the processing region R1 can be obtained.

[0079] S403. Reselect a pixel in the processing region R1 until the final effective depth and normal vector information of all pixels in the processing region R1 is obtained, and the final effective depth and normal vector information of all pixels in the processing region R1 is stored in the image feature set M1 to obtain the image feature set M1 corresponding to the processing region R1;

[0080] S404. Take the average value of all the final effective depths and normal vector information in the image feature set M1 to obtain the initial value of the three-dimensional coordinate X of the measured point A1 in the geodetic coordinate system and the initial normal vector of the physical plane where the measured point A1 is located.

[0081] In this embodiment, the initial value of the three-dimensional coordinate X is calculated as follows:

[0082] ;

[0083] Where, is the three-dimensional coordinate of the point A1 to be measured, is the coordinate of the point A1 to be measured in the camera coordinate system, is the average value of the final effective depth in the image feature set M1, is the coordinate of the surveying camera in the geodetic coordinate system, The rotation matrix of the camera coordinate system of the mapping camera relative to the geodetic reference system.

[0084] S5. Back-project the initial value of the three-dimensional coordinate X onto any adjacent photograph in the reference frame group to obtain the test point A2 corresponding to the test point A1, and set this adjacent photograph as a new base frame. Then, reselect a set of adjacent photographs in the reference frame group that have a co-viewing relationship with the new base frame as a new reference frame group, until the image feature set M1, image feature set M2, ..., and image feature set Mn corresponding to all processing areas in the reference frame group are obtained; where n is the number of adjacent photographs in the initial reference frame group.

[0085] In step S5, the initial value of the three-dimensional coordinate X is reversely projected onto any adjacent photo in the reference frame group using the following formula to obtain the test point A2 corresponding to the test point A1:

[0086] ;

[0087] Where, is the coordinate of the point A2 to be measured, is the internal reference information of the surveying and mapping camera, is the initial value of the three-dimensional coordinate X, is the three-dimensional coordinate of the surveying camera, is the transpose of the rotation matrix of the camera coordinate system of the surveying camera relative to the earth reference system, and Z is the depth value of the three-dimensional coordinate X in the camera coordinate system of the surveying camera, which is an intermediate temporary variable.

[0088] In this embodiment, the initial value of the three-dimensional coordinate X is back-projected onto any adjacent photo in the reference frame group to obtain the measured point A2 corresponding to the measured point A1, and this adjacent photo is set as a new base frame. When a group of adjacent photos in the reference frame group that have a co-viewing relationship with the new base frame are reselected as a new reference frame group, multi-core CPUs are used for parallel execution, thereby accelerating data processing speed, so that this embodiment has the technical effect of high operating efficiency.

[0089] S6. Merge the image feature set M1, the image feature set M2, ..., and the image feature set Mn to obtain a common view matching result for all pixels surrounding the test point A1. It should be noted that, in this embodiment, the common view matching result for all pixels surrounding the test point A1 is the common view matching result for all pixels within a specified two-dimensional square region surrounding the test point A1, where the position of the specified two-dimensional square region corresponds to the position of the processing region R1 corresponding to the test point A1.

[0090] In step S6, the image feature set M1, the image feature set M2, ..., and the image feature set Mn are merged using a preset similarity rule; correspondingly, the image feature set M1, the image feature set M2, ..., and the image feature set Mn are merged using a preset similarity rule to obtain a common view matching result of all pixels around the test point A1, including:

[0091] S601. Select any image feature set M from the image feature set M1, the image feature set M2, ..., and the image feature set Mn, and obtain any final effective depth and normal vector information E3 in the image feature set M, the final effective depth and normal vector information E3 corresponding to the pixel P1;

[0092] S602. Project the final effective depth and normal vector information E3 onto any reference frame in the reference frame group to obtain the corresponding pixel P2 in the reference frame;

[0093] S603. Use the coordinates of pixel P2 as an index to query the image feature set Mn to obtain the final effective depth and normal vector information E4 corresponding to pixel P2, and compare the depth difference and normal vector angle between the final effective depth and normal vector information E3 and the final effective depth and normal vector information E4 to see if they are less than a preset threshold. If so, determine that pixel P1 and pixel P2 are common view matching points, that is, it is considered that pixel P1 and pixel P2 are actually the same physical point on the surface of the object to be measured, and then proceed to the next step; the preset threshold includes an angle threshold and a depth threshold, compare the depth difference and normal vector angle between the final effective depth and normal vector information E3 and the final effective depth and normal vector information E4 to see if they are less than the preset threshold, that is, compare whether the angle between the normal vectors in the final effective depth and normal vector information E3 and the final effective depth and normal vector information E4 is less than the preset angle threshold, and compare whether the depths in the final effective depth and normal vector information E3 and the final effective depth and normal vector information E4 are less than the preset depth threshold;

[0094] S604. Delete the final effective depth and normal vector information E3 and the final effective depth and normal vector information E4 from the image feature set M and the image feature set Mn respectively;

[0095] S605. Reselect any one of the image feature sets M1, M2, ... and Mn until a common view matching result of all pixels around the point A1 to be measured is obtained.

[0096] The beneficial effects of this embodiment are as follows: high matching accuracy, which can reach the sub-pixel level; high availability, which can realize automatic matching of multi-frame common viewpoints of user-selected points in an environment without feature points. Specifically, during the implementation of this embodiment, by receiving photos of the object to be measured taken by a surveying camera at different positions, any one of the photos is used as a basic frame, and a point to be measured A1 is selected on the basic frame, and then a reference frame is selected, and the specified two-dimensional square area around the point to be measured A1 in the basic frame is used as the processing area R1, and the image feature set M1 corresponding to the processing area R1 and the initial value of the three-dimensional coordinate X of the point to be measured A1 in the geodetic coordinate system are obtained according to the reference frame group, and then the initial value of the three-dimensional coordinate X is reversely projected to any adjacent photo in the reference frame group to obtain the point to be measured A2 corresponding to the point to be measured A1. , and set this adjacent photo as a new base frame, and then reselect a group of adjacent photos in the reference frame group that have a common view relationship with the new base frame as a new reference frame group, until the image feature set M1, image feature set M2, ... and image feature set Mn corresponding to all processing areas in the reference frame group are obtained. Finally, by merging the image feature set M1, image feature set M2, ... and image feature set Mn, the common view matching result of all pixels around the test point A1 is obtained, thereby realizing the common view matching of multiple frames of photos. In this process, the matching relationship can reach the sub-pixel level, and even in areas without corner points or feature points, automatic matching of user-specified points can be completed, which has higher usability.

[0097] Furthermore, in this implementation, each time multi-frame common view matching is performed, a subregion around the user-selected test point A1 is first processed. Within this subregion, a rapid iterative search for depth and plane orientation is performed using random optimization and propagation. The accuracy of these depth and plane orientations is evaluated using homography and the NCC cost function to guide the search. Ultimately, the approximate depth information of the user-selected point is obtained. This depth information is used to calculate the processing subregion of the reference frame. Random optimization and propagation are then performed on this subregion within the reference frame. This avoids full-frame computation of the reference frame, significantly improving processing speed and enabling this method to run on low-performance embedded devices. After random optimization and propagation have been completed on all reference frames, multi-frame merging is performed.

[0098] Example 2:

[0099] Based on Example 1, this embodiment discloses an electronic device, which may be a smartphone, tablet computer, laptop computer, or desktop computer. The electronic device may be called a user terminal, portable terminal, desktop terminal, etc. As shown in FIG2 , the electronic device includes:

[0100] a memory for storing computer program instructions; and

[0101] A processor is used to execute the computer program instructions to complete the operation of the image matching method as described in any one of Example 1.

[0102] [Corrected 29.12.2023 according to Rule 91] Specifically, the processor 101 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 101 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 101 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 101 may be further integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen.

[0103] [Corrected 29.12.2023 in accordance with Rule 91] Memory 102 may include one or more computer-readable storage media, which may be non-transitory. Memory 102 may also include high-speed random access memory and non-volatile memory, such as one or more magnetic disk storage devices or flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in memory 102 is used to store at least one instruction, which is executed by processor 101 to implement the image matching method provided in Example 1 of this application.

[0104] [Corrected 29.12.2023 according to Rule 91] In some embodiments, the terminal may optionally include a communication interface 103 and at least one peripheral device. The processor 101, memory 102, and communication interface 103 may be connected via a bus or signal lines. Each peripheral device may be connected to the communication interface 103 via a bus, signal lines, or circuit boards. Specifically, the peripheral device includes at least one of a radio frequency circuit 104, a display screen 105, and a power supply 106.

[0105] [Corrected 29.12.2023 according to Rule 91] The communication interface 103 can be used to connect at least one I / O (Input / Output)-related peripheral device to the processor 101 and the memory 102. In some embodiments, the processor 101, the memory 102, and the communication interface 103 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 101, the memory 102, and the communication interface 103 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0106] [Corrected 29.12.2023 according to Rule 91] The RF circuit 104 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 104 communicates with a communication network and other communication devices via electromagnetic signals.

[0107] [Corrected 29.12.2023 according to Rule 91] The display screen 105 is used to display a UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof.

[0108] [Corrected 29.12.2023 according to Rule 91] The power supply 106 is used to supply power to various components in the electronic device.

[0109] Example 3:

[0110] Based on any one of Examples 1 to 2, this embodiment discloses a computer-readable storage medium for storing computer-readable computer program instructions, wherein the computer program instructions are configured to execute the operations of the image matching method described in Example 1 when run.

[0111] Obviously, those skilled in the art will appreciate that the modules or steps of the present invention described above can be implemented using a general-purpose computing device. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Alternatively, they can be implemented using program code executable by a computing device, which can then be stored in a storage device and executed by the computing device. Alternatively, they can be fabricated into separate integrated circuit modules, or multiple modules or steps can be fabricated into a single integrated circuit module for implementation. Thus, the present invention is not limited to any specific combination of hardware and software.

[0112] Finally, it should be noted that the above embodiments are intended only to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art will appreciate that modifications may be made to the technical solutions described in the above embodiments, or that some of the technical features may be replaced with equivalents. Such modifications or replacements do not deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. An image matching method, characterized in that: It includes: Receiving the photographed pictures of the object to be measured taken by the surveying and mapping camera at different positions; Taking any one of the photographed pictures as a base frame, and selecting a measurement point A1 on the base frame; Selecting a group of adjacent pictures having a co-visibility relationship with the base frame from the photographed pictures, and taking this group of adjacent pictures as a reference frame group; Taking the specified two-dimensional square area around the measurement point A1 in the base frame as a processing area R1, and obtaining an image feature set M1 corresponding to the processing area R1 and an initial value of the three-dimensional coordinate X of the measurement point A1 in the geodetic coordinate system according to the reference frame group; wherein, the image feature set M1 includes the final effective depth and normal vector information of all pixels in the processing area R1; Back-projecting the initial value of the three-dimensional coordinate X to any one of the adjacent pictures in the reference frame group to obtain a measurement point A2 corresponding to the measurement point A1, setting this adjacent picture as a new base frame, and then re-selecting a group of adjacent pictures having a co-visibility relationship with the new base frame in the reference frame group as a new reference frame group until obtaining the image feature sets M1, M2,..., and Mn corresponding to all processing areas in the reference frame group; wherein, n is the number of adjacent pictures in the initial reference frame group; Merging the image feature sets M1, M2,..., and Mn to obtain the co-visibility matching result of all pixels around the measurement point A1.

2. The image matching method according to claim 1, characterized in that: Obtaining the image feature set M1 corresponding to the processing area R1 and the initial value of the three-dimensional coordinate X of the measurement point A1 in the geodetic coordinate system according to the reference frame group includes: Selecting any pixel P in the processing area R1, and determining that the pixel P and a specified number of adjacent pixels adjacent to the pixel P are on the same physical plane of the object to be measured; Obtaining the final effective depth and normal vector information of all pixels in the processing area R1 according to the reference frame group; Re-selecting a pixel in the processing area R1 until obtaining the final effective depth and normal vector information of all pixels in the processing area R1, and storing the final effective depth and normal vector information of all pixels in the processing area R1 into the image feature set M1 to obtain the image feature set M1 corresponding to the processing area R1; Taking the average value of all the final effective depth and normal vector information in the image feature set M1 to obtain the initial value of the three-dimensional coordinate X of the measurement point A1 in the geodetic coordinate system.

3. The image matching method according to claim 2, characterized in that: Obtaining the final effective depth and normal vector information of all pixels in the processing area R1 according to the reference frame group includes: Randomly initializing all pixels in the processing area R1 to obtain the initial depth and normal vector of the physical plane where each pixel in the processing area R1 is located; Obtain the similarity between all pixels in the processing region R1 and the reference frame group, and when the similarity is greater than the preset similarity threshold, use the initial depth and normal vector of the physical plane where the corresponding pixel is located as the effective depth and normal vector of the pixel; Perform depth and normal vector propagation optimization and depth and normal vector random optimization on all pixels in the processing region R1 in sequence, so as to obtain the final effective depth and normal vector information of all pixels based on the effective depth and normal vector of all pixels in the processing region R1.

4. An image matching method according to claim 3, characterized in that: Obtaining the similarity between the pixel P in the processing region R1 and the reference frame group, and when the similarity is greater than the preset similarity threshold, using the initial depth and normal vector of the physical plane where the pixel P is located as the effective depth and normal vector of the pixel, includes: Perform a homography transformation on the coordinates of the adjacent sub-image block of the pixel P in the processing region R1 to obtain the transformed pixel coordinates corresponding to the adjacent sub-image block in all adjacent photos in the reference frame group, and obtain the reference frame pixel information group of the area corresponding to the transformed pixel coordinates in all adjacent photos in the reference frame group, so as to calculate the similarity between each reference frame pixel information group and the current adjacent sub-image block, and obtain similarities s1, s2,..., and sn, where n is the number of adjacent photos in the initial reference frame group; take the average value of the three largest similarities among similarities s1, s2,..., and sn as the similarity S between the pixel P and the reference frame group, and when the similarity S is greater than the preset similarity threshold, use the initial depth and normal vector of the physical plane where the pixel P is located as the effective depth and normal vector E1 of the pixel P.

5. An image matching method according to claim 4, characterized in that: During the sequential depth and normal vector propagation optimization and depth and normal vector random optimization of all pixels in the processing region R1, performing depth and normal vector propagation optimization on the pixel P in the processing region R1 includes: if the pixel P has an effective depth and normal vector E1, then assign the effective depth and normal vector E1 to the adjacent pixels of the pixel P, and use the effective depth and normal vector E1 to calculate the similarity between the current adjacent pixel and the reference frame group. If the similarity is greater than the current similarity of the current adjacent pixel, accept this assignment and update the similarity of the current adjacent pixel, and then proceed to the next step; otherwise, reject this assignment and directly perform depth and normal vector random optimization on the adjacent pixels of the pixel P; Performing depth and normal vector random optimization on the pixel P in the processing region R1 includes: setting a perturbation on the estimated value of the current effective depth and normal vector of the pixel P, and obtaining the similarity between the pixel P and the reference frame group after setting the perturbation, and then retaining the estimated value with a higher similarity as the final effective depth and normal vector information and assigning it to the pixel P.

6. An image matching method according to claim 4, characterized in that: Perform a homography transformation on any coordinate within the adjacent sub-image block of the pixel P in the processing region R1 using the following formula: ; In the formula, is the current coordinate within the adjacent sub-image block; is the transformed pixel coordinate corresponding to the current coordinate in the neighboring sub-image block for any neighboring photo in the reference frame group; matrix Obtained by the following formula: ; In the formula, And are the internal parameter information of the mapping camera; And is the rotation matrix of the mapping camera; And is the translation vector of the surveying and mapping camera; is the normal vector of the preset physical plane; is the depth vector of the preset physical plane in the camera coordinate system.

7. A method for image matching according to claim 4, characterized in that: When calculating the similarity between each reference frame pixel information group and the current adjacent sub-image block, the weighted NCC matching algorithm is used; correspondingly, the similarity between any reference frame pixel information group and the current adjacent sub-image block is: ; In the formula, is a preset NCC weight; , is the gray value of the pixel in the current adjacent sub-image block, is the coordinate of the pixel in the current adjacent sub-image block, is the gray value of the pixel at the center point of the current adjacent sub-image block, are the coordinates of the pixel at the center point of the current adjacent sub-image block, is a preset gray-scale weight parameter, is the preset coordinate weight parameter.

8. A method for image matching according to claim 1, characterized in that: Use the following formula to back-project the initial value of the three-dimensional coordinate X to any adjacent photo in the reference frame group to obtain the corresponding measurement point A2 of the measurement point A1: ; In the formula, are the coordinates of the point A2 to be measured, is the internal parameter information of the mapping camera, is the initial value of the three-dimensional coordinate X, are the three-dimensional coordinates of the mapping camera, is the transpose of the rotation matrix of the camera coordinate system of the mapping camera relative to the geodetic reference system, and Z is the depth value of the three-dimensional coordinate X in the camera coordinate system of the mapping camera.

9. A method for image matching according to claim 1, characterized in that: Back-project the initial value of the three-dimensional coordinate X to any adjacent photo in the reference frame group to obtain the corresponding measurement point A2 of the measurement point A1, and set this adjacent photo as the new base frame. Then, when re-selecting a group of adjacent photos having a co-visibility relationship with the new base frame in the reference frame group as the new reference frame group, multi-core CPU parallel execution is adopted.

10. A method for image matching according to claim 1, characterized in that: Merge the image feature sets M1, image feature sets M2,..., and image feature sets Mn to obtain the co-visibility matching results of all pixels around the measurement point A1, including: Arbitrarily select an image feature set M from the image feature sets M1, image feature sets M2,..., and image feature sets Mn, and obtain any final effective depth and normal vector information E3 in the image feature set M, and the final effective depth and normal vector information E3 corresponds to the pixel P1; Project the final effective depth and normal vector information E3 onto any reference frame in the reference frame group to obtain the corresponding pixel P2 in this reference frame; Use the coordinates of the pixel P2 as an index to query in the image feature set Mn to obtain the final effective depth and normal vector information E4 corresponding to the pixel P2, and compare whether the depth difference and normal vector angle between the final effective depth and normal vector information E3 and the final effective depth and normal vector information E4 are less than the preset threshold. If so, determine that the pixel P1 and the pixel P2 are co-visibility matching points, and then proceed to the next step; Delete the final effective depth and normal vector information E3 and the final effective depth and normal vector information E4 from the image feature set M and the image feature set Mn respectively; Arbitrarily select an image feature set from the image feature sets M1, image feature sets M2,..., and image feature sets Mn again until the co-visibility matching results of all pixels around the measurement point A1 are obtained.

Citation Information

Patent Citations

  • On-orbit real-time image stabilizing method and system for video satellite image

    CN108076341A

  • Photogrammetry method and device, and storage medium

    CN110360991A

  • Method and device for achieving registration between image frames and storage medium

    CN111681270A

  • Graphical coordinate system transform for video frames

    US20190197709A1