Image registration method, device, storage medium and electronic device
By determining the relative position and spatial sampling of the drone and the unmanned vehicle, and establishing a coarse affine model, the problem of ground and aerial image registration is solved, and accurate image registration is achieved without prior map information.
Patent Information
- Application Number
- CN202310299814.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-20
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2043-03-20
AI Technical Summary
In low-altitude or indoor environments in urban areas, the perspective angles of ground and aerial images vary greatly, and it is difficult for the prior art to effectively realize the registration of ground acquisition images and aerial acquisition images.
By acquiring the aerial images collected by the drone and the ground images collected by the drone, determining the relative position of the drone and the drone, establishing a coarse affine model, and performing spatial sampling around the initial position to determine the coarse affine model of the spatial sampling point, and finally obtaining the fine affine model by matching feature points and adjusting parameters to achieve image registration.
Without prior map information, registration of aerial images and ground images can be achieved in low-altitude urban areas or indoor environments without environmental constraints.
Smart Images

Figure CN116258753B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to an image registration method, device, storage medium, and electronic device. Background Art
[0002] With the advancement of technology, artificial intelligence (AI) has rapidly developed. Among these, ground-to-air collaborative navigation systems have garnered widespread attention. In low-altitude urban environments or indoor environments, Global Positioning System (GPS) signals are often obscured and unusable. Therefore, vision-based positioning methods have become particularly important.
[0003] In ground-air collaborative navigation applications, determining a registration method for ground-collected and aerial images—that is, establishing a correlation between the ground and aerial data—is fundamental to ensuring the feasibility of ground-air collaborative navigation. However, the significant difference in perspective between ground and aerial images complicates their registration. Therefore, achieving this registration is a pressing issue.
[0004] Based on this, this application specification provides an image registration method. Summary of the Invention
[0005] This specification provides an image registration method, device, medium, and electronic device to at least partially solve the above-mentioned problems existing in the prior art.
[0006] This manual adopts the following technical solutions:
[0007] This specification provides an image registration method, the method comprising:
[0008] Acquire aerial images collected by a drone and ground images collected by an unmanned vehicle, and determine the relative positions of the drone and the unmanned vehicle;
[0009] Determining a coarse affine model between the aerial image and the ground image according to the relative pose as the coarse affine model corresponding to the initial position of the drone when acquiring the aerial image;
[0010] Determining spatial sampling points around an initial position of the drone when the drone collects the aerial image;
[0011] Determining an aerial image corresponding to each spatial sampling point based on the aerial image corresponding to the initial position, and determining a coarse affine model corresponding to each spatial sampling point based on the coarse affine model corresponding to the initial position;
[0012] Determining, among the aerial images, an aerial image that matches the ground image as a matching image;
[0013] A spatial sampling point corresponding to the matching image is determined as a matching sampling point, and a fine affine model between the UAV and the unmanned vehicle is obtained based on a coarse affine model corresponding to the matching sampling point.
[0014] Optionally, determining the aerial image corresponding to each spatial sampling point according to the aerial image corresponding to the initial position specifically includes:
[0015] For each spatial sampling point, determining a relative position between the initial position and the spatial sampling point;
[0016] The aerial image corresponding to the spatial sampling point is determined according to the relative position and the aerial image corresponding to the initial position.
[0017] Optionally, determining a coarse affine model corresponding to each spatial sampling point according to the coarse affine model corresponding to the initial position specifically includes:
[0018] For each spatial sampling point, determining the relative position between the initial position and the spatial sampling point;
[0019] Determine the relative position of the drone and the unmanned vehicle when the drone is located at the spatial sampling point according to the relative position of the drone and the unmanned vehicle and the relative position of the initial position and the spatial sampling point;
[0020] According to the relative posture of the UAV and the unmanned vehicle when the UAV is located at the spatial sampling point, a coarse affine model corresponding to the spatial sampling point is determined.
[0021] Optionally, determining an aerial image that matches the ground image specifically includes:
[0022] extracting image features of each aerial image and extracting image features of the ground image;
[0023] An aerial image matching the ground image is determined based on the image features of each aerial image and the image features of the ground image.
[0024] Optionally, obtaining a fine affine model between the UAV and the unmanned vehicle based on the coarse affine model corresponding to the matching sampling points specifically includes:
[0025] Extracting image features of each pixel in the matching image as first features;
[0026] Determining, according to the coarse affine model corresponding to the matching sampling points, the number of first features among the first features that match the ground image;
[0027] If the number is greater than a preset threshold, the coarse affine model corresponding to the matching sampling point is determined as the fine affine model between the UAV and the unmanned vehicle;
[0028] If the number is not greater than a preset threshold, the parameters of the coarse affine model corresponding to the matching sampling points are adjusted according to the first features, and the number of first features among the first features that match the ground image is re-determined according to the adjusted coarse affine model corresponding to the matching sampling points, until the number is greater than the preset threshold.
[0029] Optionally, determining the number of first features among the first features that match the ground image specifically includes:
[0030] Extracting a second feature of each pixel in the ground image;
[0031] For each pixel point in the matching image, the pixel point is used as an aerial pixel point, and a ground pixel point corresponding to the aerial pixel point in the ground image is determined based on a coarse affine model corresponding to the matching sampling point;
[0032] According to the similarity between the first feature of the aerial pixel point and the second feature of the ground pixel point corresponding to the aerial pixel point, it is determined whether the first feature of the aerial pixel point matches the ground image.
[0033] Optionally, adjusting the parameters of the coarse affine model corresponding to the matching sampling points according to the first features specifically includes:
[0034] Selecting a plurality of feature points in the matching image as selected feature points, and extracting a second feature of each ground pixel point in the ground image;
[0035] Performing an affine transformation on each selected feature point according to a coarse affine model corresponding to the matching sampling point to obtain each affine feature corresponding to each selected feature point, and determining a ground pixel point in the ground image corresponding to each selected feature point;
[0036] The parameters of the coarse affine model corresponding to the matching sampling points are adjusted with the maximization of the similarity between the second features of the affine features and the ground pixel points in the ground image corresponding to the selected feature points as the optimization goal.
[0037] This specification provides an image registration device, comprising:
[0038] An acquisition module is used to acquire aerial images collected by the drone and ground images collected by the unmanned vehicle, and determine the relative position of the drone and the unmanned vehicle;
[0039] a coarse affine determination module, configured to determine a coarse affine model between the aerial image and the ground image according to the relative pose, as the coarse affine model corresponding to the initial position of the drone when acquiring the aerial image;
[0040] A sampling module, configured to determine, based on an initial position at which the drone collects the aerial image, spatial sampling points around the initial position;
[0041] an expansion module, configured to determine an aerial image corresponding to each spatial sampling point based on the aerial image corresponding to the initial position, and to determine a coarse affine model corresponding to each spatial sampling point based on the coarse affine model corresponding to the initial position;
[0042] An image matching module is used to determine, among the aerial images, an aerial image that matches the ground image as a matching image;
[0043] The fine affine determination module is used to determine the spatial sampling points corresponding to the matching image as matching sampling points, and obtain the fine affine model between the UAV and the unmanned vehicle based on the coarse affine model corresponding to the matching sampling points.
[0044] Optionally, the expansion module is specifically configured to determine, for each spatial sampling point, a relative position between the initial position and the spatial sampling point; and determine an aerial image corresponding to the spatial sampling point based on the relative position and the aerial image corresponding to the initial position.
[0045] Optionally, the extension module is specifically used to determine, for each spatial sampling point, the relative position of the initial position and the spatial sampling point; determine the relative position of the drone and the unmanned vehicle when located at the spatial sampling point based on the relative position of the drone and the unmanned vehicle and the relative position of the initial position and the spatial sampling point; and determine the coarse affine model corresponding to the spatial sampling point based on the relative position of the drone and the unmanned vehicle when located at the spatial sampling point.
[0046] Optionally, the image matching module is specifically used to extract image features of each aerial image and extract image features of the ground image; and determine an aerial image that matches the ground image based on the image features of each aerial image and the image features of the ground image.
[0047] Optionally, the fine affine determination module is specifically used to extract the image features of each pixel point in the matching image as each first feature; determine the number of first features matching the ground image in each first feature according to the coarse affine model corresponding to the matching sampling point; if the number is greater than a preset threshold, determine the coarse affine model corresponding to the matching sampling point as the fine affine model between the UAV and the unmanned vehicle; if the number is not greater than the preset threshold, adjust the parameters of the coarse affine model corresponding to the matching sampling point according to each first feature, and re-determine the number of first features matching the ground image in each first feature according to the adjusted coarse affine model corresponding to the matching sampling point, until the number is greater than the preset threshold.
[0048] Optionally, the fine affine determination module is specifically used to extract the second feature of each pixel point in the ground image; for each pixel point in the matching image, the pixel point is used as an aerial pixel point, and the ground pixel point corresponding to the aerial pixel point in the ground image is determined according to the coarse affine model corresponding to the matching sampling point; and based on the similarity between the first feature of the aerial pixel point and the second feature of the ground pixel point corresponding to the aerial pixel point, it is determined whether the first feature of the aerial pixel point matches the ground image.
[0049] Optionally, the fine affine determination module is specifically used to select multiple feature points in the matching image as selected feature points, and extract the second features of each ground pixel point in the ground image; perform affine transformation on each selected feature point according to the coarse affine model corresponding to the matching sampling point to obtain each affine feature corresponding to each selected feature point, and determine the ground pixel point in the ground image corresponding to each selected feature point; and adjust the parameters of the coarse affine model corresponding to the matching sampling point with the maximum similarity between each affine feature and the second feature of the ground pixel point in the ground image corresponding to each selected feature point as the optimization goal.
[0050] This specification provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned image registration method.
[0051] This specification provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned image registration method when executing the program.
[0052] At least one of the above technical solutions adopted in this specification can achieve the following beneficial effects:
[0053] In the image registration method provided in this specification, an aerial image collected by a drone and a ground image collected by an unmanned vehicle can be first obtained, and the relative position and posture of the drone and the unmanned vehicle can be determined. Based on the relative position and posture, a coarse affine model between the aerial image and the ground image can be determined. Then, based on the initial position of the drone when collecting the aerial image, the spatial sampling points around the initial position are determined, and based on the aerial image corresponding to the initial position, the aerial image corresponding to each spatial sampling point is determined, and based on the coarse affine model corresponding to the initial position, the coarse affine model corresponding to each spatial sampling point is determined. Finally, the aerial image that matches the ground image in each aerial image is determined as the matching image, and the spatial sampling points corresponding to the matching image are determined as the matching sampling points. Based on the coarse affine model corresponding to the matching sampling point, a fine affine model between the drone and the unmanned vehicle is obtained.
[0054] As can be seen from the above method, by determining a coarse affine model between the drone and the unmanned vehicle at its initial position, and a coarse affine model between the drone and the unmanned vehicle at each spatial sampling point, and then using the aerial image captured by the drone at the initial position and the determined aerial images corresponding to the drone at each spatial sampling point, an aerial image that matches the ground image is determined. Based on the coarse affine model of the spatial sampling points corresponding to the aerial image, a fine affine model between the drone and the unmanned vehicle can be derived. This registration of aerial and ground images is still possible without prior map information. Furthermore, it is not constrained by the environment, meaning that it can also be achieved in low-altitude urban areas or indoor environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] The accompanying drawings described herein are used to provide a further understanding of this specification and constitute a part of this specification. The exemplary embodiments and descriptions of this specification are used to explain this specification and do not constitute an improper limitation of this specification.
[0056] In the picture:
[0057] Figure 1 A flowchart of an image registration method in this specification;
[0058] Figure 2 A schematic diagram of spatial sampling provided in this specification;
[0059] Figure 3 A schematic diagram of an image registration device provided in this specification;
[0060] Figure 4 The corresponding Figure 1 Schematic diagram of electronic equipment. DETAILED DESCRIPTION
[0061] To make the objectives, technical solutions, and advantages of this specification more clear, the following will provide a clear and complete description of the technical solutions of this specification in conjunction with the specific embodiments of this specification and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this specification, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments in this specification without inventive effort are within the scope of protection of this specification.
[0062] The technical solutions provided by the embodiments of this specification are described in detail below with reference to the accompanying drawings.
[0063] Figure 1 This is a flow chart of an image registration method provided in this specification, which may specifically include the following steps:
[0064] S100: Acquire an aerial image collected by a drone and a ground image collected by an unmanned vehicle, and determine the relative position of the drone and the unmanned vehicle.
[0065] Generally, when realizing the registration of ground images and aerial images based on visual features, a priori maps need to be used. For example: the ground-to-air collaborative navigation system first obtains the offline visual feature map of the navigation environment, and then when the ground-to-air collaborative navigation system is running, the collected features can be matched with the features in the offline visual feature map, and the position of the unmanned vehicle is estimated through a discrete Bayesian filter, thereby completing the registration of the ground image and the aerial image. However, the image registration method that relies on the prior map cannot be applied to indoor environments and other environments that are not covered by the prior map. The image registration method that does not rely on the prior map needs to find an area with consistent perspective for the unmanned vehicle and the drone’s field of view to extract visual features, and is therefore constrained by the environment. Based on this, the present application specification provides an image registration method that can achieve image registration without using a priori maps and is not constrained by the environment.
[0066] The execution subject of the embodiments of this application specification can be any device with computing capabilities (such as a server, terminal, etc.). The technical solution of this specification is now described with the server as the execution subject.
[0067] In one or more embodiments of the present specification, the server may first obtain aerial images collected by the drone and ground images collected by the unmanned vehicle, and determine the relative positions of the drone and the unmanned vehicle.
[0068] Specifically, for the same target object, the shooting equipment (such as a camera) on the drone can shoot the target object and determine the position when shooting the target object. The position can be expressed by the spherical coordinate longitude, latitude and roll angle of the camera relative to the ground coordinate system. Similarly, the shooting equipment in the unmanned vehicle can also shoot the target object and determine the position when shooting the target object. Furthermore, the communication equipment on the drone and the unmanned vehicle can establish a connection with the server, and the server can obtain the aerial image of the target object collected by the drone and the ground image of the target object collected by the unmanned vehicle, and can determine the relative posture of the drone and the unmanned vehicle when shooting the target object based on the position of the drone when shooting the target object and the position of the unmanned vehicle when shooting the target object. In the subsequent steps, a coarse affine model between the aerial image and the ground image is determined based on the relative posture.
[0069] When determining the position of a drone when photographing a target, this position can be derived from the drone's inertial navigation output, or from the planned trajectory output by a planner. Similarly, the position of an unmanned vehicle when photographing a target can also be determined based on inertial navigation, lidar, and other sensors on the vehicle.
[0070] S102: Determine a coarse affine model between the aerial image and the ground image according to the relative posture, as the coarse affine model corresponding to the initial position of the drone when collecting the aerial image.
[0071] In general, according to the first-order Taylor formula, any smooth deformation of a plane can be approximated locally using an affine mapping, and the surface deformation of the target object caused by the camera is a planar homography, which is a smooth deformation. Therefore, the deformation of the target object's appearance caused by the camera shooting the target at different angles can be represented by an affine mapping. In other words, the local effect of the perspective transformation can be modeled by a local affine mapping, i.e., u(x,y)→u(ax+by+e,cx+dy+f), where u(x,y) represents the coordinates of each pixel in the aerial image captured by the drone, u(ax+by+e,cx+dy+f) represents the coordinates of each pixel in the ground image captured by the unmanned vehicle, and a, b, c, d, e, and f are the parameters of the affine transformation.
[0072] Therefore, in one or more embodiments of this specification, after obtaining the relative pose of the drone and the unmanned vehicle (that is, the relative pose when the camera shoots the target at different angles), the server can obtain the conversion relationship between the relative pose and the affine transformation by performing matrix decomposition on the affine mapping: That is, a coarse affine model between the aerial image and the ground image is obtained, and the coarse affine model can be used as a coarse affine model corresponding to the initial position of the UAV when collecting the aerial image.
[0073] Among them, φ and represents the visual parameters, that is, the spherical coordinate longitude and latitude of the UAV relative to the ground coordinate system of the unmanned vehicle, ψ is the roll angle parameter of the camera in the UAV relative to the camera in the unmanned vehicle, λ is the scaling factor, and A is the transformation relationship between the coordinates of the pixel points in the aerial image and the pixel points in the ground image.
[0074] It should be noted that since the positions of the drone when collecting aerial images and the unmanned vehicle when collecting ground images are both rough, the determined relative pose is also a rough relative pose, not the precise relative pose between the drone and the unmanned vehicle. Therefore, the coarse affine model determined based on the relative pose is only an initialized model, that is, the coarse affine model is an affine model with low precision.
[0075] S104: Determine spatial sampling points around the initial position according to the initial position where the drone collects the aerial image.
[0076] S106: Determine the aerial image corresponding to each spatial sampling point based on the aerial image corresponding to the initial position, and determine the coarse affine model corresponding to each spatial sampling point based on the coarse affine model corresponding to the initial position.
[0077] Because the determined relative pose between the drone and the unmanned vehicle is a rough relative pose, not a precise relative pose, the aerial image obtained after being transformed using the coarse affine model corresponding to the drone at its initial position may not necessarily match the ground image. However, there exists an ideal affine model around the initial position. To find the ideal affine model (i.e., the affine model that best matches the aerial image with the ground image), in one or more embodiments of this specification, the server may perform spatial sampling around the initial position of the drone to obtain various spatial sampling points. The server then determines the aerial image corresponding to each spatial sampling point based on the aerial image corresponding to the initial position, and determines the coarse affine model corresponding to each spatial sampling point based on the coarse affine model corresponding to the initial position.
[0078] Figure 2 A schematic diagram of spatial sampling provided in this application specification, such as Figure 2As shown, circle P1 represents the initial position of the drone, circle P2 represents the position of the unmanned vehicle, and triangles 1 to 8 represent the spatial sampling points. Specifically, the server can perform 2-degree-of-freedom spatial sampling in the φ and θ dimensions based on the initial position of the drone to obtain the spatial sampling points. For example, taking the initial position of the drone as the starting point, five samples are taken in the φ dimension with a step size of 0.1°, and five samples are taken in the θ dimension with a step size of 0.1° to obtain 121 spatial sampling points. For each spatial sampling point, the server can determine the relative position between the initial position and the spatial sampling point, so as to obtain the spatial sampling points. Figure 2 Taking triangle 1 in the figure as an example, the relative position of the drone's initial position P1 and the position of triangle 1 is determined. Based on the relative position and the aerial image corresponding to the initial position, the aerial image corresponding to the spatial sampling point is determined. Based on the relative pose of the drone and the unmanned vehicle (i.e., P1 and P2) and the relative position of the initial position and the spatial sampling point (i.e., P1 and triangle 1), the relative pose of the drone and the unmanned vehicle (i.e., triangle 1 and P2) when located at the spatial sampling point can be determined. Then, based on the relative pose of the drone and the unmanned vehicle when located at the spatial sampling point, the coarse affine model corresponding to the spatial sampling point is determined.
[0079] Furthermore, when the server determines the aerial image corresponding to the spatial sampling point based on the relative position between the initial position and the spatial sampling point and the aerial image corresponding to the initial position, it can also perform modeling based on affine mapping to obtain a transformation relationship between the aerial image corresponding to the spatial sampling point and the aerial image corresponding to the initial position. In other words, the transformation relationship between the spatial sampling point and the initial position can be determined based on the relative position between the initial position and the spatial sampling point (i.e., the affine transformation caused by the change in camera position can be determined based on the relative position). The aerial image corresponding to the initial position can then be transformed based on this transformation relationship to obtain the aerial image corresponding to the spatial sampling point.
[0080] S108: Determine, among the aerial images, an aerial image that matches the ground image as a matching image.
[0081] In one or more embodiments of the present specification, in order to find the coarse affine model that best matches the aerial image and the ground image from each coarse affine model (that is, to determine the coarse affine model that is closest to the "ideal" affine model), the server may first determine the aerial image that matches the ground image among each aerial image as the matching image.
[0082] Specifically, since some visual feature points (such as SIFT, SURF, ORB, etc.) have good robustness against noise, perspective changes and illumination changes. That is to say, when the camera shooting angle changes, these visual feature points still have a high probability of being stably extracted. Therefore, in one or more embodiments of this specification, similarity matching of images at different shooting angles can be performed based on visual feature points (that is, matching of aerial images with ground images). The server can then extract features (visual feature points) from the aerial image and also perform feature extraction on the ground image. That is, feature extraction is performed on at least some of the pixel points in the aerial image, and feature extraction is performed on at least some of the pixel points in the ground image. Then, for each aerial image, the coarse affine model of the spatial sampling points corresponding to the aerial image is used to determine whether the features of the pixel points in the aerial image match the features of the pixel points in the ground image.
[0083] Then, for each aerial image, the number of features in the aerial image's pixels that match the features in the ground image's pixels is determined. The aerial image with the largest number of matching features is then deemed the image that best matches the ground image.
[0084] In addition, the features of the aerial images collected by the drone at the initial position and the pixel coordinates corresponding to each feature, the features of the aerial images corresponding to the drone at each spatial sampling point and the pixel coordinates corresponding to each feature, the index relationship between each feature of the aerial image corresponding to each spatial sampling point and each feature of the aerial image corresponding to the initial position, the index relationship between the aerial image corresponding to each spatial sampling point and the coarse affine model corresponding to each spatial sampling point, etc. can be added to the database to facilitate the matching of the features of the aerial image with the features of the ground image, and to facilitate the server to determine the correspondence between each spatial sampling point, each aerial image, and each coarse affine model.
[0085] S110: Determine spatial sampling points corresponding to the matching image as matching sampling points, and obtain a fine affine model between the UAV and the unmanned vehicle based on the coarse affine model corresponding to the matching sampling points.
[0086] To determine the affine model that best matches the transformed aerial image with the ground image, after determining the matching image that matches the ground image in step S108, the server can determine the spatial sampling points corresponding to the matching image as matching sampling points. Based on the coarse affine models corresponding to the matching sampling points, a refined affine model between the UAV and the unmanned vehicle is derived.
[0087] Specifically, the server may first extract the image features of each pixel in the matching image as each first feature. Then, based on the coarse affine model corresponding to the matching sampling point, the server determines the number of first features in each first feature that match the ground image. If the number is greater than a preset threshold, the coarse affine model corresponding to the matching sampling point is the fine affine model between the UAV and the unmanned vehicle. If the number is not greater than the preset threshold, the parameters of the coarse affine model corresponding to the matching sampling point may be adjusted based on each first feature. Based on the adjusted coarse affine model corresponding to the matching sampling point, the server re-determines the number of first features in each first feature that match the ground image until the number exceeds the preset threshold.
[0088] In one or more embodiments of the present specification, when determining the number of first features among the first features that match the ground image, the server may extract the second feature of each pixel in the ground image and, for each pixel in the matching image, identify the pixel as an aerial pixel. Then, based on the coarse affine model corresponding to the matching sampling point, the ground pixel in the ground image corresponding to the aerial pixel is determined. Finally, based on the similarity between the first feature of the aerial pixel and the second feature of the ground pixel corresponding to the aerial pixel, the server determines whether the first feature of the aerial pixel matches the ground image.
[0089] Furthermore, when adjusting the parameters of the coarse affine model corresponding to the matching sampling point based on each first feature, the server may select multiple feature points in the matching image as selected feature points and extract the second feature of each ground pixel in the ground image. Then, based on the coarse affine model corresponding to the matching sampling point, each selected feature point is affine transformed to obtain the affine features corresponding to each selected feature point. Based on the coarse affine model corresponding to the matching sampling point, the ground pixel in the ground image corresponding to each selected feature point is determined. Finally, the parameters of the coarse affine model corresponding to the matching sampling point are adjusted, with the optimization goal of maximizing the similarity between each affine feature and the second feature of the ground pixel in the ground image corresponding to each selected feature point.
[0090] Among them, when adjusting the parameters of the coarse affine model corresponding to the matching sampling points, the least squares (LS) algorithm, the nonlinear least squares (Gauss-Newton, GN) algorithm, the least squares optimization (Levenberg-Marquardt, LM) algorithm, etc. can be used, as long as the similarity between each affine feature and the second feature of the ground pixel point corresponding to each selected feature point can be maximized. This specification does not impose any specific restrictions.
[0091] based on Figure 1In the image registration method described in this specification, a coarse affine model can be first determined based on the relative pose of the drone and the unmanned vehicle. However, since this relative pose is a rough one, spatial sampling can be performed around the drone's initial position to obtain spatial sampling points. Based on the coarse affine model corresponding to the initial position and the relative pose of the drone and the unmanned vehicle at each sampling point, the coarse affine model corresponding to each spatial sampling point is determined. The coarse affine model corresponding to each spatial sampling point is then found to be the closest coarse affine model to the "ideal" affine model. A fine affine model is then determined based on this coarse affine model. This method can achieve registration between images collected by the drone and the unmanned vehicle. It can also achieve registration of aerial and terrestrial images without a priori maps. It is also not constrained by the environment, meaning that it can still achieve registration of aerial and terrestrial images in low-altitude urban or indoor environments.
[0092] Furthermore, in one or more embodiments of the present specification, after obtaining the precise affine model between the drone and the unmanned vehicle, in order to make the precise affine model more accurate, that is, to make the aerial image more closely match the ground image after being transformed by the precise affine model, the server can also determine, based on the precise affine model, the features of the matching image that match the features of the ground image after being transformed by the precise affine model, as features to be optimized. According to each feature to be optimized and the Virtual Line Descriptor (VLD) algorithm, for any two features to be optimized in each feature to be optimized, a geometric consistency check is performed on the two features to be optimized. For the features to be optimized that pass the test, a virtual line is generated and a descriptor is calculated. The graph structure formed by the virtual line is calculated according to the VLD algorithm, and the feature points that meet the VLD graph structure constraints are retained as inliers. Based on the inliers, the precise affine model is optimized.
[0093] Since the affine model is determined based on the matching relationship between the feature points in the matching image and the feature points in the ground image, that is, the obtained affine model is only based on the correspondence between the feature points in the two images, and different feature points have a relationship in the feature space. For example: according to the affine model, it is determined that points i and j in the matching image correspond to feature points i' and j' in the ground image respectively, and point p is determined based on i and j, and point p' is determined based on i' and j'. If the affine model is accurate, the feature points i, j, and p satisfy a certain geometric relationship, and i', j', and p' should also satisfy this geometric relationship. Therefore, the affine model can be further optimized based on the relationship between the feature points.
[0094] Specifically, two feature points X and Y can be selected in the matching image, and feature points X1 and Y1 corresponding to X and Y in the ground image can be determined based on the affine transformation model. Then, point Q can be determined in the matching image based on feature points X, Y, X1, and Y1: Determine point Q1 in the ground image: Among them, s(X) and s(X1) represent the scale information of the descriptors of feature points X and X1, a(X) and a(X1) represent the main direction information of the descriptors of feature points X and X1, and R() represents the rotation matrix corresponding to the direction information. Then we can calculate the distance D1 between X and Y, the distance D2 between X and Q, the distance D3 between Y and Q, and the distance D4 between X1 and Y1, the distance D5 between X1 and Q1, and the distance D6 between Y1 and Q1. Then, by calculating The geometric consistency index is determined to be χ(XY, X1Y1) = min(η, η1). When χ(XY, X1Y1) is less than the preset χ threshold, it can be determined that X and X1, and Y and Y1 meet the geometric consistency test, that is, the affine model is accurate. The feature points that meet the geometric consistency test can then be used as inliers, and the affine model can be adjusted based on the inliers to make the affine model more accurate.
[0095] Based on the image registration method described above, the embodiment of this specification also provides a schematic diagram of an image registration device, as shown in FIG. Figure 3 shown.
[0096] Figure 3 A schematic diagram of an image registration device provided in an embodiment of this specification, the device comprising:
[0097] An acquisition module 300 is configured to acquire an aerial image captured by a drone and a ground image captured by an unmanned vehicle, and determine the relative position of the drone and the unmanned vehicle;
[0098] a coarse affine determination module 302 for determining a coarse affine model between the aerial image and the ground image according to the relative pose, as the coarse affine model corresponding to the initial position of the drone when acquiring the aerial image;
[0099] The sampling module 304 is configured to determine, based on the initial position of the drone when collecting the aerial image, spatial sampling points around the initial position;
[0100] An expansion module 306 is configured to determine an aerial image corresponding to each spatial sampling point based on the aerial image corresponding to the initial position, and to determine a coarse affine model corresponding to each spatial sampling point based on the coarse affine model corresponding to the initial position;
[0101] An image matching module 308 is configured to determine, from among the aerial images, an aerial image that matches the ground image as a matching image;
[0102] The fine affine determination module 310 is used to determine the spatial sampling points corresponding to the matching image as matching sampling points, and obtain the fine affine model between the UAV and the unmanned vehicle based on the coarse affine model corresponding to the matching sampling points.
[0103] Optionally, the expansion module 306 is specifically configured to determine, for each spatial sampling point, a relative position between the initial position and the spatial sampling point; and determine an aerial image corresponding to the spatial sampling point based on the relative position and the aerial image corresponding to the initial position.
[0104] Optionally, the expansion module 306 is specifically used to determine, for each spatial sampling point, the relative position of the initial position and the spatial sampling point; determine the relative position of the drone and the unmanned vehicle when located at the spatial sampling point based on the relative position of the drone and the unmanned vehicle and the relative position of the initial position and the spatial sampling point; and determine the coarse affine model corresponding to the spatial sampling point based on the relative position of the drone and the unmanned vehicle when located at the spatial sampling point.
[0105] Optionally, the image matching module 308 is specifically configured to extract image features of each aerial image and image features of the ground image; and determine an aerial image that matches the ground image based on the image features of each aerial image and the image features of the ground image.
[0106] Optionally, the fine affine determination module 310 is specifically used to extract the image features of each pixel point in the matching image as each first feature; determine the number of first features matching the ground image in each first feature according to the coarse affine model corresponding to the matching sampling point; if the number is greater than a preset threshold, determine the coarse affine model corresponding to the matching sampling point as the fine affine model between the UAV and the unmanned vehicle; if the number is not greater than the preset threshold, adjust the parameters of the coarse affine model corresponding to the matching sampling point according to each first feature, and re-determine the number of first features matching the ground image in each first feature according to the adjusted coarse affine model corresponding to the matching sampling point, until the number is greater than the preset threshold.
[0107] Optionally, the fine affine determination module 310 is specifically used to extract the second feature of each pixel point in the ground image; for each pixel point in the matching image, the pixel point is used as an aerial pixel point, and the ground pixel point corresponding to the aerial pixel point in the ground image is determined according to the coarse affine model corresponding to the matching sampling point; and based on the similarity between the first feature of the aerial pixel point and the second feature of the ground pixel point corresponding to the aerial pixel point, it is determined whether the first feature of the aerial pixel point matches the ground image.
[0108] Optionally, the fine affine determination module 310 is specifically used to select multiple feature points in the matching image as selected feature points, and extract the second features of each ground pixel point in the ground image; perform affine transformation on each selected feature point according to the coarse affine model corresponding to the matching sampling point to obtain each affine feature corresponding to each selected feature point, and determine the ground pixel point in the ground image corresponding to each selected feature point; and adjust the parameters of the coarse affine model corresponding to the matching sampling point with the maximum similarity between each affine feature and the second feature of the ground pixel point in the ground image corresponding to each selected feature point as the optimization goal.
[0109] The embodiments of this specification further provide a computer-readable storage medium, which stores a computer program. The computer program can be used to execute the image registration method described above.
[0110] Based on the image registration method described above, this specification also proposes Figure 4 The schematic structure diagram of the electronic device shown in FIG. Figure 4 At the hardware level, the electronic device includes a processor, an internal bus, a network interface, memory, and non-volatile storage, and may also include other hardware required for its operations. The processor reads the corresponding computer program from the non-volatile storage into the internal memory and then runs it to implement the image registration method described above.
[0111] Of course, in addition to software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0112] In the 1990s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures such as diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements today can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to hire a chip manufacturer to design and produce a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly done using "logic compiler" software. This is similar to the software compiler used when developing programs. Before compilation, the original code must also be written in a specific programming language, called a hardware description language (HDL). There is not just one HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that by simply programming the method flow in one of these hardware description languages and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.
[0113] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that in addition to implementing the controller in a purely computer-readable program code format, the controller can be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be considered as structures within the hardware component. Or even, the devices for implementing various functions can be considered as both software modules that implement the method and structures within the hardware component.
[0114] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0115] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0116] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0117] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0118] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0119] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0120] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0121] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0122] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0123] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0124] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Thus, this specification may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0125] This specification may be described in the general context of computer-executable instructions, such as program modules, executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communications network. In a distributed computing environment, program modules may be located in both local and remote computer storage media, including storage devices.
[0126] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.
[0127] The foregoing is merely an embodiment of the present invention and is not intended to limit the present invention. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention are intended to be included within the scope of the claims of this application.
Claims
1. An image registration method, characterized in that: The method comprises: Acquire aerial images collected by a drone and ground images collected by an unmanned vehicle, and determine the relative positions of the drone and the unmanned vehicle; Determining a coarse affine model between the aerial image and the ground image according to the relative pose as the coarse affine model corresponding to the initial position of the drone when acquiring the aerial image; Determining spatial sampling points around an initial position of the drone when the drone collects the aerial image; Determining an aerial image corresponding to each spatial sampling point based on the aerial image corresponding to the initial position, and determining a coarse affine model corresponding to each spatial sampling point based on the coarse affine model corresponding to the initial position; Determining, among the aerial images, an aerial image that matches the ground image as a matching image; A spatial sampling point corresponding to the matching image is determined as a matching sampling point, and a fine affine model between the UAV and the unmanned vehicle is obtained based on a coarse affine model corresponding to the matching sampling point.
2. The method according to claim 1, wherein Determining, based on the aerial image corresponding to the initial position, the aerial image corresponding to each spatial sampling point, specifically includes: For each spatial sampling point, determining the relative position between the initial position and the spatial sampling point; The aerial image corresponding to the spatial sampling point is determined according to the relative position and the aerial image corresponding to the initial position.
3. The method according to claim 1, wherein Determining the coarse affine model corresponding to each spatial sampling point according to the coarse affine model corresponding to the initial position specifically includes: For each spatial sampling point, determining the relative position between the initial position and the spatial sampling point; Determine the relative position of the drone and the unmanned vehicle when the drone is located at the spatial sampling point according to the relative position of the drone and the unmanned vehicle and the relative position of the initial position and the spatial sampling point; According to the relative posture of the UAV and the unmanned vehicle when the UAV is located at the spatial sampling point, a coarse affine model corresponding to the spatial sampling point is determined.
4. The method according to claim 1, wherein Determining an aerial image that matches the ground image specifically includes: extracting image features of each aerial image and extracting image features of the ground image; An aerial image matching the ground image is determined based on the image features of each aerial image and the image features of the ground image.
5. The method according to claim 1, wherein Obtaining a fine affine model between the UAV and the unmanned vehicle based on the coarse affine model corresponding to the matching sampling points, specifically including: Extracting image features of each pixel in the matching image as first features; determining, according to the coarse affine model corresponding to the matching sampling points, the number of first features among the first features that match the ground image; If the number is greater than a preset threshold, the coarse affine model corresponding to the matching sampling point is determined as the fine affine model between the UAV and the unmanned vehicle; If the number is not greater than a preset threshold, the parameters of the coarse affine model corresponding to the matching sampling points are adjusted according to the first features, and the number of first features among the first features that match the ground image is re-determined according to the adjusted coarse affine model corresponding to the matching sampling points, until the number is greater than the preset threshold.
6. The method according to claim 5, wherein Determining the number of first features among the first features that match the ground image specifically includes: Extracting a second feature of each pixel in the ground image; For each pixel point in the matching image, the pixel point is used as an aerial pixel point, and a ground pixel point corresponding to the aerial pixel point in the ground image is determined based on a coarse affine model corresponding to the matching sampling point; According to the similarity between the first feature of the aerial pixel point and the second feature of the ground pixel point corresponding to the aerial pixel point, it is determined whether the first feature of the aerial pixel point matches the ground image.
7. The method according to claim 5, wherein According to the first features, the parameters of the coarse affine model corresponding to the matching sampling points are adjusted, specifically including: Selecting a plurality of feature points in the matching image as selected feature points, and extracting a second feature of each ground pixel point in the ground image; Performing an affine transformation on each selected feature point according to a coarse affine model corresponding to the matching sampling point to obtain each affine feature corresponding to each selected feature point, and determining a ground pixel point in the ground image corresponding to each selected feature point; The parameters of the coarse affine model corresponding to the matching sampling points are adjusted with the maximization of the similarity between the second features of the affine features and the ground pixel points in the ground image corresponding to the selected feature points as the optimization goal.
8. An image registration device, characterized in that: The device specifically includes: An acquisition module is used to acquire aerial images collected by the UAV and ground images collected by the unmanned vehicle, and determine the relative position of the UAV and the unmanned vehicle; a coarse affine determination module, configured to determine a coarse affine model between the aerial image and the ground image according to the relative pose, as the coarse affine model corresponding to the initial position of the drone when acquiring the aerial image; A sampling module, configured to determine, based on an initial position at which the drone collects the aerial image, spatial sampling points around the initial position; an expansion module, configured to determine an aerial image corresponding to each spatial sampling point based on the aerial image corresponding to the initial position, and to determine a coarse affine model corresponding to each spatial sampling point based on the coarse affine model corresponding to the initial position; An image matching module is used to determine, among the aerial images, an aerial image that matches the ground image as a matching image; The fine affine determination module is used to determine the spatial sampling points corresponding to the matching image as matching sampling points, and obtain the fine affine model between the UAV and the unmanned vehicle based on the coarse affine model corresponding to the matching sampling points.
9. The device according to claim 8, wherein The expansion module is specifically configured to determine, for each spatial sampling point, a relative position between the initial position and the spatial sampling point; and determine an aerial image corresponding to the spatial sampling point based on the relative position and the aerial image corresponding to the initial position.
10. The device according to claim 8, wherein The expansion module is specifically configured to, for each spatial sampling point, determine the relative position of the initial position and the spatial sampling point; determine the relative position of the drone and the unmanned vehicle when located at the spatial sampling point based on the relative position of the drone and the unmanned vehicle and the relative position of the initial position and the spatial sampling point; and determine a coarse affine model corresponding to the spatial sampling point based on the relative position of the drone and the unmanned vehicle when located at the spatial sampling point.
11. The device according to claim 8, wherein The image matching module is specifically used to extract image features of each aerial image and image features of the ground image; and determine an aerial image that matches the ground image based on the image features of each aerial image and the image features of the ground image.
12. The device according to claim 8, wherein The fine affine determination module is specifically used to extract the image features of each pixel point in the matching image as each first feature; determine the number of first features matching the ground image in each first feature based on the coarse affine model corresponding to the matching sampling point; if the number is greater than a preset threshold, determine the coarse affine model corresponding to the matching sampling point as the fine affine model between the UAV and the unmanned vehicle; if the number is not greater than the preset threshold, adjust the parameters of the coarse affine model corresponding to the matching sampling point based on each first feature, and re-determine the number of first features matching the ground image in each first feature based on the adjusted coarse affine model corresponding to the matching sampling point, until the number is greater than the preset threshold.
13. The device according to claim 12, wherein The fine affine determination module is specifically used to extract the second feature of each pixel point in the ground image; for each pixel point in the matching image, the pixel point is used as an aerial pixel point, and the ground pixel point corresponding to the aerial pixel point in the ground image is determined based on the coarse affine model corresponding to the matching sampling point; based on the similarity between the first feature of the aerial pixel point and the second feature of the ground pixel point corresponding to the aerial pixel point, it is determined whether the first feature of the aerial pixel point matches the ground image.
14. The device according to claim 12, wherein The fine affine determination module is specifically used to select multiple feature points in the matching image as selected feature points, and extract the second feature of each ground pixel point in the ground image; perform affine transformation on each selected feature point according to the coarse affine model corresponding to the matching sampling point to obtain each affine feature corresponding to each selected feature point, and determine the ground pixel point in the ground image corresponding to each selected feature point; and adjust the parameters of the coarse affine model corresponding to the matching sampling point with the optimization goal of maximizing the similarity between each affine feature and the second feature of the ground pixel point in the ground image corresponding to each selected feature point.
15. A computer-readable storage medium, characterized in that The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
16. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Method and device for positioning unmanned aerial vehicle
CN106643664A
Image registration method and terminal
WO2017107700A1