Methods, devices, equipment, and storage media for restoring camera shooting pose.
By recovering the camera pose of lost image frames using image feature matching and RTK-assisted methods, the problem of holes in 3D maps caused by image loss was solved, achieving completeness and accuracy of 3D reconstruction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-30
- Publication Date
- 2026-03-13
AI Technical Summary
In existing technologies, image loss leads to inaccurate camera shooting pose, resulting in holes in the reconstructed 3D map and preventing a complete representation of the real scene.
By detecting lost image frames and matching them with reference image frames, the translation direction and rotation parameters are determined. Combined with RTK and the translation distance of historically tracked successful image frames, the camera shooting pose of the lost image frames is restored.
The camera shooting pose of the lost images was accurately restored, avoiding image loss, ensuring the integrity of the 3D map, and optimizing the composition effect of the 3D reconstruction.
Smart Images

Figure CN115311339B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of three-dimensional reconstruction technology, and in particular to a method, apparatus, device and storage medium for restoring camera shooting pose. Background Technology
[0002] Real-time 3D reconstruction refers to the process of simultaneously capturing images of a surveyed area using a camera and constructing a 3D map of that area based on those images. Since the camera's pose during image capture is crucial data for real-time 3D reconstruction, the image is tracked after acquisition to determine the camera's pose. Because the camera is in a stable shooting state, the pose between consecutive frames does not change abruptly. Therefore, when the pose of an image differs significantly from the previous frame, the camera's pose becomes inaccurate, leading to tracking loss.
[0003] In existing technologies, for tracking lost images, relocalization methods based on the bag-of-words model recover the camera pose of the image; if relocalization fails, the image is discarded. However, discarding some images results in holes in the constructed 3D map, failing to fully represent the real scene and leading to poor 3D reconstruction composition. Summary of the Invention
[0004] This application provides a method, apparatus, device, and storage medium for restoring camera shooting pose, which solves the problem of holes in 3D maps caused by image loss in the prior art. It accurately restores and tracks the camera shooting pose of lost images, avoids image loss, ensures the integrity of 3D maps, and thus optimizes the composition effect of 3D reconstruction.
[0005] Firstly, this application provides a method for restoring the pose of a camera shot, comprising:
[0006] If a pose recovery event is detected, the lost image frame and the reference image frame required for recovery are determined.
[0007] Image feature matching is performed based on the reference image frame and the lost image frame to obtain a first feature matching pair, and a first translation direction and a first rotation parameter are determined based on the first feature matching pair;
[0008] Determine the reference translation distance between the lost image frame and the reference image frame from the successfully tracked image frames;
[0009] The first translation parameter is determined based on the reference translation distance and the first translation direction;
[0010] The camera pose of the lost image frame is recovered based on the first rotation parameter and the first translation parameter.
[0011] Secondly, this application provides a camera shooting pose recovery device, comprising:
[0012] The lost image determination module is configured to determine the lost image frame and the reference image frame required for recovery if a pose recovery event is detected.
[0013] The first rotation parameter determination module is configured to perform image feature matching based on the reference image frame and the lost image frame to obtain a first feature matching pair, and determine a first translation direction and a first rotation parameter based on the first feature matching pair.
[0014] The reference translation distance determination module is configured to determine the reference translation distance between the lost image frame and the reference image frame from the successfully tracked image frames;
[0015] The first translation parameter determination module is configured to determine the first translation parameter based on the reference translation distance and the first translation direction;
[0016] The pose recovery module is configured to recover the camera shooting pose of the lost image frame based on the first rotation parameter and the first translation parameter.
[0017] Thirdly, this application provides a camera shooting pose recovery device, comprising:
[0018] One or more processors; a storage device storing one or more programs that, when executed by the one or more processors, cause the one or more processors to implement the camera pose recovery method as described in the first aspect.
[0019] Fourthly, this application provides a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the camera pose recovery method as described in the first aspect.
[0020] This application determines feature matching pairs between lost and reference image frames by performing image feature matching on the lost and reference frames. Based on these feature matching pairs, it determines the rotation parameters and translation direction between the camera poses of the lost and reference image frames. The translation distance between the camera poses of the lost and reference image frames is determined based on the translation distance between the camera poses of two successfully tracked adjacent image frames. Furthermore, the translation parameters between the camera poses of the lost and reference image frames are determined based on the translation distance and translation direction. Finally, the camera pose of the lost image frame is determined based on the translation and rotation parameters between the camera poses of the lost and reference image frames, as well as the camera pose of the reference image frame. Through these techniques, the pose of the lost image is accurately recovered, avoiding image loss and solving the problem of incomplete 3D map reconstruction due to image loss in existing technologies. This ensures the integrity of the 3D map and optimizes the composition effect of the 3D reconstruction. Attached Figure Description
[0021] Figure 1 This is a flowchart of a camera shooting pose recovery method provided in an embodiment of this application;
[0022] Figure 2 This is a schematic diagram of image frames captured by a drone along its flight path, provided in an embodiment of this application.
[0023] Figure 3 This is a flowchart illustrating the determination of a detected pose recovery event provided in an embodiment of this application;
[0024] Figure 4 This is a flowchart provided in an embodiment of the present application for determining whether the first image frame meets the tracking success conditions;
[0025] Figure 5 This is a flowchart of two-dimensional feature matching of a first image frame and a second image provided in an embodiment of this application;
[0026] Figure 6 This is a flowchart of determining the reference translation distance provided in the embodiments of this application;
[0027] Figure 7 This is a flowchart of determining the first translation distance provided in an embodiment of this application;
[0028] Figure 8 This is a flowchart of determining the first translation distance based on the difference parameters provided in an embodiment of this application;
[0029] Figure 9 This is a flowchart illustrating the camera shooting pose for recovering lost image frames according to an embodiment of this application;
[0030] Figure 10 This is a flowchart illustrating the determination of the constraint image frame according to an embodiment of this application;
[0031] Figure 11 This is a schematic diagram of the structure of a camera shooting pose recovery device provided in an embodiment of this application;
[0032] Figure 12 This is a schematic diagram of the structure of a camera shooting pose recovery device provided in an embodiment of this application. Detailed Implementation
[0033] To make the objectives, technical solutions, and advantages of this application clearer, specific embodiments of this application will be described in further detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are merely for explaining this application and not for limiting it. It should also be noted that, for ease of description, only the parts relevant to this application are shown in the drawings, not all of them. Before discussing exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe operations (or steps) as sequential processes, many of these operations can be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the operations can be rearranged. A process can be terminated when its operation is completed, but it may also have additional steps not included in the drawings. A process can correspond to a method, function, procedure, subroutine, subroutine, etc.
[0034] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0035] The camera pose recovery method provided in this embodiment can be executed by a camera pose recovery device. This device can be implemented through software and / or hardware, and can consist of two or more physical entities, or a single physical entity. For example, the camera pose recovery device can be an intelligent device equipped with a camera, such as an unmanned device, or it can be the processor of an intelligent device. Here, unmanned devices refer to devices such as drones that can automatically execute tasks based on preset parameters.
[0036] The camera pose recovery device is equipped with at least one type of operating system. Based on this operating system, the device can install at least one application. This application can be a built-in application of the operating system or an application downloaded from a third-party device or server. In this embodiment, the camera pose recovery device has at least one application capable of executing the camera pose recovery method; therefore, the device itself can also be the application.
[0037] For ease of understanding, this embodiment uses a drone as the main example to describe the method for restoring the pose of a camera shot.
[0038] In one embodiment, the UAV flies along a pre-planned flight path and controls the camera to capture images of the surveyed area at a preset shooting step size based on the heading overlap of the flight path. During flight, the UAV can construct a 3D map of the surveyed area in real time based on the images captured by the camera and the camera pose of the images. Since the UAV flies at a constant speed to capture images of the surveyed area, the camera pose deviation between two adjacent image frames will not be too large. Therefore, if the difference between the camera pose of the current image frame and the previous image frame is small, it can be determined that the camera pose of the current image frame is accurate, and thus the tracking of the current image frame is successful. Conversely, if the difference between the camera pose of the current image frame and the previous image frame is large, it may be due to an inaccurate prediction of the camera pose of the current image frame, resulting in pose deviation. In this case, it can be determined that the current image frame has experienced tracking loss. In the prior art, for images with tracking loss, a relocalization method based on the Bag-of-Words model (BOW) is used to re-determine the lost tracking image. If relocation fails, the lost image will be discarded, resulting in the loss of some image information in the surveyed area. This is especially true in 3D reconstruction scenarios of surveyed areas with complex terrain, such as farmland, or in 3D reconstruction scenarios where flight paths have low overlap. The lost image information can cause holes in the constructed 3D map, making it impossible to fully represent the real scene of the surveyed area, affecting the composition and even causing map reconstruction failure.
[0039] To address the aforementioned issues, this embodiment provides a method for restoring camera shooting pose, so as to accurately restore the camera shooting pose for tracking lost images.
[0040] Figure 1 A flowchart of a camera pose recovery method provided in an embodiment of this application is given.
[0041] refer to Figure 1The method for restoring the camera's shooting pose specifically includes:
[0042] S110. If a pose recovery event is detected, determine the lost image frame and the reference image frame required for recovery.
[0043] Here, a lost image frame refers to a lost image frame being tracked, and a reference image frame refers to a neighboring image frame that has been successfully tracked. In this embodiment, a neighboring image frame of a certain image frame can be a forward neighboring frame or a lateral neighboring frame. A forward neighboring frame can be regarded as the image frame that is closest to the image frame on the same route, and a lateral neighboring frame can be regarded as the image frame that is closest to the image frame on the lateral route of the route where the image frame is located. Figure 2 This is a schematic diagram of image frames captured by a drone along its flight path, as provided in an embodiment of this application. Figure 2 As shown, for image frame C, image frames E and F are lateral adjacent frames of image frame C, and image frames D and B are forward adjacent frames of image frame C.
[0044] The pose recovery event refers to the event that triggers the UAV to recover the pose of a lost image frame by capturing it with the camera. In one embodiment, a pose recovery event is determined to have been detected when a lost image frame is generated. (See reference...) Figure 2 Image frame B is the image following image frame A. When image frame A is successfully tracked, the camera pose for image frame B is determined based on image frame A. If image frame B fails to track, a pose recovery event is detected, and image frame B is determined to be a lost image frame, while image frame A is determined to be the reference image frame. In this case, the reference image frame is the image frame that was successfully tracked before the lost image frame.
[0045] In another embodiment, the next image frame after successful tracking of the lost image frame is used as the reference image frame. For example, Figure 3 This is a flowchart illustrating the determination of a detected pose recovery event, provided in an embodiment of this application. For example... Figure 3 As shown, the steps for determining that a pose recovery event has been detected specifically include S1101-S1103:
[0046] S1101. Determine whether the first image frame meets the preset tracking success condition. The first image frame is the image frame currently acquired by the camera.
[0047] The tracking success condition refers to the conditions that must be met for an image frame to be successfully tracked. When an image frame meets the tracking success condition, the image frame tracking is successful; when an image frame does not meet the tracking success condition, the image frame tracking is lost.
[0048] In one embodiment, Figure 4 This is a flowchart provided in an embodiment of this application for determining whether the first image frame meets the tracking success conditions. For example... Figure 4As shown, the step of determining whether the first image frame meets the tracking success condition specifically includes S11011-S11017:
[0049] S11011. Perform feature matching between the two-dimensional feature points of the first image frame and the three-dimensional feature points constructed from the second image frame to obtain a second feature matching pair; the second image frame is the previous successfully tracked image frame.
[0050] In this embodiment, a 3D-2D data association method is used to determine the matching feature points between the first image frame and the second image frame. For example, feature points are extracted from the first image frame to obtain two-dimensional feature points. Feature points are extracted from the second and third image frames, and feature matching is performed based on the extracted feature points to obtain feature matching pairs. The three-dimensional feature points of the second image frame are then constructed using triangulation based on these feature matching pairs. The third image frame is the previous image frame that was successfully tracked before the second image frame. The three-dimensional feature points of the second image frame are projected onto the first image frame and matched with the two-dimensional feature points of the first image frame to obtain the second feature matching pair.
[0051] S11012. Compare the number of second feature matching pairs with a preset number threshold.
[0052] For example, three-dimensional feature points are feature points that coexist in the second and third image frames, and two-dimensional feature points in the first image frame that match the three-dimensional feature points also exist in the second and third image frames. Since there is a high degree of overlap between the first, second, and third image frames, the two-dimensional feature points in the overlapping area are observable; these two-dimensional feature points constitute the second feature matching pairs. If the number of second feature matching pairs is small, the correlation between the first image frame and the first two image frames is low. However, since the camera pose does not change abruptly during shooting (i.e., the correlation between the three image frames is not too low), the small number of second feature matching pairs may be due to feature point matching errors or feature point extraction errors. In this case, the second feature matching pairs cannot accurately represent the camera shooting pose deviation between the first and second image frames, and therefore, the camera shooting pose of the first image frame cannot be accurately determined based on the second feature matching pairs. If the number of second feature matching pairs is large, the second feature matching pairs can accurately represent the camera shooting pose deviation between the first and second image frames, and thus, the camera shooting pose of the first image frame can be determined based on the second feature matching pairs.
[0053] The preset threshold refers to the number of second feature matching pairs when the correlation between three adjacent image frames is high. If the number of second feature matching pairs is greater than or equal to the preset threshold, it indicates that there is a high correlation between the observed first image frame and the second and third image frames, and subsequent pose tracking operations can be performed based on the second feature matching pairs. If the number of second feature matching pairs is less than the preset threshold, it indicates that the correlation between the first image frame and the second and third image frames is low, and the camera pose of the first image frame cannot be accurately determined based on the second feature matching pairs. For example, assuming that 2000 two-dimensional feature points in the second image frame can construct 300 three-dimensional feature points, and that 300 three-dimensional feature points projected onto the first image frame can generate 100 second feature matching pairs, if the preset threshold is 80, then the feature points between the first image frame and the second and third image frames are successfully associated.
[0054] S11013. In response to the comparison result that the number of second feature matching pairs is less than a preset number threshold, determine that the first image frame does not meet the tracking success condition.
[0055] For example, if the number of second feature matching pairs is less than a preset threshold, the camera shooting pose of the first image frame cannot be accurately determined based on the second feature matching pairs, and the tracking of the first image frame is determined to be lost.
[0056] S11014. In response to the comparison result that the number of second feature matching pairs is greater than or equal to a preset number threshold, determine the relative transformation parameter between the first image frame and the second image frame based on the second feature matching pairs.
[0057] For example, when the number of second feature matching pairs is greater than or equal to a preset threshold, PNP (Perspective-N-Points) motion estimation is performed based on the second feature matching pairs. Assuming the relative transformation parameter between the camera pose of the first image frame and the camera pose of the second image frame is (R, t), where R represents the rotation parameter and t represents the displacement parameter, the three-dimensional feature point in the second feature matching pair is V, the two-dimensional feature point is v, and the camera intrinsic parameter is K, then the cost function is constructed as follows:
[0058] J(R,t)=min‖vK(RV+t)‖ 2
[0059] By optimizing the cost function J(R,t) using a nonlinear optimization algorithm, the relative transformation parameter (R,t) can be determined.
[0060] S11015. Compare the relative transformation parameter with the preset deviation threshold.
[0061] The preset deviation threshold refers to the maximum difference in camera pose between two adjacent image frames when the UAV flies at a constant speed along the flight path. The relative transformation parameter is the pose deviation between the first and second image frames. For example, since the camera pose between two adjacent image frames does not change abruptly when the UAV flies at a constant speed along the flight path, if the relative transformation parameter is greater than the preset deviation threshold, it indicates that the relative transformation parameter estimated by the current motion is inaccurate. Conversely, if the relative transformation parameter is less than or equal to the preset deviation threshold, it indicates that the relative transformation parameter estimated by the current motion is accurate, and the camera pose of the first image frame determined based on the relative transformation parameter is also accurate.
[0062] S11016. In response to the comparison result that the relative transformation parameter is less than or equal to the preset deviation threshold, determine that the first image frame meets the tracking success condition.
[0063] For example, when the relative transformation parameter is accurate, the camera pose of the first image frame determined based on the relative transformation parameter is also accurate, thus confirming successful tracking of the first image frame. The camera pose of the first image frame can be obtained by multiplying the camera pose of the second image frame by the transformation matrix corresponding to the relative transformation parameter.
[0064] S11017. In response to the comparison result that the relative transformation parameter is greater than the preset deviation threshold, it is determined that the first image frame does not meet the tracking success condition.
[0065] For example, when the relative transformation parameter is inaccurate, the camera shooting pose of the first image frame determined based on the relative transformation parameter is also inaccurate, and it can be determined that the tracking of the first image frame is lost.
[0066] S1102. In response to the judgment result that the first image frame meets the tracking success condition, determine whether there is an image frame stored in the lost frame container. The image frame in the lost frame container is added when it is determined that the tracking success condition is not met during tracking based on the first image frame.
[0067] For example, a lost frame container is used to store and track lost image frames. Image frames in the lost frame container are not used to determine the camera pose for the next image frame. (See reference) Figure 2When the drone flies along the flight path, the shooting sequence is shown by the arrows. Dashed boxes indicate lost image frames, and solid boxes indicate successfully tracked image frames. This embodiment describes the tracking process from image frame A to image frame D. Image frame A is tracked. After successful tracking of image frame A, the lost frame container is checked for any image frames. If no image frame is found in the lost frame container, tracking of image frame B continues based on image frame A. If tracking of image frame B fails, image frame B is stored in the lost frame container, and tracking of image frame C continues based on image frame A. If tracking of image frame C fails, image frame C is stored in the lost frame container, and tracking of image frame D continues based on image frame A. After successful tracking of image frame D, the lost frame container is checked to see if image frame C and image frame B are stored. Image frame C is the lost image frame, and image frame D is the next image frame after successful tracking of image frame C.
[0068] S1103. In response to the judgment result that the lost frame container stores an image frame, determine that a pose recovery event has been detected.
[0069] In this embodiment, a pose recovery event is detected when the lost frame container stores image frames. Since the last image frame added to the lost frame container is the image frame preceding the most recently tracked image frame, it can be determined that the last image frame added to the lost frame container is the lost image frame, and the first image frame is determined as the reference image frame. Figure 2 Image frame C is determined to be the lost image frame, and image frame D is determined to be the reference image frame.
[0070] It should be noted that if the reference image frame is the next image frame after the successfully tracked lost image frame, then this reference image frame is determined based on the successfully tracked image frames preceding the lost image frame. For example, image frame D is determined based on image frame A, and the camera pose of image frame A constrains the camera pose of image frame D. Therefore, when subsequently recovering the camera pose of the lost image frame based on the reference image frame, the camera poses of image frames D and A form a forward-backward constraint on the camera pose of image frame C. This forward-backward constraint can improve the accuracy of recovering the camera pose of image frame C.
[0071] S120. Perform image feature matching based on the reference image frame and the lost image frame to obtain a first feature matching pair, and determine the first translation direction and the first rotation parameter based on the first feature matching pair.
[0072] In one embodiment, feature matching is performed between the 3D feature points constructed from the reference image frame and the 2D feature points of the lost image frame to obtain feature matching pairs. Based on these feature matching pairs, PNP motion estimation is performed to determine the relative transformation matrix between the lost image frame and the reference image frame. However, tracking failure in the lost image frame is caused by feature matching errors or motion estimation errors. Therefore, tracking failure may occur again using 3D-2D data association methods and PNP motion estimation methods.
[0073] To address this, this embodiment proposes a 2D-2D data association method for feature matching. The success rate of 2D-2D matching between adjacent image frames is much higher than that of 3D-2D matching. For example, if 2000 two-dimensional feature points are extracted from each image frame, the number of successfully matched feature pairs in 2D-2D matching is approximately 500. For instance, Figure 5 This is a flowchart illustrating two-dimensional feature matching of a first image frame and a second image, provided in an embodiment of this application. For example... Figure 5 As shown, the step of performing two-dimensional feature matching on the first image frame and the second image specifically includes S1201-S1202:
[0074] S1201. Match the two-dimensional feature points of the reference image frame with the two-dimensional feature points of the lost image frame to obtain the first feature matching pair.
[0075] S1202, filter out first feature matching pairs that do not meet geometric constraints through random consistency.
[0076] This embodiment uses image frame D and image frame C as examples to describe the reference image frame and the lost image frame, respectively. Two-dimensional feature points are extracted from image frame D and image frame C respectively. The two-dimensional feature points of image frame D are matched with the two-dimensional feature points of image frame C to obtain the first feature matching pair between image frame D and image frame C. The first feature matching pair that does not meet the geometric constraints is filtered out using the Random Consensus Algorithm (RANSAC) to recover the camera pose of image frame C using the remaining first feature matching pairs. It can be understood that if two two-dimensional feature points in the first feature matching pair do not meet the geometric constraints, it indicates that there is a matching error between the two feature points. Therefore, the first feature matching pair that does not meet the geometric constraints is filtered out to avoid affecting the accuracy of subsequent pose recovery.
[0077] In this embodiment, based on the first feature matching pair, the essential matrix between the reference image frame and the lost image frame is determined using polar geometric constraints. The essential matrix is a 3*3 matrix with five degrees of freedom: three rotational degrees of freedom and three translational degrees of freedom, with the scale degree of freedom removed from the translational degree of freedom. The first relative transformation matrix between the reference image frame and the lost image frame is T1 = (R1, t1), where R1 is the first rotation parameter and t1 is the first translation parameter.
[0078] It should be noted that since the scale is unknown in the essential matrix, only the unit vector t of the first rotation parameter R1 and the first translation parameter t1 in the first relative transformation matrix T1 can be determined from the essential matrix. uv1 Therefore, based on the first feature matching pair, only the first translation direction and the first rotation parameter R1 of the first translation parameter t1 can be determined. However, the first feature matching pair cannot determine the first translation distance of the first translation parameter t1. Therefore, in this embodiment, it is also necessary to determine the first translation distance, and then determine the first translation parameter t1 based on the first translation distance and the first translation direction. The first translation distance is the distance between the camera shooting pose of the reference image frame and the lost image frame.
[0079] S130. Determine the reference translation distance between the lost image frame and the reference image frame from the successfully tracked image frames.
[0080] In existing technologies, drones are equipped with positioning sensors such as RTK (Real-Time Kinematics). RTK positioning accuracy is at the centimeter level. Based on the world coordinates of the lost image frame and the reference image frame acquired by RTK, the coordinate distance between the lost image frame and the reference image frame can be determined. This coordinate distance is then used as a reference value for the first translation distance measured by RTK in world coordinates. However, in some cases, RTK positioning is not very accurate, and relying entirely on RTK may introduce larger errors.
[0081] To address this, this embodiment proposes determining the translational distance between an image frame and its adjacent frames based on the camera's shooting pose of historically successfully tracked image frames. This translational distance is then used as a reference value for the first translational distance measured by the camera in camera coordinates, thereby improving the accuracy of the first translational distance estimation. The reference translational distance is precisely the reference value of the first translational distance measured by the camera in camera coordinates.
[0082] It is understandable that when the reference image frame and the lost image frame are on the same flight path, the reference image frame is the flight path adjacent frame of the lost image frame. Since the UAV takes images at regular intervals while traveling at a constant speed along a certain flight path, this ensures that two adjacent image frames on the same flight path satisfy the flight path overlap requirement. Therefore, the translational distance between two adjacent image frames on the same flight path that were successfully tracked in the past is approximately the first translational distance between the reference image frame and the lost image frame. This translational distance can be used as a reference value for the first translational distance. It should be noted that two adjacent image frames on the same flight path are not necessarily on the same flight path as the lost image frame.
[0083] When the reference image frame and the lost image frame are not on the same flight path, the reference image frame is the lateral neighbor of the lost image frame. The distance between adjacent flight paths is determined based on the lateral overlap, meaning that the distance between adjacent flight paths is the same, and the distance between adjacent flight paths is approximately equal to the translation distance between the image frame and the lateral neighbor. Therefore, the translation distance between a historically successfully tracked image frame and its lateral neighbor is approximately the first translation distance between the reference image frame and the lost image frame, and the translation distance between a historically successfully tracked image frame and its lateral neighbor can be used as a reference value for the first translation distance.
[0084] In summary, when the lost image frame and the reference image frame are laterally adjacent frames, the reference translation distance can be determined based on historically successfully tracked laterally adjacent image pairs; when the lost image frame and the reference image frame are directionally adjacent frames, the reference translation distance can be determined based on historically successfully tracked directionally adjacent image pairs. Specifically, a laterally adjacent image pair consists of an image frame and its corresponding laterally adjacent frame, and a directionally adjacent image pair consists of an image frame and its corresponding directionally adjacent frame.
[0085] This embodiment describes an example where the reference image frame and the lost image frame are on the same flight path. In this embodiment, the translation distance between two corresponding image frames is determined based on the translation parameters of adjacent image pairs, and a reference translation distance is determined based on the translation distance. An adjacent image pair consists of two adjacent image frames from the same flight path that were successfully tracked in the past. For example, the image frame IDs are generated according to their shooting order, meaning the IDs of adjacent image frames are consecutive. The image frames that were successfully tracked in the past are traversed, and an image sequence for the flight path is generated by sorting the image frames based on their IDs. Two image frames with adjacent IDs in this image sequence are grouped into an adjacent image pair, resulting in the adjacent image pairs pairs1 to pairs from the first flight path to the flight path where the lost image frame is located. n pairs j ={(I j1 I j2 ), (I j2 I j3 ), ..., (I j(i-1) I ji ), ..., (I j(n-1) I jn )},I ji For the i-th image frame in the image sequence that was successfully tracked on the j-th route, I nn When the lost image frame is the image frame that was successfully tracked before the image frame on the same flight path as the lost image frame, i.e., when the lost image frame is image frame C, I nn Image frame A. Based on image frame I. i and image frame I i-1 The relative transformation parameter T between them (i-1)i =(R (i-1)i , t(i-1)i Determine the translation parameter t between two image frames. (i-1)i Therefore, the translation distance between the two image frames is determined to be ||t. (i-1)i || 2 After determining the translation distance of all adjacent image pairs, the mean or median of the translation distance can be used as the reference translation distance.
[0086] In another embodiment, since calculating the translation distance of all adjacent image pairs is computationally too expensive, it can be done in pairs1 to pairs2. n At least five pairs of adjacent images are selected to calculate the translation distance, and then a reference translation distance is determined based on the median or mean of the calculated translation distances. For example, Figure 6 This is a flowchart illustrating the determination of a reference translation distance provided in an embodiment of this application. For example... Figure 6 As shown, the steps for determining the reference translation distance specifically include S1301-S1302:
[0087] S1301. Select at least two pairs of first adjacent images that are closest to the lost image frame, and randomly select at least three pairs of second adjacent images from the remaining pairs of adjacent images.
[0088] For example, the closer the adjacent image pair is to the lost image frame, the more closely the translation distance approximates the first translation distance, meaning the higher the confidence level of that adjacent image pair. Therefore, from pairs n Select (I) n(n-1) I nn ) and (I n(n-2) I n(n-1) () as the first adjacent image pair. From pair1 to pair2 n Three pairs of the remaining adjacent image pairs are randomly selected as the second adjacent image pairs.
[0089] S1302. Determine the reference translation distance based on the translation distance of the first adjacent image pair and the second adjacent image pair.
[0090] For example, the translation distance between the first adjacent image pair and the second adjacent image pair is calculated, and the median of the translation distance is taken as the reference translation distance.
[0091] S140. The first translation parameter is determined based on the reference translation distance and the first translation direction.
[0092] For example, the first translation parameter is determined by using the reference translation distance as the final first translation distance, such that the scale of the first relative transformation matrix is consistent with the scale of the adjacent image pair.
[0093] In another embodiment, the UAV aligns the camera and RTK scales during the shooting process. This involves constraining the camera's shooting pose using world coordinates acquired by the RTK to improve the accuracy of the camera's shooting pose in image frames. It should be noted that in monocular visual SLAM, when there are small errors in the reference translation distance calculated based on adjacent image pairs, these scale errors will be amplified after accumulating multiple image frames, thus reducing the accuracy of the reference translation distance calculated based on adjacent image pairs. Therefore, after the camera and RTK scales are aligned, the reference value of the first translation distance measured by the RTK is closer to the true value of the first translation distance than the reference value measured by the camera. Therefore, the reference value of the first translation distance measured by the RTK is used as the final first translation distance to calculate the first translation parameter. However, the camera and RTK scale alignment is performed only after at least two image frames along the flight path have been captured. Therefore, the first translation distance can be determined after determining whether the camera and RTK scales are aligned. For example, Figure 7 This is a flowchart illustrating the determination of the first translation distance provided in an embodiment of this application. For example... Figure 7 As shown, the step of determining the first translation distance specifically includes S1401-S1402:
[0094] S1401. Determine the coordinate translation distance of the positioning sensor based on the world coordinates of the lost image frame and the reference image frame acquired by the positioning sensor.
[0095] Here, the coordinate translation distance is the reference value of the first translation distance in the world coordinate system measured by RTK. Assume the world coordinates of the lost image frame are RTK... lost The world coordinates of the reference image frame are RTK. ref Then determine the coordinate translation distance s rtk =||RTK ref -RTK lost || 2 .
[0096] S1402. Determine the first translation distance based on the difference parameter between the reference translation distance and the coordinate translation distance, and determine the first translation parameter based on the first translation distance and the first translation direction.
[0097] The difference parameter can be the difference or ratio between the reference translation distance and the coordinate translation distance. This embodiment uses the ratio of the coordinate translation distance to the reference translation distance as an example. For instance, Figure 8 This is a flowchart illustrating the determination of the first translation distance based on difference parameters, provided in an embodiment of this application. For example... Figure 8 As shown, the step of determining the first translation distance based on the difference parameter specifically includes S14021-S14023:
[0098] S14021. Determine the ratio of the coordinate translation distance to the reference translation distance as the difference parameter, and compare the difference parameter with the preset difference range.
[0099] The preset difference range refers to the range of variation in the ratio when the RTK and camera are scale-aligned. The preset difference range is (1-δ, 1+δ), where δ is a very small positive number. For example, the difference parameter... Where s vis The reference translation distance is denoted as . When ratio ∈ (1-δ, 1+δ), it indicates that the difference parameters between the camera and RTK are small, meaning the UAV has already performed scale alignment between the RTK and the camera. When ratio < 1-δ or ratio > 1+δ, it indicates that the difference parameters between the camera and RTK are large, meaning the UAV has not yet performed scale alignment between the RTK and the camera.
[0100] S14022. In response to the comparison result where the difference parameter is within the preset difference range, determine the first translation distance as the coordinate translation distance.
[0101] For example, if the UAV has already undergone RTK and camera scale alignment, the coordinate translation distance is used as the final first translation distance, thereby determining the first translation parameter t1 = t uv *s rtk .
[0102] S14023. In response to the comparison result where the difference parameter exceeds the preset difference range, determine the first translation distance as the reference translation distance.
[0103] For example, if the drone has not yet performed RTK and camera scale alignment, using a reference translation distance can ensure that the scale of the first translation distance and the translation distance of adjacent image pairs are consistent. Therefore, the reference translation distance is used as the final first translation distance, and thus the first translation parameter t1 = t is determined. uv *s vis .
[0104] S150. Recover the camera shooting pose of the lost image frame based on the first rotation parameter and the first translation parameter.
[0105] For example, the camera shooting pose of the reference image frame is multiplied by the transformation matrix corresponding to the first relative transformation parameter to obtain the camera shooting pose of the lost image frame.
[0106] In this embodiment, after the camera pose of the lost image frame is successfully restored, the lost image frame is deleted from the lost frame container. It is then determined whether the lost frame container still stores image frames. If it does, the last stored image frame is used as the lost image frame, and the currently successfully restored image frame is used as the reference image frame. Figure 2After recovering the camera shooting pose of image frame C, image frame C is deleted from the lost frame container, and image frame B is used as the lost image frame, while image frame C is used as the reference image frame, so as to recover the camera shooting pose of image frame B based on image frame C.
[0107] In another embodiment, the camera pose of the lost image frame is determined based on the next image frame, and the camera pose of the lost image frame is unstable. Based on the principle of the most stable triangle structure, the triangle constraint relationship can be determined by constraint image frames that are not on the same straight line as the lost image frame and the reference image frame. A joint optimization equation can then be constructed based on the triangle constraint relationship to recover a more stable camera pose. It should be noted that when the reference image frame and the lost image frame are on the same flight path, a lateral adjacent frame is obtained from the lateral flight path of the lost image frame as a constraint image frame to construct the triangle constraint relationship; when the reference image frame and the lost image frame are not on the same flight path, a heading adjacent frame is obtained from the flight path of the lost image frame as a constraint image frame to construct the triangle constraint relationship.
[0108] This embodiment describes an example where the reference image frame and the lost image frame are on the same flight path. Figure 10 This is a flowchart illustrating the camera shooting pose for recovering lost image frames according to an embodiment of this application. Figure 10 As shown, the steps for restoring the camera shooting pose of the lost image frame specifically include S1501-S1504:
[0109] S1501. From the image frames on the side path of the lost image frame, determine the image frame closest to the lost image frame as the constraint image frame.
[0110] In one embodiment, the image frame closest to the lost image frame on the lateral route is determined as the constraint image frame based on the distance between each image frame on the lateral route and the lost image frame. For example, Figure 10 This is a flowchart illustrating the determination of a constraint image frame according to an embodiment of this application. Figure 10 As shown, the step of determining the lateral adjacent frame specifically includes S15011-S15012:
[0111] S15011. Determine the positional distance between the lost image frame and the image frame on the corresponding lateral route based on the world coordinates of the lost image frame and the world coordinates of the image frame on the corresponding lateral route.
[0112] For example, based on the world coordinates of the lost image frame acquired by RTK and the world coordinates of the image frames on the sideline route, the positional distance between the lost image frame and each image frame on the sideline route is calculated respectively.
[0113] S15012. Based on the location distance, determine the image frame on the lateral flight path that is closest to the lost image frame as the constrained image frame.
[0114] For example, the positional distance of the lost image frame is compared with that of each image frame on the lateral route, and the image frame with the smallest positional distance is determined as the constrained image frame.
[0115] S1502. Perform image feature matching on the lost image frame and the constrained image frame to obtain a third feature matching pair, and determine the second translation direction and the second rotation parameter based on the third feature matching pair.
[0116] For example, two-dimensional feature matching is performed on the lost image frame and the constrained image frame to obtain a third feature matching pair. The essential matrix between the lost image frame and the constrained image frame is determined based on the third feature matching pair. The unit vector t of the second rotation parameter R2 and the second translation parameter t2 between the lost image frame and the constrained image frame is then determined based on the essential matrix. uv2 For details, please refer to step S120.
[0117] S1503. Determine the second translation distance based on the translation distance between the image frame and the closest image frame on the lateral route, and determine the second translation parameter according to the second translation direction and the second translation distance.
[0118] For example, the distance between two adjacent flight paths in the flight path planning can be determined as the second translation distance. Alternatively, the positional distance between the lost image frame and the constrained image frame can be determined as the second translation distance.
[0119] In another embodiment, a second translation distance is determined based on the translation distance of historically successfully tracked lateral neighbor image pairs. The implementation process of this embodiment is largely the same as the process of determining the first translation distance, and specific steps S130-S140 can be referred to.
[0120] In this embodiment, based on the flight path where the historically successfully tracked image frames are located, and based on the camera shooting pose of the image frame and the camera shooting poses of each image frame on the lateral flight path, the translation distance between the image frame and each image frame on the lateral flight path is determined, and lateral adjacent image pairs are determined based on this distance. From the lateral adjacent image pairs, at least two first lateral adjacent image pairs with IDs closest to the lost image frame are selected, and at least three second lateral adjacent image pairs are randomly selected from the remaining lateral adjacent image pairs. Based on the translation distances of the first and second lateral adjacent image pairs, a reference value for the second translation distance in the camera coordinate system is determined. The positional distance between the lateral adjacent frame and the lost image frame is determined as the reference value for the second translation distance in the world coordinates measured by RTK. The final second translation distance can be determined based on the determination result of whether the previous camera and RTK are aligned.
[0121] S1504. Recover the camera shooting pose of the lost image frame based on the first rotation parameter, the second rotation parameter, the first translation parameter, and the second translation parameter.
[0122] For example, a first constraint equation is constructed based on the camera pose of the reference image frame, a first rotation parameter, and a first translation parameter. A second constraint equation is constructed based on the camera pose of the constraint image frame, a second rotation parameter, and a second translation parameter. The first and second constraint relationships are added to the window optimization equation, and the optimization equation is solved to obtain the camera pose of the lost image frame. The window optimization equation is constructed based on the association relationship between the reference image frame and image frames with which there is a correlation.
[0123] In summary, the camera pose recovery method provided in this application involves performing image feature matching between the lost image frame and the reference image frame to determine the feature matching pair between them. Based on the feature matching pair, the rotation parameters and translation direction between the camera poses of the lost and reference image frames are determined. The translation distance between the camera poses of the lost and reference image frames is determined based on the translation distance between the camera poses of two successfully tracked adjacent image frames. Furthermore, the translation parameters between the camera poses of the lost and reference image frames are determined based on the translation distance and translation direction. Finally, the camera pose of the lost image frame is determined based on the translation and rotation parameters between the camera poses of the lost and reference image frames, as well as the camera pose of the reference image frame. Through these technical means, the pose of the tracked lost image is accurately recovered, avoiding image loss and solving the problem of incomplete 3D map reconstruction due to image loss in existing technologies. This ensures the integrity of the 3D map and optimizes the composition effect of the 3D reconstruction.
[0124] Based on the above embodiments, Figure 11 This is a schematic diagram of a camera shooting pose recovery device provided in an embodiment of this application. (Reference) Figure 11 The camera pose recovery device provided in this embodiment specifically includes: a lost image determination module 21, a first rotation parameter determination module 22, a reference translation distance determination module 23, a first translation parameter determination module 24, and a pose recovery module 25.
[0125] The lost image determination module is configured to determine the lost image frame and the reference image frame required for recovery if a pose recovery event is detected.
[0126] The first rotation parameter determination module is configured to perform image feature matching based on the reference image frame and the lost image frame to obtain a first feature matching pair, and determine the first translation direction and the first rotation parameter based on the first feature matching pair.
[0127] The reference translation distance determination module is configured to determine the reference translation distance between the lost image frame and the reference image frame from the successfully tracked image frames;
[0128] The first translation parameter determination module is configured to determine the first translation parameter based on the reference translation distance and the first translation direction;
[0129] The pose recovery module is configured to recover the camera pose of the lost image frame based on a first rotation parameter and a first translation parameter.
[0130] Based on the above embodiments, when the reference image frame and the lost image frame are on the same flight path, the reference translation distance determination module includes: a reference translation distance determination submodule, configured to determine the translation distance between two corresponding image frames based on the translation parameters of adjacent image pairs, and to determine the reference translation distance based on the translation distance; the adjacent image pairs consist of two adjacent image frames on the same flight path that were successfully tracked in the past.
[0131] Based on the above embodiments, the reference translation distance determination submodule includes: an image pair selection unit, configured to select at least two sets of first adjacent image pairs closest to the lost image frame, and randomly select at least three sets of second adjacent image pairs from the remaining multiple sets of adjacent image pairs; and a reference translation distance determination unit, configured to determine a reference translation distance based on the translation distance of the first adjacent image pairs and the second adjacent image pairs.
[0132] Based on the above embodiments, the first translation parameter determination module includes: a coordinate translation distance determination submodule, configured to determine the coordinate translation distance of the positioning sensor based on the world coordinates of the lost image frame and the reference image frame acquired by the positioning sensor; a first translation distance determination submodule, configured to determine a first translation distance based on the difference parameter between the reference translation distance and the coordinate translation distance; and a first translation parameter determination submodule, configured to determine a first translation parameter based on the first translation distance and the first translation direction.
[0133] Based on the above embodiments, the first translation distance determination submodule includes: a first comparison unit configured to determine the ratio of coordinate translation distance to reference translation distance as a difference parameter, and compare the difference parameter with a preset difference range; a first determination unit configured to determine the first translation distance as a coordinate translation distance in response to a comparison result where the difference parameter is within the preset difference range; and a second determination unit configured to determine the first translation distance as a reference translation distance in response to a comparison result where the difference parameter exceeds the preset difference range.
[0134] Based on the above embodiments, the pose recovery module includes: a lateral image determination submodule, configured to determine the constraint image frame closest to the lost image frame from image frames on the lateral flight path of the lost image frame; a third parameter determination submodule, configured to perform image feature matching on the lost image frame and the constraint image frame to obtain a third feature matching pair, and determine a second translation direction and a second rotation parameter based on the third feature matching pair; a fourth parameter determination submodule, configured to determine a second translation distance based on the translation distance between the image frame and the closest image frame on the lateral flight path, and determine a second translation parameter based on the second translation direction and the second translation distance; and a pose recovery submodule, configured to recover the camera shooting pose of the lost image frame based on the first rotation parameter, the second rotation parameter, the first translation parameter, and the second translation parameter.
[0135] Based on the above embodiments, the lateral image determination submodule includes: a position distance determination unit, configured to determine the position distance between the lost image frame and the image frame on the lateral route based on the world coordinates of the lost image frame and the world coordinates of the corresponding image frame on the lateral route; and a lateral image determination unit, configured to determine the image frame on the lateral route closest to the lost image frame as the constraint image frame based on the position distance.
[0136] Based on the above embodiments, the lost image determination module includes: a tracking judgment submodule, configured to determine whether a first image frame meets a preset tracking success condition, wherein the first image frame is an image frame currently acquired by the camera; a container detection submodule, configured to, in response to the judgment result that the first image frame meets the tracking success condition, determine whether an image frame is stored in the lost frame container, wherein an image frame in the lost frame container is determined to not meet the tracking success condition when tracking is performed based on the first image frame and is added to the container; and an event detection submodule, configured to, in response to the judgment result that an image frame is stored in the lost frame container, determine that a pose recovery event has been detected.
[0137] Based on the above embodiments, the lost image determination module includes: a lost image determination submodule, configured to determine that the last image frame added to the lost frame container is a lost image frame, and to determine that the first image frame is a reference image frame.
[0138] Based on the above embodiments, the tracking judgment submodule includes: a first feature matching unit, configured to perform feature matching between two-dimensional feature points of the first image frame and three-dimensional feature points constructed from the second image frame to obtain a second feature matching pair; the second image frame is the previous successfully tracked image frame; a second comparison unit, configured to compare the number of second feature matching pairs with a preset number threshold; and a first tracking judgment unit, configured to determine that the first image frame does not meet the tracking success condition in response to the comparison result that the number of second feature matching pairs is less than the preset number threshold.
[0139] Based on the above embodiments, the tracking judgment submodule further includes: a second tracking judgment unit configured to determine the relative transformation parameter between the first image frame and the second image frame according to the second feature matching pair in response to a comparison result where the number of second feature matching pairs is greater than or equal to a preset number threshold; a third comparison unit configured to compare the relative transformation parameter with a preset deviation threshold; and a third tracking judgment unit configured to determine that the first image frame meets the tracking success condition in response to a comparison result where the relative transformation parameter is less than or equal to the preset deviation threshold.
[0140] Based on the above embodiments, the first parameter determination module includes: a second feature matching submodule, configured to match the two-dimensional feature points of the reference image frame with the two-dimensional feature points of the lost image frame to obtain a first feature matching pair; and a matching pair filtering submodule, configured to filter out the first feature matching pairs that do not meet the geometric constraints through random consistency.
[0141] The camera pose recovery device provided in this application, as described above, determines the feature matching pair between the lost image frame and the reference image frame by performing image feature matching between them, and then determines the rotation parameters and translation direction between the camera poses of the lost image frame and the reference image frame based on the feature matching pair. The translation distance between the camera poses of the lost image frame and the reference image frame is determined based on the translation distance between the camera poses of two successfully tracked adjacent image frames. Furthermore, the translation parameters between the camera poses of the lost image frame and the reference image frame are determined based on the translation distance and translation direction. The camera pose of the lost image frame can be determined based on the translation and rotation parameters between the camera poses of the lost image frame and the reference image frame, as well as the camera pose of the reference image frame. Through these technical means, the pose of the tracked lost image is accurately recovered, avoiding image loss and solving the problem of incomplete 3D map reconstruction due to image loss in the prior art, ensuring the integrity of the 3D map and thus optimizing the composition effect of the 3D reconstruction.
[0142] The camera pose recovery device provided in this application embodiment can be used to execute the camera pose recovery method provided in the above embodiment, and has corresponding functions and beneficial effects.
[0143] Figure 12 This is a schematic diagram of a camera shooting pose recovery device provided in an embodiment of this application, with reference to... Figure 12 The camera pose recovery device includes a processor 31, a memory 32, a communication device 33, an input device 34, and an output device 35. The number of processors 31 and the number of memories 32 in the camera pose recovery device can be one or more. The processor 31, memory 32, communication device 33, input device 34, and output device 35 of the camera pose recovery device can be connected via a bus or other means.
[0144] The memory 32, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as program instructions / modules corresponding to the camera shooting pose recovery method in any embodiment of this application (e.g., the lost image determination module 21, the first rotation parameter determination module 22, the reference translation distance determination module 23, the first translation parameter determination module 24, and the pose recovery module 25 in the camera shooting pose recovery device). The memory 32 may mainly include a program storage area and a data storage area, wherein the program storage area may store the operating system and at least one application program required for a function; the data storage area may store data created according to the use of the device, etc. In addition, the memory 32 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the device via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0145] The communication device 33 is used for data transmission.
[0146] The processor 31 executes various functional applications and data processing of the device by running software programs, instructions and modules stored in the memory 32, thereby realizing the above-mentioned method for restoring the camera shooting pose.
[0147] Input device 34 can be used to receive input digital or character information, and to generate key signal inputs related to user settings and function control of the device. Output device 35 may include display devices such as a display screen.
[0148] The camera pose recovery device provided above can be used to perform the camera pose recovery method provided in the above embodiments, and has corresponding functions and beneficial effects.
[0149] This application embodiment also provides a storage medium containing computer-executable instructions. When executed by a computer processor, the computer-executable instructions are used to perform a camera shooting pose recovery method. The camera shooting pose recovery method includes: if a pose recovery event is detected, determining a lost image frame and a reference image frame required for recovery; performing image feature matching based on the reference image frame and the lost image frame to obtain a first feature matching pair, and determining a first translation direction and a first rotation parameter based on the first feature matching pair; determining a reference translation distance between the lost image frame and the reference image frame from successfully tracked image frames; determining a first translation parameter based on the reference translation distance and the first translation direction; and recovering the camera shooting pose of the lost image frame based on the first rotation parameter and the first translation parameter.
[0150] Storage medium – any type of memory device or storage device. The term “storage medium” is intended to include: mounting media, such as CD-ROM, floppy disk, or magnetic tape devices; computer system memory or random access memory, such as DRAM, DDR RAM, SRAM, EDO RAM, Rambus RAM, etc.; non-volatile memory, such as flash memory, magnetic media (e.g., hard disk or optical storage); registers or other similar types of memory elements, etc. Storage medium may also include other types of memory or combinations thereof. Furthermore, storage medium may reside in a first computer system in which the program is executed, or it may reside in a different second computer system connected to the first computer system via a network (such as the Internet). The second computer system can provide program instructions to the first computer for execution. The term “storage medium” can include two or more storage media residing in different locations (e.g., in different computer systems connected via a network). Storage medium may store program instructions (e.g., specifically implemented as a computer program) executable by one or more processors.
[0151] Of course, the computer-executable instructions provided in the embodiments of this application are not limited to the camera shooting pose recovery method described above, but can also execute related operations in the camera shooting pose recovery method provided in any embodiment of this application.
[0152] The camera shooting pose recovery device, storage medium, and camera shooting pose recovery equipment provided in the above embodiments can execute the camera shooting pose recovery method provided in any embodiment of this application. For technical details not described in detail in the above embodiments, please refer to the camera shooting pose recovery method provided in any embodiment of this application.
[0153] The above description is merely a preferred embodiment and the technical principles employed in this application. This application is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions that can be made by those skilled in the art will not depart from the scope of protection of this application. Therefore, although this application has been described in detail through the above embodiments, this application is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of this application. The scope of this application is determined by the scope of the claims.
Claims
1. A method for restoring the pose of a camera shot, characterized in that, include: If a pose recovery event is detected, the lost image frame and the reference image frame required for recovery are determined. Image feature matching is performed based on the reference image frame and the lost image frame to obtain a first feature matching pair, and a first translation direction and a first rotation parameter are determined based on the first feature matching pair; Determine the reference translation distance between the lost image frame and the reference image frame from the successfully tracked image frames; The first translation parameter is determined based on the reference translation distance and the first translation direction; The camera pose of the lost image frame is recovered based on the first rotation parameter and the first translation parameter.
2. The method for restoring the camera shooting pose according to claim 1, characterized in that, When the reference image frame and the lost image frame are on the same flight path; Determining the reference translation distance between the lost image frame and the reference image frame from the successfully tracked image frames includes: The translation distance between two corresponding image frames is determined based on the translation parameters of adjacent image pairs, and the reference translation distance is determined based on the translation distance. The adjacent image pair consists of two adjacent image frames of the same route that were successfully tracked in the past.
3. The method for restoring the camera shooting pose according to claim 2, characterized in that, Determining the reference translation distance based on the translation distance includes: Select at least two pairs of first adjacent images that are closest to the lost image frame, and randomly select at least three pairs of second adjacent images from the remaining multiple pairs of adjacent images; The reference translation distance is determined based on the translation distances of the first adjacent image pair and the second adjacent image pair.
4. The method for restoring the camera shooting pose according to any one of claims 1-3, characterized in that, The determination of the first translation parameter based on the reference translation distance and the first translation direction includes: The coordinate translation distance of the positioning sensor is determined based on the world coordinates of the lost image frame and the reference image frame acquired by the positioning sensor. The first translation distance is determined based on the difference parameter between the reference translation distance and the coordinate translation distance; The first translation parameter is determined based on the first translation distance and the first translation direction.
5. The method for restoring the camera shooting pose according to claim 4, characterized in that, Determining the first translation distance based on the difference parameter between the reference translation distance and the coordinate translation distance includes: The ratio of the coordinate translation distance to the reference translation distance is determined as the difference parameter, and the difference parameter is compared with a preset difference range; In response to the comparison result that the difference parameter is within a preset difference range, the first translation distance is determined to be the coordinate translation distance; In response to a comparison result where the difference parameter exceeds a preset difference range, the first translation distance is determined as the reference translation distance.
6. The method for restoring the camera shooting pose according to claim 2, characterized in that, The step of restoring the camera pose of the lost image frame based on the first rotation parameter and the first translation parameter includes: From the image frames on the side path of the lost image frame, determine the constraint image frame that is closest to the lost image frame; The lost image frame and the constrained image frame are matched for image features to obtain a third feature matching pair, and the second translation direction and the second rotation parameter are determined based on the third feature matching pair. The second translation distance is determined based on the translation distance between the image frame and the nearest image frame on the lateral flight path, and the second translation parameter is determined based on the second translation direction and the second translation distance; The camera pose of the lost image frame is recovered based on the first rotation parameter, the second rotation parameter, the first translation parameter, and the second translation parameter.
7. The method for restoring the camera shooting pose according to claim 6, characterized in that, Determining the constrained image frame closest to the lost image frame from image frames along the lateral route of the lost image frame includes: The positional distance between the lost image frame and the image frame on the corresponding lateral flight path is determined based on the world coordinates of the lost image frame and the world coordinates of the image frame on the corresponding lateral flight path. Based on the location distance, the image frame on the lateral flight path that is closest to the lost image frame is determined as the constrained image frame.
8. The method for restoring the camera shooting pose according to claim 1, characterized in that, Prior to the detection of the pose recovery event, the following are included: Determine whether the first image frame meets the preset tracking success condition, wherein the first image frame is the image frame currently acquired by the camera; In response to the determination result that the first image frame satisfies the tracking success condition, it is determined whether there is an image frame stored in the lost frame container. When tracking is performed based on the first image frame, the image frame in the lost frame container is determined not to meet the tracking success condition and is added to the container. In response to the determination that the lost frame container stores an image frame, a pose recovery event is determined to have been detected.
9. The method for restoring the camera shooting pose according to claim 8, characterized in that, The process of determining the lost image frame and the reference image frame required for recovery includes: The last image frame added to the lost frame container is determined to be the lost image frame, and the first image frame is determined to be the reference image frame.
10. The method for restoring the camera shooting pose according to claim 8, characterized in that, The step of determining whether the first image frame meets the preset tracking success conditions includes: The two-dimensional feature points of the first image frame are matched with the three-dimensional feature points constructed from the second image frame to obtain the second feature matching pair; the second image frame is the previous image frame that was successfully tracked. Compare the number of the second feature matching pairs with a preset number threshold; In response to the comparison result that the number of the second feature matching pairs is less than a preset number threshold, it is determined that the first image frame does not meet the tracking success condition.
11. The method for restoring the camera shooting pose according to claim 10, characterized in that, The step of determining whether the first image frame meets the preset tracking success conditions includes: In response to a comparison result where the number of the second feature matching pairs is greater than or equal to a preset number threshold, the relative transformation parameter between the first image frame and the second image frame is determined based on the second feature matching pairs. The relative transformation parameter is compared with a preset deviation threshold. In response to the comparison result that the relative transformation parameter is less than or equal to the preset deviation threshold, it is determined that the first image frame satisfies the tracking success condition.
12. The method for restoring the camera shooting pose according to claim 1, characterized in that, The step of performing image feature matching based on the reference image frame and the lost image frame to obtain a first feature matching pair includes: The two-dimensional feature points of the reference image frame are matched with the two-dimensional feature points of the lost image frame to obtain the first feature matching pair; First feature matching pairs that do not meet geometric constraints are filtered out using random consistency.
13. A device for restoring the shooting pose of a camera, characterized in that, include: The lost image determination module is configured to determine the lost image frame and the reference image frame required for recovery if a pose recovery event is detected. The first rotation parameter determination module is configured to perform image feature matching based on the reference image frame and the lost image frame to obtain a first feature matching pair, and determine a first translation direction and a first rotation parameter based on the first feature matching pair. The reference translation distance determination module is configured to determine the reference translation distance between the lost image frame and the reference image frame from the successfully tracked image frames; The first translation parameter determination module is configured to determine the first translation parameter based on the reference translation distance and the first translation direction; The pose recovery module is configured to recover the camera shooting pose of the lost image frame based on the first rotation parameter and the first translation parameter.
14. A device for restoring the shooting pose of a camera, characterized in that, include: One or more processors; A storage device for storing one or more programs that, when executed by one or more processors, cause the one or more processors to implement the camera pose recovery method as described in any one of claims 1-12.
15. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to perform the camera pose recovery method as described in any one of claims 1-12.
Citation Information
Patent Citations
Method and device for determining posture, storage medium and electronic device
CN109035334A
Visual positioning method, visual positioning device, storage medium and electronic equipment
CN113096185A