Method of generating a three-dimensional map and method of determining a pose of a user terminal using the generated three-dimensional map

CN116508061BActive Publication Date: 2026-09-18SK TELECOM CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180073278.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-08-04
Filing Date
2021-08-19
Publication Date
2026-09-18
Estimated Expiration
2041-08-19

AI Technical Summary

Technical Problem

使用摄像机的定位方法具有相对高的定位精度,但是需要预先构建具有大数据容量的3D地图,并且存在用于位置测量的计算量大的问题

Benefits of technology

[0021]According to embodiments of this disclosure, 3D maps can be generated by using keyframes selected from images taken from the location of the user terminal, thereby reducing the data size of the 3D map.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116508061B_ABST
    Figure CN116508061B_ABST
Patent Text Reader

Abstract

A method for determining a pose of a camera included in a user terminal according to an embodiment of the present application can include the steps of determining a prediction mode for predicting a second pose of the camera in a second query image obtained by the user terminal photographing a place where the user terminal is located after a first query image based on whether a first pose of the camera has been effectively predicted in the first query image, and determining the second pose based on the determined prediction mode.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to methods for generating 3D maps and methods for determining the pose of a user terminal using the generated 3D maps. Background Technology

[0002] GPS information is typically used to measure the location of user terminals in applications such as navigation.

[0003] However, using GPS information for location measurement is not feasible in GPS-shadowed areas (e.g., in tunnels). Methods for measuring the location of user terminals in GPS-shadowed areas include radio wave-based positioning methods (e.g., positioning based on base station / AP location and signal strength, positioning based on signal time difference or angle of incidence, positioning based on signal pattern matching in a grid unit), camera-based positioning methods, and LiDAR sensor-based positioning methods, among others.

[0004] Among these, radio wave-based positioning methods using smartphones and camera-based positioning methods are suitable for positioning user terminals such as smartphones. Camera-based positioning methods offer relatively high positioning accuracy, but require the pre-construction of large-scale 3D maps and involve significant computational demands for location measurement. Summary of the Invention

[0005] Technical issues

[0006] The problem this disclosure aims to solve is to provide a method for generating a 3D map using keyframes selected from images taken from the location of a user terminal and for determining the pose information in the current frame based on whether valid pose information is obtained from previous frames.

[0007] Furthermore, the problems to be solved in this disclosure are not limited to those described above, and those skilled in the art can clearly understand from the following description another problem that is not described.

[0008] Technical solutions to the problem

[0009] According to an aspect of this disclosure, a method for determining the pose of a camera included in a user terminal is provided. The method includes: determining a prediction pattern for predicting a second pose of the camera in a second query image captured by the user terminal at the location of the user terminal, based on whether a first pose of the camera is effectively predicted in a first query image captured by the user terminal at the location of the user terminal; and determining the second pose based on the determined prediction pattern.

[0010] In this paper, determining the second pose includes: when the first pose is effectively predicted, using the first pose to select one or more matching frames from a plurality of keyframes included in the pre-generated 3D map and captured at the location as matching candidates for the second query image.

[0011] In this paper, the determination of the second pose includes: if the first pose is not effectively predicted, using a pre-learned neural network to determine the global features of the second query image; and selecting one or more matching frames from the plurality of keyframes as matching candidates with the second query image by comparing the global features of the second query image with the global features of each of the multiple keyframes included in the pre-generated 3D map and captured at the location.

[0012] In this paper, the determination of the second pose includes: if the first pose is effectively predicted but a preset condition is met, then by comparing the global features of the second query image determined using a pre-learned neural network with the global features of each of a plurality of keyframes included in a pre-generated 3D map and capturing the location, one or more frames are selected as matching frames from the plurality of keyframes using the first pose, and the preset condition includes at least one of the following: the number of poses effectively predicted by the camera is equal to or greater than a preset threshold number; and the distance difference between the second pose determined according to the prediction mode when the first pose is effectively predicted and the second pose determined according to the prediction mode when the first pose is not effectively predicted is equal to or greater than a preset threshold distance.

[0013] In this document, determining the second pose includes: selecting one or more matching images from a plurality of keyframes included in a pre-generated 3D map and captured at the location that match the second query image; and determining the second pose in the second query image by using local feature matching between local features included in one or more matching images and local features included in the second query image. In this document, the 3D map is generated using a plurality of keyframes selected from images captured at the location by the location capture device or the terminal pose determination device using predetermined criteria, and the predetermined criteria include at least one of the time interval and the movement distance of the location capture device or the terminal pose determination device.

[0014] In this paper, a 3D map is generated by further using an image generated by rotating at least one of a plurality of keyframes at one or more angles.

[0015] In this paper, determining the second pose includes: performing local feature matching between a first local feature extracted from the second query image and a second local feature in each of one or more keyframes included in the pre-generated 3D map and captured at the location; and using the result of the local feature matching to determine the second pose in the second query image, wherein the second local feature is the local feature with the shortest distance when calculating the distance between the first local feature and a local feature that matches each of the first local features in a plurality of local features included in each of the one or more keyframes.

[0016] The method further includes: obtaining information from the user terminal, together with obtaining the second query image, about whether the first pose was effectively predicted, the first pose information, information about the last pose of the user terminal predicted when the previous query image was not effectively predicted, and the number of images in the previous images in which the pose was determined to be effectively predicted.

[0017] The method further includes: obtaining information from the user terminal, along with obtaining the second pose, about whether the second pose was effectively predicted, and determining the number of images in which the pose was effectively predicted among the images preceding the second query image and the second query image.

[0018] According to another aspect of this disclosure, an apparatus for determining a terminal pose is provided. The apparatus includes: a transceiver configured to receive from a user terminal a first query image capturing the location of the user terminal; and a processor configured to determine a prediction pattern for predicting a second pose of the camera in a second query image based on whether a first pose of the user terminal's camera is effectively predicted in the first query image, and to determine the second pose based on the determined prediction pattern, the second query image being received from the user terminal after the first query image and capturing the location therein.

[0019] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided for storing a computer program programmed to perform a method for determining the pose of a camera included in a user terminal. The method includes: determining a prediction pattern for predicting a second pose of the camera in a second query image captured by the user terminal at the location of the user terminal, based on whether a first pose of the camera is effectively predicted in a first query image captured by the user terminal at the location of the user terminal; and determining the second pose based on the determined prediction pattern.

[0020] Beneficial effects of the present invention

[0021] According to embodiments of this disclosure, 3D maps can be generated by using keyframes selected from images taken from the location of the user terminal, thereby reducing the data size of the 3D map.

[0022] Furthermore, according to the embodiments of this disclosure, when valid pose information is obtained in a previous frame, global feature matching is not performed, and the number of local features used for local feature matching is reduced, thereby reducing the computational load for obtaining pose information of the user terminal. Attached Figure Description

[0023] Figure 1 This is a block diagram illustrating a system for determining terminal posture according to an embodiment of the present disclosure.

[0024] Figure 2 This is a block diagram illustrating a 3D map generating apparatus for generating 3D maps according to an embodiment of the present disclosure.

[0025] Figure 3 This is a block diagram conceptually illustrating the functionality of a terminal posture determination model according to an embodiment of the present disclosure.

[0026] Figure 4A and Figure 4B An example of generating an image by rotating keyframes according to an embodiment of this disclosure is shown.

[0027] Figure 5 An example of comparing query images and matching frames according to an embodiment of this disclosure is shown.

[0028] Figure 6 An example of matching local features included in a query image with local features included in a matching image, according to an embodiment of the present disclosure, is shown.

[0029] Figure 7 An embodiment of the present disclosure is shown, which involves finding local feature pairs by matching local features included in a query image with local features included in a matching frame.

[0030] Figure 8A and Figure 8B An example is shown where local feature matching is performed only on a predetermined number of local feature pairs in order of shortest distance, for local features included in the query image and local features included in the matching frame that have the shortest distance.

[0031] Figure 9 This is a block diagram illustrating a method for training a global feature-based neural network according to an embodiment of the present disclosure.

[0032] Figure 10 This is a block diagram illustrating a method for determining a neural network by training local features according to an embodiment of the present disclosure.

[0033] Figure 11 This is a flowchart illustrating a method for generating a 3D map according to an embodiment of the present disclosure.

[0034] Figure 12A and Figure 12B This is a flowchart illustrating a method for generating a 3D map included in a user terminal according to an embodiment of the present disclosure. Detailed Implementation

[0035] The advantages and features of this disclosure, as well as methods for achieving these advantages and features, will become clear from the following description taken in conjunction with the accompanying drawings. However, the embodiments are not limited to those described, as embodiments can be implemented in various forms. It should be noted that this disclosure is provided for the purpose of full disclosure and to allow those skilled in the art to understand the exhaustive scope of the embodiments. Therefore, the embodiments will be defined only by the scope of the appended claims.

[0036] In describing embodiments of this disclosure, detailed descriptions of known components or functions will be omitted if it is determined that such descriptions unnecessarily obscure the main points of this disclosure. Furthermore, the terminology described below is defined in consideration of the functionality of embodiments of this disclosure and may vary depending on the intent or practice of the user or operator. Therefore, its definition can be based on the entire contents of this specification.

[0037] Figure 1 This is a block diagram illustrating a system for determining terminal posture according to an embodiment of the present disclosure.

[0038] Reference Figure 1 The terminal attitude determination system 10 may include a user terminal 100 and a terminal attitude determination device 200.

[0039] User terminal 100 may include a first processor 110, a camera 120, a first transceiver 130, and a first memory 140.

[0040] The first processor 110 can control the overall functions of the user terminal 100.

[0041] Camera 120 can capture images around user terminal 100.

[0042] The first transceiver 130 can send images captured by the camera 120 to the terminal attitude determination device 200, and receive information about the location of the user terminal 100 from the terminal attitude determination device 200.

[0043] According to this embodiment, the first transceiver 130 may transmit, along with images captured by the camera 120, information regarding whether the attitude information in the previous (immediately preceding) frame was effectively predicted, the attitude information in the previous (immediately preceding) frame, the attitude information of the user terminal 100 estimated immediately before operation in the terminal attitude tracking mode (i.e., the most recent attitude information of the user terminal 100 estimated in the initial attitude estimation mode), and the number of operations in the terminal attitude tracking mode (e.g., consecutive or non-consecutive).

[0044] The first memory 140 can capture images using the camera 120, send the images to the terminal attitude determination device 200 using the first transceiver 130, and store one or more models for receiving information about the location of the user terminal 100 from the terminal attitude determination device 200.

[0045] In this specification, "model" refers to the software (computer program code) or a collection thereof used to perform the functions of each device, and can be implemented by a series of commands.

[0046] The first processor 110 can execute one or more programs stored in the first memory 140, capture images, send the captured images to the terminal attitude determination device 200, and receive information about the location of the user terminal 100 from the terminal attitude determination device 200.

[0047] The terminal attitude determination device 200 can use images received from the user terminal 100 and pre-stored 3D maps to predict the attitude information of the user terminal 100.

[0048] In this specification, attitude information may include information about the position and orientation of the user terminal 100 (i.e., the position and orientation of the camera 120 in the user terminal 100). For example, attitude information may be represented using six degrees of freedom (DOF) (i.e., position information including X, Y, and Z, and rotation information including pitch, roll, and yaw).

[0049] According to this embodiment, the terminal attitude determination device 200 can be implemented as a server (e.g., a cloud server), but is not limited thereto. That is, the terminal attitude determination device 200 can be implemented as any device that predicts the attitude information of the user terminal 100 by receiving images from the user terminal 100.

[0050] The terminal attitude determination device 200 may include a second processor 210, a second transceiver 220, and a second memory 230.

[0051] The second processor 210 can control the overall functions of the terminal attitude determination device 200.

[0052] The second transceiver 220 can receive images captured by the camera 120 from the user terminal 100 and send the attitude information of the user terminal 100 predicted by the terminal attitude determination model 260 to the user terminal 100.

[0053] The second memory 230 can store the 3D map 250 and the terminal attitude determination model 260.

[0054] The second processor 210 can execute the terminal pose determination model 260 to predict the pose information of the user terminal 100 based on the image and 3D map 250 received from the user terminal 100.

[0055] Reference Figure 2 This document describes in detail the method for generating a 3D map 250 using a 3D map generation device, and will refer to... Figure 3 Describe in detail the functions of the terminal attitude determination model 260.

[0056] In this specification, for ease of description, it is described that the 3D map 250 is stored in the second memory 230, but this is not a limitation. That is, the terminal attitude determination system 10 may include a database (not shown) different from the terminal attitude determination device 200, and the 3D map 250 may be stored in the database (not shown). In this case, the terminal attitude determination device 200 can receive the 3D map 250 from the database (not shown) and predict the attitude information of the user terminal 100.

[0057] Figure 2 This is a block diagram illustrating a 3D map generating apparatus for generating 3D maps according to an embodiment of the present disclosure.

[0058] Reference Figure 1 and Figure 2 The 3D map generation device 300 may include a third processor 310, a camera unit 320, and a third memory 330.

[0059] According to this embodiment, the 3D map generation device 300 can be the same as the terminal pose determination device 200, but is not limited thereto. That is, the terminal pose determination device 200 can generate a 3D map 250 and use the generated 3D map 250 to determine the pose information of the camera 120 included in the user terminal 100. According to this embodiment, the 3D map generation device 300 can generate a 3D map 250, and the terminal pose determination device 200 can use the 3D map 250 generated by the 3D map generation device 300 to determine the pose information of the camera 120 included in the user terminal 100.

[0060] The third processor 310 can control the overall functions of the 3D map generation device 300.

[0061] Camera unit 320 may include one or more cameras. Camera unit 320 may use one or more cameras to generate images of the locations from which a 3D map will be generated.

[0062] According to this embodiment, when the camera unit 320 includes multiple cameras, the multiple cameras can be mounted to face different directions so as to capture images in a 360-degree position.

[0063] In this specification, for ease of description, it is described that the 3D map generation device 300 uses one or more cameras included in the camera unit 320 to capture locations, but it is not limited thereto. That is, according to this embodiment, the 3D map generation device 300 may not include the camera unit 320, and in this case, the 3D map generation device 300 may receive (or input) images captured by an external camera unit.

[0064] Furthermore, according to this embodiment, in addition to the camera unit 320, the 3D map generation device 300 may include a sensor unit (not shown). According to this embodiment, the sensor unit (not shown) may include a LiDAR. The 3D map generation device 300 can generate an image of the location and information about features included in the image by using the camera unit 320 and the sensor unit (not shown).

[0065] The third memory 330 can store the 3D map generation model 350.

[0066] The third processor 310 can execute the 3D map generation model 350 to generate a 3D map that includes an image of the location and information about the feature points included in the image.

[0067] The 3D map generation model 350 may include a keyframe selection unit 351, a keyframe rotation unit 353, a global feature determination unit 355, a local feature determination unit 357, a pose calculation unit 359, and a 3D map generation unit 361.

[0068] Figure 2The keyframe selection unit 351, keyframe rotation unit 353, global feature determination unit 355, local feature determination unit 357, pose calculation unit 359, and 3D map generation unit 361 shown are conceptual divisions of the functions of the 3D map generation model 350 to facilitate explanation of the functions of the 3D map generation model 350, but are not limited thereto. That is, according to this embodiment, the functions of the keyframe selection unit 351, keyframe rotation unit 353, global feature determination unit 355, local feature determination unit 357, pose calculation unit 359, and 3D map generation unit 361 can be combined / separated, and can be implemented as a series of commands included in one or more programs.

[0069] The keyframe selection unit 351 can select keyframes from multiple frames included in the image captured using the camera unit 320 for creating a 3D map.

[0070] More specifically, the 3D map generating device 300 can capture images while moving in a zigzag pattern to the location where a 3D map is to be created, and the keyframe selection unit 351 can select keyframes according to predetermined criteria.

[0071] Pre-defined criteria may include the time interval and movement distance of the 3D map generation device 300.

[0072] In other words, among the frames included in the image captured by the 3D map generation device 300, some adjacent frames may be frames of essentially the same scene, depending on the moving speed of the 3D map generation device 300. Therefore, the keyframe selection unit 351 can select keyframes according to predetermined criteria, thereby generating the 3D map 250 without using overlapping frames.

[0073] The keyframe rotation unit 353 can generate one or more images by rotating the keyframe selected by the keyframe selection unit 351 by one or more angles, and use the generated images as keyframes.

[0074] This is because even when using multiple cameras to capture images, there are limitations to capturing the location in all directions, and recognition performance degrades if the field of view changes significantly between images captured simultaneously. When the keyframe rotation unit 353 uses an image rotated one frame in each direction as a keyframe, the uncaptured area between frames is minimized, and the rotated image and the original frame are captured at the same position, thus preventing a decrease in recognition performance.

[0075] Therefore, the keyframe rotation unit 353 can use Equation 1 below to rotate the selected keyframe by one or more angles.

[0076] Formula 1

[0077] m′=KRK -1 m

[0078] Here, K can represent the preset parameters inside the camera, R can represent the rotation transformation matrix, m can represent the coordinates of the keyframe before rotation, and m' can represent the coordinates of the keyframe after rotation.

[0079] In other words, the keyframe rotation unit 353 can generate an image by rotating the selected keyframe using preset parameters and a preset rotation transformation matrix inside the camera.

[0080] For example, further reference Figure 4A and Figure 4B , Figure 4A The keyframes selected by the keyframe selection unit 351 according to predetermined criteria are shown, while Figure 4B An image showing the keyframe rotation unit 353 rotating a keyframe using Equation 1 is shown.

[0081] When Figure 4A When a keyframe is selected, the keyframe rotation unit 353 can generate a value by rotating the keyframe upwards by a predetermined angle. Figure 4B The image. However, if the size of the keyframe before rotation and the size of the image after rotation are the same, then region B, which was not included in the keyframe, can be included in the image after the keyframe is rotated. In this case, region B, which was not included in the keyframe, can be represented in black, as shown. Figure 4B As shown.

[0082] The keyframes selected by the keyframe selection unit 351 and the keyframes generated by the keyframe rotation unit 353 (hereinafter referred to as keyframes) may include global feature information, one or more local feature information and pose information.

[0083] As described below, the global feature information, local feature information, and pose information included in the keyframe can be determined by the global feature determination unit 355, the local feature determination unit 357, and the pose calculation unit 359.

[0084] First, the global feature determination unit 355 can determine global features represented as high-dimensional vectors of keyframes. Global features can be identifiers used to distinguish keyframes from other keyframes.

[0085] According to this embodiment, the global feature determination unit 355 may include a pre-learned global feature determination neural network (e.g., NetVLAD) to output global features of the input keyframe when the keyframe is input. In this case, when the keyframe is input to the previously learned global feature determination neural network, the global feature determination unit 355 can use the vector output from the global feature determination neural network to determine the global features.

[0086] Reference Figure 9 A method for training a global feature determination neural network, including a global feature determination unit 355, is described.

[0087] The local feature determination unit 357 can extract one or more local features included in the keyframe and determine a high-dimensional vector representing the location information and descriptor information of the local features.

[0088] Location information can include the position of local features on the keyframe (2D position) and the absolute position of local features (e.g., (x, y, z) coordinates on Earth) (3D position).

[0089] Additionally, descriptor information is used to distinguish local features from other local features (i.e., another local feature included in the same keyframe or a local feature included in another keyframe), and it can imply the correlation between a local feature and the pixels surrounding the local feature.

[0090] According to this embodiment, the local feature determination unit 357 may include a pre-learned local feature determination neural network (e.g., SuperPoint) to output a high-dimensional vector representing the local features of the input keyframe when the keyframe is input. In this case, when the keyframe is input to the local feature determination neural network, the local feature determination unit 357 can use the vector output from the local feature determination neural network to determine the local features.

[0091] Reference Figure 10 A method for training a local feature determination neural network, including a local feature determination unit 357, is described.

[0092] The pose calculation unit 359 can input the local features determined by the local feature determination unit 357 into a preset SLAM algorithm, and calculate the pose information of the key frame and the coordinate information of the local features included in the key frame (e.g., 3D absolute coordinates on Earth).

[0093] More specifically, the pose calculation unit 359 can calculate the pose information of a keyframe by comparing local features included in the keyframe with local features included in other keyframes in which pose information is predefined.

[0094] For each keyframe selected or generated by the keyframe selection unit 351 and the keyframe rotation unit 353, as described above, it may include global features, one or more local features, and attitude information determined by the global feature determination unit 355, the local feature determination unit 357, and the attitude calculation unit 359.

[0095] The 3D map generation unit 361 can generate a 3D map by storing keyframes together with global features, local features, and pose information corresponding to the keyframes. That is, the 3D map generation unit 361 can generate a 3D map using keyframes that include global features of the keyframes, one or more local features included in the keyframes, and coordinate information of the local features included in the keyframes.

[0096] The 3D map generation device 300 can store the generated 3D map in the terminal attitude determination device 200 (or database).

[0097] Figure 3 This is a block diagram conceptually illustrating the functionality of a terminal posture determination model according to an embodiment of the present disclosure.

[0098] Reference Figure 1 and Figure 3 The terminal attitude determination model 260 can use images received from the user terminal 100 and a pre-stored 3D map 250 to predict the attitude information of the user terminal 100.

[0099] Therefore, the terminal pose determination model 260 may include a prediction mode determination unit 261, a matching frame selection unit 263, a local feature matching unit 265, and a camera pose determination unit 267.

[0100] Describing by conceptually dividing the functions of the terminal attitude determination model 260 Figure 3 The prediction mode determination unit 261, matching frame selection unit 263, local feature matching unit 265, and camera pose determination unit 267 shown are provided to easily illustrate the function of the terminal pose determination model 260, but are not limited thereto. That is, according to this embodiment, the functions of the prediction mode determination unit 261, matching frame selection unit 263, local feature matching unit 265, and camera pose determination unit 267 can be combined / separated, and can be implemented as a series of commands included in the program.

[0101] The prediction mode determination unit 261 can determine a prediction mode for predicting the attitude of the user terminal 100 based on whether the attitude information of the user terminal 100 in the previous frame has been effectively predicted.

[0102] More specifically, when the pose information of user terminal 100 in a previous frame is not effectively predicted, prediction mode determination unit 261 can operate in initial pose estimation mode to predict the pose information of camera 120 when the position of user terminal is unknown. And when the pose information of user terminal 100 in the previous frame is effectively predicted, prediction mode determination unit 261 determines the user when the position of terminal 100 is known. Terminal pose tracking mode prediction includes the pose information of camera 120 in user terminal 100 and the pose of user terminal 100 in the previous frame. When the information is effectively predicted but a preset condition is met, prediction mode determination unit 261 can operate in an intermediate mode combining initial pose estimation mode and terminal pose tracking mode.

[0103] In other words, if the pose information of user terminal 100 in a previous frame is not effectively predicted, then the terminal pose determination model 260 does not have the correct pose information of user terminal 100 at the current location, and therefore must perform all the processing required to determine the pose information of user terminal 100. However, if the pose information of user terminal 100 in a previous frame is effectively predicted, then the terminal pose determination model 260 can use the pose information of user terminal 100 in the previous frame to determine the pose information of user terminal 100 in the current frame.

[0104] Therefore, the terminal attitude determination model 260 can reduce the amount of computation required to determine the attitude information of the user terminal 100, and also improve the accuracy of the attitude information of the user terminal 100.

[0105] When the prediction mode determined by the prediction mode determination unit 261 is the initial pose estimation mode, the matching frame selection unit 263 can use the previously learned global feature determination neural network to determine the global features of the query image received from the user terminal 100. According to this embodiment, the global feature determination neural network can be the same neural network included in the global feature determination unit 355.

[0106] The matching frame selection unit 263 can compare the global features of the query image with the global features of multiple keyframes stored in the 3D map 250, and select one or more keyframes with the smallest global feature difference from the query image as matching frames. Here, a matching frame can refer to a frame used to determine the pose information of the query image through local features that match the query image.

[0107] In other words, the matching frame selection unit 263 can calculate the distance (e.g., Euclidean distance) between the global features of the query image and the global features of multiple keyframes stored in the 3D map 250, and select a predetermined number of keyframes as matching frames in the order of the shortest distance.

[0108] For example, further reference Figure 5 The matching frame selection unit 263 can calculate the distance between the global features of the query image Qi and the global features of multiple keyframes stored in the 3D map 250, and select three predetermined keyframes as matching frames MF1, MF2 and MF3 in the order of the shortest distance.

[0109] On the other hand, when the prediction mode determined by the prediction mode determination unit 261 is the terminal attitude tracking mode, the matching frame selection unit 263 can use the attitude information of the camera 120 predicted in the previous frame to select one or more matching frames.

[0110] More specifically, the matching frame selection unit 263 can use Equation 2 below to select one or more matching frames.

[0111]

Formula 2

[0112] T Relative =(T Prev ) -1 T KeyFrame.

[0113] Here, T KeyFrame It can be a matrix representing the pose information of camera 120 predicted in a previous frame, T Prev This can represent the transformation matrix used to determine the pose information of camera 120 in the current frame using pose information of camera 120 in the previous frame, and T Relative It can be a matrix representing the rotation and position changes of camera 120 from the previous frame to the current frame.

[0114] In other words, the matching frame selection unit 263 can use the pose information T of the camera 120 predicted in the previous frame. KeyFrame and transformation matrix T Prev inverse transform T Prev -1 To determine the relative pose information T of camera 120 in the current frame. Relative .

[0115] Matching frame selection unit 263 can select frames that are normalized relative attitude information T. RelativeKeyframes whose obtained value (i.e., the change in position of camera 120 from the previous frame to the current frame) is less than or equal to a preset threshold distance and whose rotation change of camera 120 from the previous frame to the current frame (e.g., the change in rotation angle when the matrix is ​​represented as axis angle) is less than or equal to a preset threshold angle are used as matching frames.

[0116] Finally, when the prediction mode determined by the prediction mode determination unit 261 is an intermediate mode, in order to make a trade-off between efficiency and accuracy, the matching frame selection unit 263 can select a matching frame by combining the initial attitude estimation mode and the terminal attitude tracking mode.

[0117] The preset conditions may include at least one of the following: the number of operations in the terminal attitude tracking mode is equal to or greater than a threshold; and the difference (e.g., distance difference) between the attitude information of the camera 120 calculated in the terminal attitude tracking mode and the attitude information of the camera 120 calculated in the initial attitude estimation mode is equal to or greater than a preset threshold.

[0118] In other words, even if the attitude information is determined to be effectively predicted, if the number of operations in the terminal attitude tracking mode is equal to or greater than a threshold, or if the difference (e.g., distance difference) between the attitude information of the camera 120 calculated in the terminal attitude tracking mode and the attitude information of the camera 120 calculated in the initial attitude estimation mode is equal to or greater than a preset threshold, since the effectiveness of the predicted attitude information cannot be guaranteed, the prediction mode determination unit 261 can select a matching frame by combining the initial attitude estimation mode and the terminal attitude tracking mode.

[0119] More specifically, when the prediction mode determined by the prediction mode determination unit 261 is an intermediate mode, the matching frame selection unit 263 can select one or more frames selected using Equation 2 and frames selected using the distance difference of global features as matching frames. In this case, the number of frames selected using the distance difference of global features can be less than the number of matching frames selected in the initial attitude estimation mode.

[0120] As described above, since the matching frame selection unit 263 selects matching frames by dividing the prediction mode of the user terminal 100, the matching frame selection unit 263 can reduce the amount of computation required to select matching frames to perform local feature matching with the query image.

[0121] Subsequently, the local feature matching unit 265 can extract one or more local features from the query image and can perform local feature matching, which matches the local features extracted from the query image with the local features included in the matching frame selected by the matching frame selection unit 263.

[0122] The local feature matching unit 265 can perform local feature matching and calculate the distance (e.g., Euclidean distance) between local features included in the query image and local features included in the matching frame. The distance between local features included in the query image and local features included in the matching frame implies the similarity between the local features included in the query image and local features included in the matching frame, and it can be the distance (i.e., similarity) between local feature descriptors included in the matching frame.

[0123] For example, further reference Figure 6 The local feature matching unit 265 can calculate the distance between all local features included in the query image Qi and all local features included in the second matching frame MF2, and perform local feature matching by matching the local feature with the shortest distance.

[0124] Meanwhile, due to errors in the local feature matching process, the local feature matching unit 265 may incorrectly match local features. Therefore, in order to reduce matching errors, the local feature matching unit 265 may determine the validity of the local feature matching based on the distance difference between local features included in the query image and local features included in the matching frame.

[0125] More specifically, if the ratio of a first distance between local features included in the query image and local features included in the matching frame that have the shortest distance to each other and a second distance between local features included in the query image and local features included in the matching frame that have the second shortest distance to each other is equal to or less than a predetermined threshold, then the local feature matching unit 265 can determine that the matching result is valid.

[0126] On the other hand, if the ratio of the first distance to the second distance is greater than (or equal to or greater than) a predetermined threshold, the local feature matching unit 265 can determine that the result of the local feature matching is invalid, and can perform local feature matching between the query image and the matching frame again.

[0127] This is because the matching frame is the keyframe with the most similar global features to the query image, and the matching frames are images that are similar to each other or at least have no significant differences. That is, since the matching frames are similar to each other, it can be expected that there are no significant differences in the local feature matching between the query image and the matching frames, and if there are significant differences in the local feature matching between each query image and each matching frame for each matching frame, it can be assumed that the matching between local features has been performed incorrectly.

[0128] Furthermore, depending on the position and orientation of the user terminal 100, the query image may include many local features. In this case, the local feature matching unit 265 can perform as many local feature matches as possible (including the number of local features in the query image * the number of local features in the matching frames * the number of matching frames), and if all matching pairs are used to determine the pose information, the computational load performed by the camera pose determination unit 267 may increase.

[0129] Therefore, in order to reduce the amount of computation used for local feature matching, the local feature matching unit 265 can calculate the distance between the local features included in the query image and the local features included in and matched in each matching frame, select a predetermined number of local features from the local features included in the query image in order of the shortest distance to the local features included in and matched in the matching frame, and use the local features selected from the local features included in the query image to determine the pose information to be performed by the camera pose determination unit 267.

[0130] Figure 7 A method for matching local features included in a query image with local features included in a matching frame, according to another embodiment of the present disclosure, is shown.

[0131] For example, refer to Figure 7 The local feature matching unit 265 can determine the local features with the shortest distance as local feature pairs from N matching frames MF-1, MF-2, ... for the local features lf-1a, lf-2a, lf-3a, and lf-4a of the query image Qi. Subsequently, the local feature matching unit 265 can arrange multiple local feature pairs in descending order of distance and select a predetermined number of local feature pairs from the multiple local feature pairs, thereby reducing the amount of computation performed by the camera pose determination unit 267.

[0132] For example, the local feature matching unit 265 can match the local feature lf-1b of the first matching frame MF-1, which is closest to lf-1a among the local features of the query image QI, as a first local feature pair. Subsequently, the local feature matching unit 265 can match the local feature lf-2b of the first matching frame MF-1, which is closest to lf-2a among the local features of the query image QI, as a second local feature pair. The local feature matching unit 265 can repeat the matching process to match all local features of the query image QI with the local features of N matching frames MF-1, MF-2, ... to obtain multiple local feature pairs.

[0133] In this embodiment, multiple local feature pairs can be (lf-1a, lf-1b), (lf-1a, lf-1c), (lf-2a, lf-2b), (lf-2a, lf-2c), (lf-3a, lf-2b), (lf-3a, lf-3c), (lf-4a, lf-4b), and (lf-4a, lf-4c).

[0134] Subsequently, the local feature matching unit 265 can arrange multiple local feature pairs in descending order of distance, and can select a predetermined number of local feature pairs from the multiple local feature pairs. For example, if the predetermined number is 4, and the result of arranging multiple local feature pairs in descending order of distance is (lf-1a, lf-1b), (lf-4a, lf-4b), (lf-1a, lf-1c), (lf-2a, lf-2c), (lf-3a, lf-3c), (lf-3a, lf-2b), (lf-4a, lf-4c), (lf-2a, lf-2b), the local feature matching unit 265 executes (lf-1a, lf-1b), (lf-4a, lf-4b), (lf-1a, lf-1c), and (lf-2a, lf-2c) as pose to determine local feature pairs.

[0135] By using this method, the amount of computation performed when determining the pose can be reduced because the number of local feature pairs used by the camera pose determination unit 267 to determine the pose is reduced, and the probability that incorrect matching pairs may be used to determine the pose information can be reduced.

[0136] For example, further reference Figure 8A and Figure 8B , Figure 8A The results of performing local feature matching on all local features included in the query image are shown, and Figure 8B The results show that for local features included in the query image and local features included in the matching frame and included in the matched query image, local feature matching is performed only on a predetermined number of local feature pairs in order of the shortest distance, for local feature pairs that have the shortest distance.

[0137] In other words, such as Figure 8A and Figure 8B As shown, by reducing the number of local feature pairs used by the camera pose determination unit 267 to determine the pose, not only can the computational load of pose determination be reduced, but the probability of using incorrect local feature matching pairs to determine the pose can also be reduced.

[0138] The camera pose determination unit 267 can determine the pose information of the query image by using pose determination local feature pairs between the query image and the matching frame.

[0139] More specifically, the camera pose determination unit 267 can determine the pose information of the query image by using the two-dimensional coordinates of local features included in the query image (i.e., coordinates on the query image) and the three-dimensional coordinates of the matching local features (i.e., the 3D absolute coordinates of the local features on the Earth).

[0140] For example, the camera pose determination unit 267 can predict the pose information of the query image (i.e., the pose information of the camera 120 when the query image is captured) by inputting the two-dimensional coordinates of local features included in the query image (i.e., coordinates on the query image) and the three-dimensional coordinates of local features included in the matching frame and matched (i.e., the 3D absolute coordinates of the local features on the earth) into a pose prediction algorithm (e.g., the P3P algorithm).

[0141] Additionally, the camera pose determination unit 267 can determine the number of correctly matched local features in the local feature matching between the query image and the matching image by using an error elimination algorithm (e.g., the Random Sample Consensus (RANSAC) algorithm).

[0142] When the number of correctly matched local features in the local feature matching between the query image and the matching image is equal to or greater than (or greater than) a preset threshold, the camera pose determination unit 267 can identify the pose information of the query image predicted based on the pose prediction algorithm as valid information, and determine the predicted pose information as the pose information of the query image.

[0143] On the other hand, when the number of correctly matched local features in the local feature matching between the query image and the matching image is less than (or equal to or less than) a preset threshold, the camera pose determination unit 267 can identify the pose information predicted based on the pose prediction algorithm as invalid information, and determine the pose information of the query image again by resetting the prediction mode.

[0144] The terminal attitude determination model 260 can send the determined attitude information of the query image to the user terminal 100.

[0145] At this time, according to this embodiment, the terminal attitude determination model 260 can send information about whether the attitude information has been effectively predicted, along with the attitude information of the query image, and the number of times the terminal has operated in attitude tracking mode (e.g., consecutive or non-consecutive times).

[0146] The number of times the terminal operates in attitude tracking mode can be initialized under the control of the prediction mode determination unit 261. For example, the prediction mode determination unit 261 can initialize the number of times the prediction mode operates as an intermediate mode or an initial attitude estimation mode.

[0147] Figure 9This is a block diagram illustrating a method for training a global feature decision neural network according to an embodiment of the present disclosure.

[0148] Reference Figure 2 and Figure 9 If an image is input (e.g., a keyframe used to generate a 3D map 250), the global feature determination neural network 356 included in the global feature determination unit 355 can be trained to output a vector used to distinguish the image from other images.

[0149] More specifically, when the global feature determination neural network 356 receives a reference image as input data and a reference vector as input label data, it can be trained to output a high-dimensional vector as global features of the reference image.

[0150] In addition, the global feature determination neural network 356 can be trained by receiving backpropagation values ​​as feedback to reduce the difference between the reference vector and the output high-dimensional vector.

[0151] Figure 10 This is a block diagram illustrating a method for determining a neural network by training local features according to an embodiment of the present disclosure.

[0152] Reference Figure 2 and Figure 10 If an image is input (e.g., a keyframe used to generate a 3D map 250), the local feature determination neural network 358, included in the local feature determination unit 357, can be trained to output a high-dimensional vector representing the local features contained in the image.

[0153] More specifically, when the local feature determination neural network 358 receives a reference vector of local features included in the reference image as input data and the reference image as input label data, the local feature decision neural network 358 can be trained to output a high-dimensional vector of local features included in the reference image.

[0154] In addition, the local feature determination neural network 358 can be trained by further receiving backpropagation values ​​as feedback to reduce the difference between the reference vector and the output high-dimensional vector.

[0155] Figure 11 This is a flowchart illustrating a method for generating a 3D map according to an embodiment of the present disclosure.

[0156] Reference Figure 2 and Figure 11The keyframe selection unit 351 can select keyframes for generating a 3D map from multiple frames included in the image captured by the camera unit 320 based on predetermined criteria (S1000), and the keyframe rotation unit 353 can generate one or more images by rotating the keyframes selected by the keyframe selection unit 351 by one or more angles, and use the generated images as keyframes (S1010).

[0157] When a keyframe is input to a pre-learned global feature determination neural network, the global feature determination unit 355 can determine global features based on the vector output from the global feature determination neural network (S1020).

[0158] When a keyframe is input to a pre-learned local feature determination neural network, the local feature determination unit 357 can use the vector output from the local feature determination neural network to determine local features (S1030).

[0159] The pose calculation unit 359 can input the local features determined by the local feature determination unit 357 into a preset SLAM algorithm, and can calculate the pose information of the key frame and the coordinate information of the local features included in the key frame (e.g., 3D absolute coordinates on Earth) (S1040).

[0160] The 3D map generation unit 361 can generate a 3D map 250 (S1050) by storing keyframes together with global features, local features and coordinate information corresponding to the keyframes.

[0161] Figure 12A and Figure 12B This is a flowchart illustrating a method for generating a 3D map included in a user terminal according to an embodiment of the present disclosure.

[0162] Reference Figure 3 , Figure 12A and Figure 12B The prediction mode determination unit 261 can determine the prediction mode for predicting the attitude of the user terminal 100 based on whether the attitude information of the user terminal 100 in the previous frame has been effectively predicted.

[0163] When the prediction mode determined by the prediction mode determination unit 261 is the initial pose estimation mode, the matching frame selection unit 263 can use the previously learned global feature determination neural network to determine the global features of the query image received from the user terminal 100 (S1110), and can select one or more keyframes with the smallest difference between the query image and the global features from the multiple keyframes stored in the 3D map 250 as matching frames by comparing the global features of the query image with the global features of multiple keyframes stored in the 3D map 250 (S1120).

[0164] On the other hand, when the prediction mode determined by the prediction mode determination unit 261 is the terminal attitude tracking mode, the matching frame selection unit 263 can use the attitude information of the camera 120 predicted in the previous frame to select one or more matching frames (S1130).

[0165] Finally, when the prediction mode determined by the prediction mode determination unit 261 is an intermediate mode, the matching frame selection unit 263 can select one or more frames selected by using the pose information of the camera 120 predicted in the previous frame and the frame selected by using the distance difference of global features as matching frames (S1140).

[0166] Subsequently, the local feature matching unit 265 can extract one or more local features from the query image and can perform local feature matching, which matches the local features extracted from the query image with the local features included in the matching frame selected by the matching frame selection unit 263 (S1150).

[0167] The camera pose determination unit 267 can use the result of local feature matching between the query image and the matching frame to predict the pose information of the query image (S1160).

[0168] If the number of correctly matched local features in the local feature matching between the query image and the matching image is equal to or greater than (or greater than) a preset threshold ("Yes" in S1170), then the camera pose determination unit 267 can identify the pose information of the query image predicted based on the pose prediction algorithm as valid information, and determine the predicted pose information as the pose information of the query image (S1180).

[0169] On the other hand, if the number of correctly matched local features in the local feature matching between the query image and the matching image is less than (or equal to or less than) a predetermined threshold ("No" in S1170), the camera pose determination unit 267 can identify the pose information of the query image predicted based on the pose prediction algorithm as invalid information, and the prediction mode determination unit 261 can determine the pose information of the query image again by resetting the prediction mode.

[0170] Each flowchart of this disclosure can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the steps of the flowchart. These computer program instructions can also be stored in a computer-usable or computer-readable medium that can direct the computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-usable or computer-readable medium can produce an article of writing that includes instructions for implementing the functions specified in the boxes of the flowchart. The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide processing for implementing the functions specified in the boxes of the flowchart.

[0171] Each step in the flowchart can represent a module, code segment, or code section comprising one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the boxes may not occur in the order shown in the figures. For example, two boxes shown consecutively may actually execute substantially simultaneously, or the two boxes may sometimes execute in reverse order depending on the functions involved.

[0172] The above description is merely an exemplary description of the technical scope of this disclosure, and those skilled in the art will understand that various changes and modifications can be made without departing from the original characteristics of this disclosure. Therefore, the embodiments disclosed in this disclosure are intended to illustrate, not limit, the technical scope of this disclosure, and the technical scope of this disclosure is not limited by the embodiments. The scope of protection of this disclosure should be interpreted based on the appended claims, and it should be understood that all technical scopes included within their equivalents are included within the scope of protection of this disclosure.

Claims

1. A method for determining the pose of a camera included in a user terminal, comprising: Based on whether the first pose of the camera is effectively predicted in a first query image of the location captured by the user terminal, a prediction pattern is determined for predicting the second pose of the camera in a second query image of the location captured by the user terminal after the first query image; as well as The second pose is determined based on the determined prediction pattern. Determining the second posture includes: When the first pose is effectively predicted, one or more matching frames are selected from a plurality of keyframes included in the pre-generated 3D map and captured at the location as matching candidates for the second query image, using the first pose. Perform local feature matching between a first local feature extracted from the second query image and a second local feature included in the one or more matching frames; and Further, based on the results of the local feature matching, the second pose in the second query image is determined. Determining the second posture further includes: If the first pose is not predicted effectively, a pre-learned neural network is used to determine the global features of the second query image, and By comparing the global features of the second query image with the global features of each of a plurality of keyframes included in a pre-generated 3D map and captured at the location, one or more matching frames are selected from the plurality of keyframes as matching candidates for the second query image. Determining the second posture further includes: If the first pose is predicted effectively but a preset condition is met, then by comparing the global features of the second query image determined using a pre-learned neural network with the global features of each of a plurality of keyframes included in a pre-generated 3D map and captured at the location, one or more frames are selected as matching frames from the plurality of keyframes using the first pose. The preset conditions include at least one of the following: The number of poses effectively predicted by the camera is equal to or greater than a preset threshold number; and the distance difference between the second pose determined according to the prediction mode when the first pose is effectively predicted and the second pose determined according to the prediction mode when the first pose is not effectively predicted is equal to or greater than a preset threshold distance.

2. The method according to claim 1, wherein, Determining the second attitude includes: When the first pose is effectively predicted, the 3D map is generated using multiple keyframes selected from images of the location captured by the location capture device or terminal pose determination device using predetermined criteria. The predetermined criteria include at least one of the time interval and the travel distance of the location capture device or the terminal attitude determination device.

3. The method according to claim 2, wherein, The 3D map is generated by further using an image generated by rotating at least one of the plurality of keyframes at one or more angles.

4. The method according to claim 1, wherein, Determining the second attitude includes: If the first pose is predicted effectively but a preset condition is met, local feature matching is performed between the first local features extracted from the second query image and the second local features of each keyframe included in the pre-generated 3D map and captured at the location; and The result of the local feature matching is used to determine the second pose in the second query image, and Wherein, when calculating the distance between the first local feature and each of the local features that match the first local feature among a plurality of local features included in each of the one or more keyframes, the second local feature is the local feature with the shortest distance.

5. The method according to claim 4, wherein, Performing local feature matching includes: In the matching pairs of the first local feature and the second local feature, local feature matching is performed only on a predetermined number of matching pairs in the order of the shortest calculated distance between the first local feature and the second local feature.

6. The method according to claim 1, further comprising: Together with the second query image obtained from the user terminal, information is obtained regarding whether the first pose was effectively predicted, first pose information, information regarding the last pose of the user terminal predicted when it was not effectively predicted in the previous query image, and the number of images in the previous images in which the pose was determined to be effectively predicted.

7. The method according to claim 1, further comprising: Information about whether the second pose was effectively predicted is obtained from the user terminal, along with the second pose itself, and the number of images in which the pose was effectively predicted is determined among the second query image and the images preceding the second query image.

8. A device for determining the attitude of a terminal, comprising: The transceiver is configured to receive a first query image from the user terminal, capturing the location of the user terminal; as well as The processor is configured to: determine a prediction pattern for predicting a second pose of the camera in a second query image based on whether a first pose of the user terminal's camera is effectively predicted in the first query image, and determine the second pose based on the determined prediction pattern, wherein the second query image is received from the user terminal after the first query image and in which the position is captured. Wherein, when the first pose is effectively predicted, the processor is further configured to: use the first pose to select one or more matching frames from a plurality of keyframes included in a pre-generated 3D map and captured at the location as matching candidates for the second query image. Wherein, when the first pose is not effectively predicted, the processor is further configured to: determine global features of the second query image using a pre-learned neural network, and select one or more matching frames as matching candidates from the plurality of keyframes by comparing the global features of the second query image with the global features of each of the multiple keyframes included in the pre-generated 3D map and captured at the location. When the first pose is effectively predicted but a preset condition is met, the processor is further configured to: select one or more frames from the plurality of keyframes using the first pose as matching frames for the second query image by comparing the global features of the second query image determined using a pre-learned neural network with the global features of each keyframe in a plurality of keyframes included in a pre-generated 3D map and captured at the location. The preset conditions include at least one of the following: the number of poses effectively predicted by the camera is equal to or greater than a preset threshold number; and the distance difference between the second pose determined according to the prediction mode when the first pose is effectively predicted and the second pose determined according to the prediction mode when the first pose is not effectively predicted is equal to or greater than a preset threshold distance.

9. A non-transitory computer-readable storage medium for storing a computer program, said computer program comprising commands for causing a processor to perform the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Camera positioning method and device, terminal and storage medium

    CN110148178A

  • Method for tightly-coupling visual slam, terminal and computer readable storage medium

    WO2019169540A1