Visual positioning method, apparatus, device, and medium
By using a multi-camera visual localization method, and matching the first global descriptor of the multi-camera at the current moment with the second global stitched descriptor at historical moments, the problem of insufficient visual localization accuracy of monocular cameras is solved, and high-precision visual localization effect is achieved.
Patent Information
- Application Number
- CN202310871351.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-14
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2043-07-14
AI Technical Summary
Existing monocular camera-based visual positioning methods have low accuracy and cannot meet the requirements for high-precision visual positioning.
Multiple images are acquired using a multi-view camera at the current moment. The first global descriptor corresponding to each image is determined and matched with the second global stitching descriptor in the pre-constructed visual sparse map. The scene pose data of the local scene in the target scene is determined by stitching and dot product calculation.
It improves the accuracy and efficiency of visual positioning, and realizes fast, low-cost, high-precision visual positioning.
Smart Images

Figure CN119313728B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of visual positioning technology, and in particular to a visual positioning method, apparatus, device and medium. Background Technology
[0002] Visual positioning technology is required for positioning in many fields such as virtual reality (VR), autonomous driving, and robot navigation. For example, when a user wears VR glasses, it is necessary to locate the pose of the local scene the user sees in the real scene in real time, and then render the local scene based on the pose of the local scene in the real scene so that the user can view the rendered image of the local scene.
[0003] Current visual localization methods generally involve acquiring scene images using a monocular camera, constructing a scene map based on these images, extracting keyframes and matching features from the constructed scene map, and then using Plug-and-Play (PnP) technology for pose calculation to achieve visual localization. However, monocular camera-based localization methods have low accuracy and cannot meet the requirements for high-precision visual localization. Therefore, proposing a high-precision visual localization method is a pressing technical problem that needs to be solved. Summary of the Invention
[0004] In order to solve the above-mentioned technical problems, or at least partially solve the above-mentioned technical problems, this disclosure provides a visual positioning method, apparatus, device and medium.
[0005] In a first aspect, this disclosure provides a visual positioning method, the method comprising:
[0006] When a multi-view camera acquires multiple current images at the current moment, it determines the first global descriptor corresponding to each of the multiple current images, wherein the scene corresponding to the multiple current images is a local scene in the target scene;
[0007] The first global descriptors corresponding to the multiple current images are stitched together to determine the first global stitched descriptor of the multi-view camera at the current time.
[0008] Obtain a visual sparse map corresponding to the target scene, wherein the visual sparse map is pre-built from multiple historical images collected by the multi-view camera at different historical times, and the visual sparse map contains a second global stitching descriptor corresponding to the multi-view camera at different historical times.
[0009] Based on the first global stitching descriptor and the second global stitching descriptors corresponding to different historical moments, multiple historical target images that match the multiple current images are determined from the multiple historical images acquired by the multi-view camera at the different historical moments.
[0010] Based on the multiple current images and the multiple historical target images, the scene pose data of the local scene in the target scene is determined.
[0011] Secondly, this disclosure provides a visual positioning device, the device comprising:
[0012] The first determining module is used to determine the first global descriptor corresponding to each of the multiple current images when the multi-view camera acquires multiple current images at the current moment, wherein the scene corresponding to the multiple current images is a local scene in the target scene;
[0013] The stitching module is used to stitch together the first global descriptors corresponding to the multiple current images respectively, and determine the first global stitched descriptor of the multi-view camera at the current time;
[0014] The acquisition module is used to acquire a visual sparse map corresponding to the target scene. The visual sparse map is obtained in advance by mapping multiple historical images collected by the multi-view camera at different historical times. Furthermore, the visual sparse map contains a second global stitching descriptor corresponding to the multi-view camera at different historical times.
[0015] The second determining module is used to determine, based on the first global stitching descriptor and the second global stitching descriptors corresponding to the different historical times, multiple historical target images that match the multiple current images from the multiple historical images acquired by the multi-view camera at the different historical times.
[0016] The localization module is used to determine the scene pose data of the local scene in the target scene based on the multiple current images and the multiple historical target images.
[0017] Thirdly, this disclosure provides a computer-readable storage medium storing instructions that, when executed on a terminal device, cause the terminal device to implement the above-described method.
[0018] Fourthly, this disclosure provides an apparatus comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the method described above.
[0019] Fifthly, this disclosure provides a computer program product comprising a computer program / instruction that, when executed by a processor, implements the method described above.
[0020] The technical solution provided in this disclosure has at least the following advantages compared with the prior art:
[0021] This disclosure provides a visual positioning method, apparatus, device, and medium. The method includes: when a multi-view camera acquires multiple current images at a current time, determining first global descriptors corresponding to the multiple current images respectively, wherein the scene corresponding to the multiple current images is a local scene in a target scene; stitching the first global descriptors corresponding to the multiple current images respectively to determine a first global stitched descriptor of the multi-view camera at the current time; acquiring a visual sparse map corresponding to the target scene, wherein the visual sparse map is pre-built from multiple historical images acquired by the multi-view camera at different historical times, and the visual sparse map includes second global stitched descriptors corresponding to the multi-view camera at different historical times; based on the first global stitched descriptor and the second global stitched descriptors corresponding to different historical times, determining multiple historical target images matching the multiple current images from the multiple historical images acquired by the multi-view camera at different historical times; and based on the multiple current images and the multiple historical target images, determining scene pose data of the local scene in the target scene. By using the above method, the first global stitching descriptor of the multi-view camera at the current moment and the second global stitching descriptor of the multi-view camera at different historical moments can be used to determine multiple historical target images and multiple current images with matching relationships for visual localization. In this way, visual localization based on multi-view cameras is realized quickly and at low cost, while improving the accuracy of visual localization. Attached Figure Description
[0022] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0023] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 A flowchart illustrating a visual positioning method provided in an embodiment of this disclosure;
[0025] Figure 2 A logical schematic diagram of a visual positioning method provided in an embodiment of this disclosure;
[0026] Figure 3 This is a schematic diagram of the structure of a visual positioning device provided in an embodiment of the present disclosure;
[0027] Figure 4 This is a schematic diagram of the structure of a visual positioning device provided in an embodiment of this disclosure. Detailed Implementation
[0028] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.
[0029] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.
[0030] Figure 1 A flowchart illustrating a visual positioning method provided in an embodiment of this disclosure is shown. Figure 1 As shown, the visual positioning method includes the following steps.
[0031] S110. When the multi-view camera acquires multiple current images at the current moment, determine the first global descriptor corresponding to each of the multiple current images, wherein the scene corresponding to the multiple current images is a local scene in the target scene.
[0032] In this embodiment, scenarios such as users wearing VR glasses for virtual reality experiences, autonomous vehicles performing autonomous driving, and robots navigating all require determining the pose data of the device's local scene within the global scene. Specifically, multiple cameras are used to acquire images of the device's local scene in real time, which are then used as the current images. Each camera captures an image with the same timestamp, and the pose data of the device's local scene within the global scene is calculated based on these multiple current images to achieve visual positioning of the device.
[0033] In this embodiment, Simultaneous Localization and Mapping (SLAM) technology is used to achieve visual localization. SLAM includes a mapping process and a localization process. Specifically, in the mapping process, images of the global scene are first acquired using a camera, and then the global scene images are processed to construct an offline image of the global scene. In the localization process, images of the local scene are first acquired in real time using a camera, and then the pose data of the local scene in the global scene is determined based on the coordinate data of feature points in the local scene images and the coordinate data of feature points in the offline images of the global scene. The pose data is then used as the local scene localization result, thereby achieving the visual localization effect.
[0034] The number of multi-view cameras must be at least two. The multi-view cameras can be installed directly in front of, on top of, or on the left or right sides of devices such as VR glasses, autonomous vehicles, and robots; there are no restrictions on this.
[0035] In this context, the target scene refers to the global scene in which the device is located, while the local scene is a portion of the scene within the field of view of the multi-view camera. For example, when a user wears a VR device for a virtual reality experience, the target scene is the entire living room, while the local scene is the TV wall or the ceiling within the living room. Similarly, during autonomous driving, the global scene is the road scene of XX city, while the local scene is the road scene of YY street.
[0036] Furthermore, after acquiring multiple current images, descriptor calculations are performed on each of the multiple current images to obtain the first global descriptor corresponding to each of the multiple current images.
[0037] In this embodiment of the disclosure, the method for determining the first global descriptor includes, but is not limited to, the following: extracting feature points for each current image to obtain two-dimensional feature points for each current image; and calculating a descriptor for each current image based on the two-dimensional feature points for each current image to obtain the first global descriptor corresponding to each current image.
[0038] Specifically, a preset feature point detection algorithm is used to extract feature points for each current image to obtain two-dimensional feature points for each current image. Then, the neighborhood region of each feature point is divided into blocks, the gradient histogram of each block is calculated, and the first global descriptor corresponding to each current image is calculated based on the gradient histogram.
[0039] Here, the descriptor can be a vector. Optionally, the first global descriptor corresponding to the current image acquired by each camera is denoted as I. p Furthermore, the first global descriptor corresponding to each current image is a 1*n vector.
[0040] Optionally, the preset feature point detection algorithm may include, but is not limited to, ORB feature point detection algorithm, pattern recognition algorithm, etc.
[0041] In other embodiments, a first global descriptor that can characterize each current image can be directly determined without detecting two-dimensional feature points of each current image.
[0042] S120. The first global descriptors corresponding to multiple current images are stitched together to determine the first global stitched descriptor of the multi-view camera at the current time.
[0043] To improve visual positioning accuracy, after determining the first global descriptors corresponding to multiple current images, these first global descriptors are stitched together to form the first global stitched descriptor of the multi-view camera at the current moment. This facilitates subsequent matching of multiple current images with multiple historical images at different historical moments based on the first global stitched descriptor.
[0044] In this embodiment, optionally, S120 specifically includes: stitching together the first global descriptors corresponding to multiple current images according to the camera identifier of each camera to obtain the first candidate stitched descriptor of the multi-camera at the current time; and normalizing the first candidate stitched descriptor to obtain the first global stitched descriptor.
[0045] Here, the camera identifier is used to represent the position or sequence number of each camera. Optionally, based on the camera identifier, the first global descriptors corresponding to multiple current images are concatenated row by row, and after normalization processing, a first global concatenation descriptor set is formed, denoted as G1(I). p If p is 4, then the first global concatenation descriptor set is denoted as G1(I1, I2, I3, I4), that is, G1(I1, I2, I3, I4) is a 1*4n vector.
[0046] S130. Obtain the visual sparse map corresponding to the target scene. The visual sparse map is obtained by pre-constructing multiple historical images collected by the multi-view camera at different historical times. The visual sparse map also includes the second global stitching descriptor corresponding to the multi-view camera at different historical times.
[0047] In this embodiment, before locating the local scene where the device is located, the multi-view camera is used to collect multiple historical images of the target scene in different historical periods, and the multiple historical images are processed to generate a visual sparse map corresponding to the target scene. Then, the visual sparse map is used to locate the local scene where the device is located.
[0048] In this embodiment, the generation of the visual sparse map includes, but is not limited to, the following methods: acquiring multiple historical images captured by a multi-view camera at the same time; extracting feature points from the multiple historical images to obtain two-dimensional feature points for each of the multiple historical images; calculating descriptors for the multiple historical images based on the two-dimensional feature points corresponding to each of the multiple historical images to obtain second global descriptors corresponding to each of the multiple historical images; and generating a visual sparse map containing second global stitching descriptors corresponding to the multi-view camera at different historical times based on the second global descriptors corresponding to the multiple historical images.
[0049] In other embodiments, a second global descriptor that can characterize each historical image can be directly determined without detecting two-dimensional feature points of each historical image.
[0050] Specifically, based on the second global descriptors corresponding to multiple historical images, a visual sparse map is generated containing the second global stitched descriptors corresponding to the multi-camera at different historical moments. This includes: for multiple historical images acquired at each historical moment, stitching the second global descriptors corresponding to the multiple historical images acquired at each historical moment according to the camera identifier of each camera to obtain the second candidate stitched descriptor of the multi-camera at each historical moment; and normalizing the second candidate stitched descriptors corresponding to the multi-camera at different historical moments to obtain the second global stitched descriptor.
[0051] Optionally, the second global descriptor corresponding to the historical images acquired by each camera is denoted as I. q Furthermore, the second global descriptor corresponding to each historical image is a 1*n vector. Based on the camera identifier, the second global descriptors corresponding to multiple historical images acquired at each historical moment are concatenated row-wise and normalized to form a set of second global concatenated descriptors, denoted as G2(I). q If q is 4, then the second global concatenation descriptor set is denoted as G2(I1, I2, I3, I4), that is, G2(I1, I2, I3, I4) is a 1*4n vector.
[0052] In this embodiment, the visual sparse map may further include two-dimensional feature points and their corresponding three-dimensional feature points for each historical image. Specifically, the method for determining the three-dimensional feature points corresponding to the two-dimensional feature points of each historical image includes: first, determining the local descriptor of the two-dimensional feature points in each historical image; then, based on the local descriptor of the two-dimensional feature points and the second global descriptor corresponding to each historical image, projecting the two-dimensional feature points into the global scene to find the corresponding three-dimensional feature points; next, performing triangulation calculation and feature point optimization processing on the two-dimensional feature points and their corresponding three-dimensional feature points in any two historical images, thereby optimizing the pose data of the three-dimensional feature points in the global scene, and using the optimized three-dimensional feature points as the three-dimensional feature points corresponding to the two-dimensional feature points of each historical image.
[0053] The method for determining the local descriptor of the two-dimensional feature point in each historical image includes: dividing the neighborhood region of the two-dimensional feature point in each historical image into blocks, calculating the gradient histogram of each block, and calculating the local descriptor of the two-dimensional feature point in each historical image based on the gradient histogram.
[0054] S140. Based on the first global stitching descriptor and the second global stitching descriptor corresponding to different historical moments, determine multiple historical target images that match multiple current images from multiple historical images acquired by the multi-view camera at different historical moments.
[0055] In this embodiment, after obtaining the first global stitching descriptor and the second global stitching descriptor corresponding to different historical times, the multiple current images are matched with the multiple historical images collected at different historical times based on the first global stitching descriptor and the second global stitching descriptor, so as to find the multiple historical images collected at a certain historical time that match the multiple current images from the multiple historical images, and use them as multiple historical target images.
[0056] To accurately select multiple historical target images, S140 specifically includes: calculating the dot product of the first global stitching descriptor and the second global stitching descriptors corresponding to different historical times; obtaining the second global stitching descriptor that generates the maximum dot product from the second global stitching descriptors corresponding to different historical times; and using the multiple historical images corresponding to the second global stitching descriptor that generates the maximum dot product as multiple historical target images.
[0057] The dot product is the product of the lengths of one vector and its projection onto another vector. It reflects the "matching degree" between two vectors. In other words, a larger dot product indicates a higher degree of matching between multiple historical images representing a given historical moment and multiple current images; conversely, a smaller dot product indicates a lower degree of matching between multiple historical images representing a given historical moment and multiple current images.
[0058] Therefore, by using the dot product as a parameter to measure the degree of matching, multiple historical target images that best match multiple current images can be found from multiple historical images collected at multiple historical moments. This facilitates visual localization based on multiple historical target images and helps improve the visual localization effect.
[0059] S150. Based on multiple current images and multiple historical target images, determine the scene pose data of the local scene in the target scene.
[0060] In this embodiment, the current image captured by each camera is matched with the historical target image, and based on the matching result, the scene pose data of the local scene in the target scene is determined, and the scene pose data is used as the local scene localization result.
[0061] Therefore, by performing visual localization based on multiple historical target images that have the highest matching degree with multiple current images, the local scene localization accuracy is improved.
[0062] This disclosure provides a visual localization method, which includes: when a multi-view camera acquires multiple current images at a current time, determining first global descriptors corresponding to the multiple current images respectively, wherein the scene corresponding to the multiple current images is a local scene in a target scene; stitching the first global descriptors corresponding to the multiple current images respectively to determine a first global stitched descriptor of the multi-view camera at the current time; acquiring a visual sparse map corresponding to the target scene, wherein the visual sparse map is pre-built from multiple historical images acquired by the multi-view camera at different historical times, and the visual sparse map includes second global stitched descriptors corresponding to the multi-view camera at different historical times; based on the first global stitched descriptor and the second global stitched descriptors corresponding to different historical times, determining multiple historical target images that match the multiple current images from the multiple historical images acquired by the multi-view camera at different historical times; and based on the multiple current images and the multiple historical target images, determining the scene pose data of the local scene in the target scene. By using the above method, the first global stitching descriptor of the multi-view camera at the current moment and the second global stitching descriptor of the multi-view camera at different historical moments can be used to determine multiple historical target images and multiple current images with matching relationships for visual localization. In this way, visual localization based on multi-view cameras is realized quickly and at low cost, while improving the accuracy of visual localization.
[0063] In another embodiment of this disclosure, for multiple current images and multiple historical target images, firstly, matching point pairs for each camera are determined, and then, based on the matching point pairs for each camera, scene pose data of the local scene in the target scene is determined by an optimization iteration method.
[0064] Optionally, in this embodiment of the disclosure, S150 specifically includes:
[0065] S1501. Match the two-dimensional feature points in the current image acquired by each camera with the three-dimensional feature points in the historical target image acquired by each camera to obtain the matching point pairs for each camera.
[0066] S1502. Based on the matching point pairs of each camera, determine the scene pose data of the local scene in the target scene.
[0067] Specifically, S1501 includes: matching the two-dimensional feature points in the current image acquired by each camera with the two-dimensional feature points in the historical target image acquired by each camera based on the first local descriptor of the two-dimensional feature points in the current image acquired by each camera and the second local descriptor of the two-dimensional feature points in the historical target image acquired by each camera to determine two-dimensional feature point pairs; and determining matching point pairs based on the two-dimensional feature point pairs and the three-dimensional feature points corresponding to the two-dimensional feature points in the historical target image.
[0068] The determination of the first local descriptor for two-dimensional feature points in the current image includes: dividing the neighborhood region of the two-dimensional feature points in the current image into blocks, calculating the gradient histogram of each block, and calculating the first local descriptor based on the gradient histogram. Similarly, the determination of the second local descriptor for two-dimensional feature points in historical target images acquired by each camera includes: dividing the neighborhood region of the two-dimensional feature points in historical target images acquired by each camera into blocks, calculating the gradient histogram of each block, and calculating the second local descriptor based on the gradient histogram.
[0069] Therefore, images from each camera can be matched based on local descriptors to accurately determine matching point pairs for each camera.
[0070] To avoid reducing the accuracy of feature point matching due to camera shake or other factors, the multi-view camera acquires multiple consecutive image frame sequences at the current moment. The multiple current images are the current image frame sequences in the multiple consecutive image frame sequences. Accordingly, S130 specifically includes: matching the two-dimensional feature points of the multiple current images in the current image frame sequence acquired by each view camera with the three-dimensional feature points in the historical target images acquired by each view camera to determine multiple first image pairs for each view camera, and determining the position of each feature point in each first image pair; for each other image sequence frame in the multiple consecutive image frame sequences other than the current image frame sequence, based on the relative positional relationship between each other image sequence frame and the current image frame sequence and the position of each feature point in the first image pair, determining multiple second image pairs for each view camera, wherein the multiple first image pairs and the multiple second image pairs constitute the image pairs of each view camera.
[0071] The relative positional relationship can be the relative pose between adjacent image frames.
[0072] Therefore, when multiple consecutive image frame sequences are acquired, multiple first image pairs are determined for each camera based on any one image frame sequence. Then, based on the relative positions between adjacent image frame sequences and the positions of each feature point in the matching point pair corresponding to each camera, multiple second image pairs for each camera are determined. Thus, it is unnecessary to perform dot product calculations on all image frame sequences, improving the matching efficiency of the two types of images acquired by each camera. Furthermore, feature point matching is performed on the two-dimensional and three-dimensional feature points in all image pairs of each camera, thereby improving the feature matching accuracy. This avoids deviations in the feature point matching process due to camera shake or other factors, which is beneficial for improving visual positioning accuracy.
[0073] Specifically, S1502 includes: using a preset error equation to calculate the error of the first pose data corresponding to the two-dimensional feature points and the second pose data corresponding to the three-dimensional feature points in multiple matching point pairs, and obtaining the current error parameter; iteratively adjusting the current error value until the current error value meets the iteration stop condition, obtaining the target transformation pose between the two-dimensional feature points and the three-dimensional feature points in multiple matching point pairs, and using the target transformation pose as the scene pose data of the local scene in the target scene.
[0074] During localization, when the multi-camera system is activated to acquire images, the current localization coordinate system of the multi-camera system is determined, and the first pose data corresponding to the two-dimensional feature points is determined within this coordinate system. During mapping, when the multi-camera system is activated to acquire images, a historical mapping coordinate system is determined, and the second pose data corresponding to the three-dimensional feature points is determined within this coordinate system. Optionally, both the current localization coordinate system and the historical mapping coordinate system can be map coordinate systems.
[0075] Specifically, the first pose data corresponding to the two-dimensional feature points and the second pose data corresponding to the three-dimensional feature points in the matching point pair of each camera are used as input data for the preset error equation. The current error value is calculated using the preset error equation, and then the current error value is iteratively adjusted until the current error value is equal to a specific threshold or tends to stabilize. Then, it is determined that the current error value meets the iteration stopping condition, and the target transformation pose is obtained.
[0076] Optionally, the current error value can be a reprojection error value or other forms of error value. A specific threshold can be 0.
[0077] Optionally, the preset error equation can be expressed in the following form:
[0078]
[0079] Where cost is the current error value, which can be the reprojection error value, M is the number of images acquired by the multi-view camera, and N is the number of images acquired by the multi-view camera. M It represents the number of matching point pairs between any current image and historical images that have a matching relationship. π(*) is a mapping function that specifically projects feature points in the image coordinate system onto the pixels of the image using camera intrinsic parameters. and It is the first pose data of the two-dimensional feature points in the matching point pair in the current positioning coordinate system. It is the second pose data of the 3D feature points in the matching point pair in the historical mapping coordinate system. It is the pixel coordinate of the 3D feature point in the matching point pair within the current image where the 2D feature point in the matching point pair is located, t wm and R wm It is a change in the pose of the target.
[0080] Therefore, for the first pose data corresponding to the two-dimensional feature points and the second pose data corresponding to the three-dimensional feature points in multiple matching point pairs of each camera, the first pose data and the second pose data can be processed by using a preset error equation and an iterative method to obtain the scene pose data of the local scene in the target scene, so as to achieve accurate local scene positioning.
[0081] In other cases, for the first camera in a multi-camera setup, a first transformed pose can be determined using a preset error equation. Then, based on the extrinsic parameters of the first camera, the extrinsic parameters of at least one other camera, and the first transformed pose, a second transformed pose corresponding to at least one other camera is calculated. Finally, the first transformed pose and the second transformed pose corresponding to at least one other camera are fused to obtain the scene pose data of the local scene in the target scene.
[0082] In summary, when visual localization is required, the matching point pairs of each camera are first determined. Then, based on the matching point pairs of each camera, the scene pose data of the local scene in the target scene is determined by an optimization and iteration method. This improves the efficiency and accuracy of visual localization.
[0083] In yet another embodiment of this disclosure, the visual positioning method is explained in its entirety. The specific visual positioning method includes a mapping process and a positioning process. For ease of understanding, Figure 2 A logical diagram of the visual positioning method is shown.
[0084] like Figure 2 As shown, the mapping process of the visual positioning method includes:
[0085] S1. Acquire multiple historical images captured by the multi-view camera at different historical moments.
[0086] Specifically, multiple cameras are used to acquire images of the target scene at different historical moments, resulting in multiple historical images.
[0087] S2. Generate a visual sparse map based on multiple historical images.
[0088] In some cases, firstly, multiple historical images are acquired by using a multi-camera system at different historical moments. Then, feature points are extracted from each historical image to obtain two-dimensional feature points. Based on the two-dimensional feature points of each historical image, descriptors are calculated for each historical image to obtain a second global descriptor and a second local descriptor. Based on the second global descriptors of multiple historical images acquired at the same historical moment, the second global stitching descriptor corresponding to the multi-camera system at different historical moments is calculated. Simultaneously, based on the second global descriptor and two-dimensional feature points of each historical image, triangulation calculation and feature point optimization are performed on multiple historical images to determine the three-dimensional feature points corresponding to the two-dimensional feature points of each historical image. Finally, a visual sparse map is generated based on the two-dimensional feature points of each historical image, the second global descriptor and the second local descriptor corresponding to each historical image, the second global stitching descriptor corresponding to the multi-camera system at different historical moments, and the three-dimensional feature points corresponding to the two-dimensional feature points of each historical image.
[0089] Specifically, generating the second global stitching descriptor includes: for multiple historical images acquired at each historical moment, stitching together the second global descriptors corresponding to the multiple historical images acquired at each historical moment according to the camera identifier of each camera, to obtain the second candidate stitching descriptor of the multi-camera at each historical moment; and normalizing the second candidate stitching descriptors corresponding to the multi-camera at different historical moments to obtain the second global stitching descriptor.
[0090] See also Figure 2 The positioning process of the visual positioning method includes:
[0091] S3. Acquire multiple current images captured by the multi-view camera at the current moment.
[0092] Specifically, multiple cameras are used to acquire images of a local scene within the target scene at the current moment, resulting in multiple current images. Optionally, these multiple current images can be current image frame sequences from multiple consecutive image frame sequences, or images from a single image frame sequence.
[0093] S4. Determine the first global descriptor corresponding to each of the multiple current images.
[0094] Specifically, feature points are extracted for each current image to obtain two-dimensional feature points for each current image. Based on the two-dimensional feature points of each current image, descriptor calculation is performed for each current image to obtain the first global descriptor corresponding to each current image.
[0095] S5. The first global descriptors corresponding to multiple current images are stitched together to determine the first global stitched descriptor of the multi-view camera at the current moment.
[0096] Specifically, based on the camera identifier of each camera, the first global descriptors corresponding to multiple current images are stitched together to obtain the first candidate stitched descriptor of the multi-camera system at the current time; the first candidate stitched descriptor is then normalized to obtain the first global stitched descriptor.
[0097] S6. Based on the first global stitching descriptor and the second global stitching descriptor corresponding to different historical times, determine multiple historical target images that match multiple current images from multiple historical images acquired by the multi-view camera at the different historical times.
[0098] Specifically, the dot product of the first global stitching descriptor and the second global stitching descriptors corresponding to different historical times is calculated; the second global stitching descriptor that generates the maximum dot product is obtained from the second global stitching descriptors corresponding to different historical times; and the multiple historical images corresponding to the second global stitching descriptor that generates the maximum dot product are used as multiple historical target images.
[0099] S7. Based on multiple current images and multiple historical target images, determine the scene pose data of the local scene in the target scene.
[0100] Specifically, the two-dimensional feature points in the current image captured by each camera are matched with the three-dimensional feature points in the historical target image captured by each camera to obtain the matching point pairs for each camera; based on the matching point pairs for each camera, the scene pose data of the local scene in the target scene is determined.
[0101] Specifically, the process involves matching two-dimensional feature points in the current image captured by each camera with three-dimensional feature points in the historical target image captured by each camera to obtain matching point pairs for each camera. This includes: matching two-dimensional feature points in the current image captured by each camera with two-dimensional feature points in the historical target image captured by each camera based on the first local descriptor of the two-dimensional feature points in the current image captured by each camera and the second local descriptor of the two-dimensional feature points in the historical target image captured by each camera to determine two-dimensional feature point pairs; and determining matching point pairs based on the two-dimensional feature point pairs and the corresponding three-dimensional feature points in the historical target image.
[0102] Specifically, based on the matching point pairs of each camera, the scene pose data of the local scene in the target scene is determined, including: using a preset error equation to calculate the error of the first pose data corresponding to the two-dimensional feature points and the second pose data corresponding to the three-dimensional feature points in multiple matching point pairs, and obtaining the current error value; iteratively adjusting the current error value until the current error value meets the iteration stop condition, obtaining the target transformation pose between the two-dimensional feature points and the three-dimensional feature points in multiple matching point pairs, and using the target transformation pose as the scene pose data of the local scene in the target scene.
[0103] Based on the same inventive concept as the above-described method embodiments, this disclosure also provides a visual positioning device, with reference to... Figure 3 This is a schematic diagram of the structure of a visual positioning device provided in an embodiment of the present disclosure. The visual positioning device 300 includes:
[0104] The first determining module 310 is used to determine the first global descriptor corresponding to each of the multiple current images when the multi-view camera acquires multiple current images at the current moment, wherein the scene corresponding to the multiple current images is a local scene in the target scene;
[0105] The stitching module 320 is used to stitch together the first global descriptors corresponding to the multiple current images respectively, and determine the first global stitching descriptor of the multi-view camera at the current time;
[0106] The acquisition module 330 is used to acquire a visual sparse map corresponding to the target scene. The visual sparse map is obtained in advance by mapping multiple historical images collected by the multi-view camera at different historical times. Furthermore, the visual sparse map includes a second global stitching descriptor corresponding to the multi-view camera at different historical times.
[0107] The second determining module 340 is used to determine, based on the first global stitching descriptor and the second global stitching descriptors corresponding to the different historical times, multiple historical target images that match the multiple current images from the multiple historical images acquired by the multi-view camera at the different historical times.
[0108] The positioning module 350 is used to determine the scene pose data of the local scene in the target scene based on the plurality of current images and the plurality of historical target images.
[0109] In one optional implementation, the splicing module 320 includes:
[0110] The stitching unit is used to stitch together the first global descriptors corresponding to the multiple current images according to the camera identifier of each camera, so as to obtain the first candidate stitched descriptor of the multi-camera at the current time.
[0111] The normalization unit is used to normalize the first candidate splicing descriptor to obtain the first global splicing descriptor.
[0112] In one optional implementation, the visual sparse map further includes second global descriptors corresponding to the plurality of historical images; the device further includes:
[0113] The first determining module is used to stitch together the second global descriptors corresponding to the multiple historical images acquired at each historical moment, based on the camera identifier of each camera, to obtain the second candidate stitched descriptor of the multi-camera at each historical moment.
[0114] The second determining module is used to normalize the second candidate stitching descriptors corresponding to the multi-view camera at different historical moments to obtain the second global stitching descriptor.
[0115] In one optional implementation, the second determining module 340 includes:
[0116] The calculation unit is used to calculate the dot product between the first global splicing descriptor and the second global splicing descriptors corresponding to the different historical times;
[0117] The first acquisition unit is used to acquire the second global concatenation descriptor that generates the maximum dot product from the second global concatenation descriptors corresponding to the different historical times.
[0118] The second acquisition unit is used to take the multiple historical images corresponding to the second global stitching descriptor that generates the maximum dot product as the multiple historical target images.
[0119] In one optional implementation, the positioning module 350 includes:
[0120] The matching point pair determination unit is used to match the two-dimensional feature points in the current image acquired by each camera with the three-dimensional feature points in the historical target image acquired by each camera to obtain the matching point pairs for each camera.
[0121] The positioning unit is used to determine the scene pose data of the local scene in the target scene based on the matching point pairs of each camera.
[0122] In one optional implementation, the matching point pair determination unit is specifically used for:
[0123] Based on the first local descriptor of the two-dimensional feature points in the current image captured by each camera and the second local descriptor of the two-dimensional feature points in the historical target image captured by each camera, the two-dimensional feature points in the current image captured by each camera are matched with the two-dimensional feature points in the historical target image captured by each camera to determine two-dimensional feature point pairs.
[0124] The matching point pair is determined based on the two-dimensional feature point pair and the three-dimensional feature points corresponding to the two-dimensional feature points in the historical target image.
[0125] In one optional implementation, the positioning unit is specifically used for:
[0126] The error is calculated by using a preset error equation to calculate the error of the first pose data corresponding to the two-dimensional feature points and the second pose data corresponding to the three-dimensional feature points in the multiple matching point pairs, and the current error value is obtained.
[0127] The current error value is iteratively adjusted until the current error value meets the iteration stop condition, thereby obtaining the target transformation pose between the two-dimensional feature points and the three-dimensional feature points in the multiple matching point pairs, and the target transformation pose is used as the scene pose data of the local scene in the target scene.
[0128] This disclosure provides a visual positioning device. When a multi-view camera acquires multiple current images at a current time, it determines a first global descriptor corresponding to each of the multiple current images, wherein the scene corresponding to the multiple current images is a local scene in a target scene; it stitches the first global descriptors corresponding to the multiple current images to determine a first global stitched descriptor of the multi-view camera at the current time; it acquires a visual sparse map corresponding to the target scene, wherein the visual sparse map is pre-built from multiple historical images acquired by the multi-view camera at different historical times, and the visual sparse map includes a second global stitched descriptor corresponding to the multi-view camera at different historical times; based on the first global stitched descriptor and the second global stitched descriptors corresponding to different historical times, it determines multiple historical target images that match the multiple current images from the multiple historical images acquired by the multi-view camera at different historical times; and based on the multiple current images and the multiple historical target images, it determines the scene pose data of the local scene in the target scene. By using the above method, the first global stitching descriptor of the multi-view camera at the current moment and the second global stitching descriptor of the multi-view camera at different historical moments can be used to determine multiple historical target images and multiple current images with matching relationships for visual localization. In this way, visual localization based on multi-view cameras is realized quickly and at low cost, while improving the accuracy of visual localization.
[0129] In addition to the methods and apparatus described above, this disclosure also provides a computer-readable storage medium storing instructions that, when executed on a terminal device, cause the terminal device to implement the visual positioning method of this disclosure.
[0130] This disclosure also provides a computer program product, which includes a computer program / instructions. When the computer program / instructions are executed by a processor, they implement the visual positioning method of this disclosure.
[0131] In addition, this disclosure also provides a visual positioning device, see [link to relevant documentation]. Figure 4 As shown, the visual positioning device may include:
[0132] The device includes a processor 401, a memory 402, an input device 403, and an output device 404. The visual positioning device may have one or more processors 401. Figure 4 Taking a processor as an example. In some embodiments of this disclosure, the processor 401, memory 402, input device 403, and output device 404 can be connected via a bus or other means, wherein, Figure 4 Taking the example of a connection between China and Israel via a bus.
[0133] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing of the visual positioning device by running the software programs and modules stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function, etc. In addition, the memory 402 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. The input device 403 can be used to receive input digital or character information, and generate signal inputs related to the observation user settings and function control of the visual positioning device.
[0134] Specifically in this embodiment, the processor 401 will load the executable files corresponding to the processes of one or more applications into the memory 402 according to the following instructions, and the processor 401 will run the applications stored in the memory 402 to realize the various functions of the above-mentioned visual positioning device.
[0135] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.
[0136] The above are merely specific embodiments of this disclosure, enabling those skilled in the art to understand or implement this disclosure. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to these embodiments, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method of visual positioning, characterized by The method comprises: When the multi-view camera collects a plurality of current images at a current time, determining first global descriptors corresponding to the plurality of current images respectively, wherein the scenes corresponding to the plurality of current images are local scenes in a target scene; stitching the first global descriptors corresponding to the plurality of current images respectively to determine a first global stitched descriptor of the multi-view camera at the current time; obtaining a visual sparse map corresponding to the target scene, wherein the visual sparse map is obtained by previously mapping a plurality of historical images respectively collected by the multi-view camera at different historical times, and the visual sparse map contains second global stitched descriptors respectively corresponding to the multi-view camera at different historical times; based on the first global stitched descriptor and the second global stitched descriptors respectively corresponding to the different historical times, determining a plurality of historical target images matched with the plurality of current images from the plurality of historical images respectively collected by the multi-view camera at the different historical times; based on the plurality of current images and the plurality of historical target images, determining scene pose data of the local scene in the target scene.
2. The method of claim 1, wherein, The method comprises: According to the camera identifier of each camera, the first global descriptors corresponding to the plurality of current images are stitched to obtain a first candidate stitched descriptor of the multi-view camera at the current time; the first candidate stitched descriptor is normalized to obtain the first global stitched descriptor.
3. The method of claim 1, wherein, The visual sparse map further comprises second global descriptors respectively corresponding to the plurality of historical images; the method further comprises: for a plurality of historical images collected at each historical time, according to the camera identifier of each camera, the second global descriptors respectively corresponding to the plurality of historical images collected at the each historical time are stitched to obtain a second candidate stitched descriptor of the multi-view camera at the each historical time; the second candidate stitched descriptors respectively corresponding to the multi-view camera at the different historical times are normalized to obtain the second global stitched descriptors.
4. The method of claim 1, wherein, The method comprises: calculating the dot product of the first global stitched descriptor and the second global stitched descriptors respectively corresponding to the different historical times; from the second global stitched descriptors respectively corresponding to the different historical times, obtaining a second global stitched descriptor generating the maximum dot product; the plurality of historical images corresponding to the second global stitched descriptor generating the maximum dot product are taken as the plurality of historical target images.
5. The method of claim 1, wherein, The method comprises: match the two-dimensional feature points in the current image collected by each camera with the three-dimensional feature points in the historical target image collected by each camera, to obtain a matching point pair of each camera; determine scene pose data of the local scene in the target scene based on the matching point pair of each camera.
6. The method of claim 5, wherein, The matching of the two-dimensional feature points in the current image collected by each camera with the three-dimensional feature points in the historical target image collected by each camera comprises: match the two-dimensional feature points in the current image collected by each camera with the two-dimensional feature points in the historical target image collected by each camera based on a first local descriptor of the two-dimensional feature points in the current image collected by each camera and a second local descriptor of the two-dimensional feature points in the historical target image collected by each camera, to determine a two-dimensional feature point pair; determine the matching point pair based on the two-dimensional feature point pair and the three-dimensional feature points corresponding to the two-dimensional feature points in the historical target image.
7. The method of claim 5, wherein, The determination of the scene pose data of the local scene in the target scene based on the matching point pair of each camera comprises: perform error calculation on first pose data corresponding to the two-dimensional feature point pair and second pose data corresponding to the three-dimensional feature point pair in a plurality of matching point pairs by using a preset error equation, to obtain a current error value; iteratively adjust the current error value until the current error value meets an iteration stop condition, to obtain a target transformation pose between the two-dimensional feature points and the three-dimensional feature points in the plurality of matching point pairs, and take the target transformation pose as the scene pose data of the local scene in the target scene.
8. A visual positioning device, characterized by The apparatus comprises: a first determination module configured to determine first global descriptors corresponding to a plurality of current images respectively when a multi-camera collects the plurality of current images at a current time, wherein a scene corresponding to the plurality of current images is a local scene in a target scene; a stitching module configured to stitch the first global descriptors corresponding to the plurality of current images respectively, to determine a first global stitched descriptor of the multi-camera at the current time; an acquisition module configured to acquire a visual sparse map corresponding to the target scene, wherein the visual sparse map is obtained by previously mapping a plurality of historical images respectively collected by the multi-camera at different historical times, and the visual sparse map contains second global stitched descriptors respectively corresponding to the different historical times; a second determination module configured to determine a plurality of historical target images matched with the plurality of current images from the plurality of historical images respectively collected by the multi-camera at the different historical times based on the first global stitched descriptor and the second global stitched descriptors respectively corresponding to the different historical times; a positioning module configured to determine scene pose data of the local scene in the target scene based on the plurality of current images and the plurality of historical target images.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores instructions, and when the instructions run on a terminal device, the terminal device implements the method in any one of claims 1-7.
10. An apparatus, comprising: comprise: A memory, a processor, and a computer program stored on the memory and loadable on the processor, the processor implementing the method according to any one of claims 1 to 7 when executing the computer program.
11. A computer program product, characterised in that, The computer program product comprises computer programs / instructions which, when executed by a processor, implement the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Visual positioning method and system
CN110084853A
Visual positioning method and device, electronic equipment and storage medium
CN111563922A