Neural network-based localization
The neural network trained by NeRF technology performs descriptor matching and stereoscopic rendering, which solves the problems of high computational cost and insufficient accuracy of existing visual positioning methods, and realizes efficient and accurate device positioning under limited resources.
Patent Information
- Application Number
- CN202380077055.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-09
- Filing Date
- 2023-03-30
- Publication Date
- 2025-07-04
AI Technical Summary
The existing visual positioning method is high in computing cost and memory requirements, and the camera pose estimation is not accurate enough, especially when the observation direction is largely different from the training reference image.
Neural radiation field (NeRF) technology is used to train neural networks, generate descriptor diagrams through the first neural network, and use the second neural network for stereoscopic rendering and matching, combining perspective N-points and random sampling consistency algorithm for pose estimation, reducing dependence on reference three-dimensional graphs and improving the accuracy of pose estimation.
Highly accurate device positioning is achieved under limited computing resources, reducing memory requirements, improving robustness to changes in observation directions, and reducing the probability of positioning failure.
Smart Images

Figure CN120266160A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the positioning of a movable device (e.g., an autonomous vehicle) based on a neural network for processing sensor data provided by a sensor device, the movable device including a sensor device, e.g., a camera device or a LIDAR camera device. Background Art
[0002] For the operation of vehicles such as automobiles, Automated Guided Vehicles (AGVs), and autonomous mobile robots, or other mobile devices such as smartphones, positioning is an important task. For example, in automotive applications, a Light Detection and Ranging (LIDAR) camera sensing system is applied, which includes one or more LIDAR devices for acquiring a time series of 3D point cloud data sets of a sensed object, and one or more camera devices for capturing a time series of 2D images of the object. In an automotive environment, the LIDAR camera sensing system can form an Advanced Driver Assistant System (ADAS).
[0003] Structure-based and learning-based visual positioning methods are well known. Structure-based visual positioning relies on a database of reference images collected for a navigation environment. According to the reference images, for example, through a Structure from Motion (SfM) algorithm, a three-dimensional map of triangulated key points with corresponding descriptors is reconstructed. A positioning algorithm is used to calculate in real time the actual position of a device including a camera in the three-dimensional map according to a query image captured by the camera. A feature vector of dimensions given by a plurality of key points is extracted from the query image and matched with a reference feature vector extracted from the reference image and represented by the three-dimensional map to obtain an estimated camera pose required for positioning.
[0004] This structure-based visual positioning technology provides a relatively accurate camera pose, but has high computational costs and large memory requirements.
[0005] The learning-based visual localization method utilizes a deep neural network trained based on reference images and reconstructed 3D reference maps. Local descriptors can be used to map similar keypoint patches to clusters in the feature space. The local descriptors can be general, but instead of being predefined, they are learned by the deep neural network (see M. Jahrer et al., "Learning Local Descriptors for Recognition and Matching", Winter Workshop on Computer Vision, Moravske Toplice, Slovenia, February 4 - 6, 2008; P. Napoletano, "Visual Descriptors for Content-Based Remote Sensing Image Retrieval", International Journal of Remote Sensing, 2018, 39:5, pp. 1343 - 1376; A. Moreau et al., "ImPosing: Implicit Pose Encoding for Visual Localization", IEEE / CVF Winter Conference on Applications of Computer Vision (WACV 2023), Waikoloa Village, USA, January 2023, pp. 2893 - 2902).
[0006] However, despite recent engineering advancements, learning-based visual localization methods still seemingly suffer from inaccurate camera pose estimation. For example, for a query image captured from an observation direction that is significantly different from the observation directions of the reference images used for training. Summary of the Invention
[0007] In view of the above, a basic objective of the present application is to provide a technique for accurately positioning a device based on sensor data provided by a sensor device, which can be appropriately implemented in an embedded computing system with limited computing resources.
[0008] The above and other objectives are achieved by the subject matter of the independent claims. Other implementations are apparent from the dependent claims, the description, and the drawings.
[0009] According to a first aspect, a method for determining the position of a device including a sensor device is provided, comprising the steps of: the sensor device acquiring sensor data representing the environment of the device; a first neural network generating a first descriptor map based on the sensor data; inputting input data into a second neural network different from the first neural network based on the sensor data; the second neural network outputting a descriptor based on the input data; performing stereo rendering on the descriptor to obtain a second descriptor map; matching (comparing to find matches) the first descriptor map with the second descriptor map; and determining the pose of the sensor device based on the matching.
[0010] Obviously, the term "neural network" in this article refers to an artificial neural network. The device can be a vehicle, for example, a fully or partially autonomous car, an autonomous mobile robot, or an Automated Guided Vehicle (AGV). The sensor device can be a camera device or a Light Detection and Ranging (LIDAR) device. The camera device can be, for example, a time-of-flight camera, a depth camera, etc., while the LIDAR device can be, for example, a Micro-Electro-Mechanical System (MEMS) LIDAR device, a solid-state LIDAR device, etc.
[0011] The first neural network is trained to extract descriptors / features from query sensor data obtained from the sensor device (e.g., data obtained during the movement of the device in the environment). The second neural network is trained to process data based on the sensor data to obtain local descriptors, for example, local descriptors for each pixel in an image captured by a camera or local descriptors for each point in a three-dimensional point cloud captured by a LIDAR device.
[0012] The first neural network can be a (deep) convolutional neural network (CNN), which can be based on one of the neural network architectures known in the art for learning feature extraction (see examples given in the following literature: M. Jahrer et al., "Learning Local Descriptors for Recognition and Matching", 2008 Winter Conference on Computer Vision, Moravske Toplice, Slovenia, February 4 - 6; P. Napoletano, "Visual Descriptors for Content-Based Remote Sensing Image Retrieval", International Journal of Remote Sensing, 2018, 39:5, pp. 1343 - 1376; A. Moreau et al., "ImPosing: Implicit Pose Encoding for Visual Localization", IEEE / CVF Winter Conference on Applications of Computer Vision (WACV2023), January 2023, Waikoloa Village, USA, pp. 2893 - 2902).
[0013] The second neural network may include a (deep) multi-layer perceptron (MLP) (fully connected feed-forward) neural network. Specifically, the second neural network may be trained based on the neural radiance field (NeRF) technology proposed by B. Mildenhall et al. in the paper titled "Nerf: Representing Scenes as Neural Radiance Fields for View Synthesis" ("Computer Vision - ECCV 2020", 16th European Conference, Glasgow, UK, August 23 - 28, 2020, Springer, Cham, 2020) or any improvement thereof, which has now become a popular view synthesis tool.
[0014] The input data of the Visual NeRF neural network represents the 3D position (x, y, z) and the viewing direction / angle of the camera device The neural network trained by NeRF outputs a neural field including view-related color values (e.g., RGB) and volume density values σ. In this way, the MLP realizes F Θ : and the weights Θ that are optimized during the training process. The neural field can be queried at multiple positions along the ray for stereoscopic rendering (see the detailed description below). The neural network representation of the environment captured by the sensor device is given by the neural field for subsequent stereoscopic rendering, which can produce a rendered image or a point cloud, etc. The difference between the rendered image and the mapped sensor data is minimized during the training process. If 3D point cloud data is being processed instead of an image, the NeRF neural network can be used to render the point cloud. Although in synthetic image generation applications, the camera poses of multiple images input to the NeRF neural network are known, in localization applications, the camera device (or LIDAR device) pose is iteratively determined starting from an initial estimate (pose prior) (see the detailed description below).
[0015] However, by including additional descriptors in the implicit function (see the detailed description below), this neural network trained by NeRF or its variant can be used or included in the second neural network used in the method according to the first aspect. In this case, the input data based on the sensor data includes the three-dimensional position and the viewing direction (collectively forming the pose), while the output data includes the descriptors and the volume density.
[0016] For example, the second neural network can be trained based on the highly evolved NeRF-W technology (appearance code) that reliably accounts for lighting dynamics (see R. Martin-Brualla et al., "NeRF in the Wild: Neural Radiance Fields for Unconstrained Photo Collections", 2021 IEEE International Conference on Computer Vision and Pattern Recognition (Computer Vision and Pattern Recognition, CVPR), June 19 - 25, 2021, pp. 7210 - 7219, Computer Vision Foundation, IEEE, 2021). In an implementation using such a neural network trained with NeRF-W, the input data of the second neural network includes appearance embeddings (optionally, transient embeddings); see also the detailed description below.
[0017] The method for determining the device location by the first neural network and the second neural network according to the first aspect can locate the device very accurately without powerful and expensive computing resources and memory resources. Contrary to the structure-based visual localization techniques in the art, there is no need to match with a 3D map generated from reference sensor data that requires a large amount of memory space, but the matching process is based on the output of a trained neural network that can be implemented with relatively low memory requirements.
[0018] In an implementation of the method according to the first aspect, the descriptor is independent of the viewing direction (or viewpoint) of the sensor device. This property is similar to the property of volume density, but different from the color output by a neural network trained with NeRF. Since the descriptor is independent of the viewing direction, the positioning failures caused by the significant difference between the textureless content or reference data obtained according to the viewing direction and the query sensor data obtained during the positioning process can be significantly reduced.
[0019] According to an implementation, the descriptor represents the local content of the input data and the 3D position of the data points (pixels in an image or points in a point cloud) of the input data. This implementation does not rely on features including key points and associated descriptors, but can use each point of the input data, which can improve the accuracy of the matching result.
[0020] According to one implementation, the method according to the first aspect or any of its implementations further includes: the second neural network obtains a depth map, wherein the matching of the first descriptor map and the second descriptor map is based on the obtained depth map. The depth map is obtained by performing volume rendering on the values of the volume density output by the second neural network. The information of the depth map can be used to avoid the matching of the descriptor structures of the descriptor maps that are geometrically far apart from each other (at least a certain predetermined distance threshold), that is, to ensure that the descriptor structure of one descriptor map in the descriptor maps that are considered geometrically far apart is not similar to the descriptor structure of another descriptor map in the descriptor map.
[0021] According to one implementation, the pose of the sensor device is iteratively determined starting from a pose prior (an initial estimate of the pose). The iteration can be performed based on the Perspective-N-Point (PnP) method in combination with the Random Sample Consensus (RANSAC) algorithm to obtain a robust estimate of the pose by rejecting outlier matches. The pose prior can be obtained by matching the global image descriptor with an image retrieval database or an implicit graph known in the art.
[0022] According to one implementation, the first neural network and the second neural network are jointly trained for the environment based on matching the first training descriptor map obtained by the first neural network with the second training descriptor map obtained by performing volume rendering on the training descriptors output by the second neural network. When the neural networks are jointly trained in this way, it can be ensured that during the matching process in actual localization, the corresponding descriptor structures are recognized as matching each other.
[0023] Specifically, according to this implementation, for the environment / scene where localization is to be performed, descriptor / feature extraction is learned. In this implementation, the first neural network does not use an off-the-shelf feature extractor, but is trained for a specific scene, which may further improve the accuracy of pose estimation and thus the accuracy of the localization result.
[0024] The implementation of the method described in the first aspect may include training a first neural network and a second neural network. According to one implementation, the sensor device is the camera device; the method further includes: jointly training the first neural network and the second neural network for the environment, wherein the joint training includes: training the camera device to obtain training image data of different training poses of the training camera device; inputting training input data into the first neural network based on the training image data, and inputting training pose data according to the different training poses of the training camera device into the second neural network based on the training image data; the first neural network outputs a first training descriptor map based on the training input data; the second neural network outputs training color data, training volume density data, and training descriptor data; rendering the training color data, the training volume density data, and the training descriptor data to respectively obtain a rendered training image, a rendered training depth map, and a rendered second training descriptor map. In addition, the method according to this implementation includes: minimizing a first objective function representing the difference between the rendered training image and a corresponding pre-stored reference image, or maximizing a first objective function representing the similarity between the rendered training image and a pre-stored reference image; minimizing a second objective function representing the difference between the first training descriptor map and a corresponding rendered second training descriptor map, or maximizing a second objective function representing the similarity between the first descriptor map and a corresponding rendered second training descriptor map.
[0025] This process can perform efficient joint training on the first neural network and the second neural network based on the pose of the camera device estimated by matching descriptor maps with each other (i.e., based on mutually matching descriptor maps) for accurate positioning.
[0026] For the case where the sensor device is a LIDAR device, a similar training process can be performed. In this implementation, the method includes: jointly training the first neural network and the second neural network for the environment, where the joint training includes: training the LIDAR device to obtain training three-dimensional point cloud data of different training poses of the training LIDAR device; inputting training input data into the first neural network based on the training three-dimensional point cloud data, and inputting training pose data according to the different training poses of the training LIDAR device into the second neural network based on the training three-dimensional point cloud data; the first neural network outputs a first training descriptor map based on the training input data; the second neural network outputs training volume density data and training descriptor data; rendering the training volume density data and the training descriptor data to respectively obtain a rendered training depth map and a rendered second training descriptor map. In addition, the method according to this embodiment includes: minimizing a first objective function representing the difference between the rendered training depth map and a corresponding pre-stored reference depth map, or maximizing a first objective function representing the similarity between the rendered training depth map and the pre-stored reference depth map; minimizing a second objective function representing the difference between the first training descriptor map and a corresponding rendered second training descriptor map, or maximizing a second objective function representing the similarity between the first descriptor map and a corresponding rendered second training descriptor map.
[0027] According to another implementation, the above training process further includes: applying a loss function based on the rendered training depth map to inhibit the minimization or maximization of the second objective function for a data point of a first training descriptor map in the first training descriptor maps, where the geometric distance between the data point of the first training descriptor map in the first training descriptor maps and a data point of a corresponding rendered second training descriptor map in the rendered second training descriptor maps is greater than a predetermined threshold. Considering this loss function can avoid comparing different descriptor structures of the descriptor maps with each other, which do not actually represent common features of the environment.
[0028] At least one step of the method according to the first aspect and its implementations can be performed at a device site of an embedded computing system with limited computing resources, or at a remote site provided with data required for the processing / locating device.
[0029] According to a second aspect, there is provided a computer program product including computer-readable instructions, which are used to execute or control the steps of the method according to the first aspect or any of its implementations when the computer-readable instructions run on a computer. The computer can be installed in a device (such as a vehicle).
[0030] According to a third aspect, a positioning device is provided. The method according to the first aspect and any of its implementation manners can be implemented in the positioning device according to the third aspect. The positioning device according to the third aspect and any of its implementation manners can provide the same advantages as those described above.
[0031] The positioning device according to the third aspect includes: a sensor device (e.g., a camera device or a Light Detection and Ranging (LIDAR) device) for acquiring sensor data representing the environment of the sensor device; a first neural network for generating a first descriptor map based on the sensor data; a second neural network different from the first neural network for outputting a descriptor based on input data, where the input data is based on the sensor data. The positioning device according to the third aspect further includes a processing unit configured to perform the following operations: perform stereo rendering on the descriptor to obtain a second descriptor map; match the first descriptor map with the second descriptor map; and determine the pose of the sensor device based on the matching.
[0032] According to one implementation manner, the descriptor is independent of the viewing direction of the sensor device. The descriptor can represent the local content of the input data and the three-dimensional positions of the data points of the input data.
[0033] According to one implementation manner, the second neural network is further configured to obtain a depth map, and the processing unit is further configured to match the first descriptor map with the second descriptor map based on the depth map.
[0034] According to one implementation manner, the processing unit is further configured to iteratively determine the pose of the sensor device starting from a pose prior.
[0035] According to one implementation manner, the first neural network and the second neural network are jointly trained for the environment based on matching a first training descriptor map obtained through the first neural network with a second training descriptor map obtained by performing stereo rendering on the training descriptor output by the second neural network.
[0036] According to a fourth aspect, a vehicle is provided, including the positioning system according to the third aspect or any of its implementation manners. For example, the vehicle is (specifically, fully or partially autonomous) an automobile, an autonomous mobile robot, or an Automated Guided Vehicle (AGV).
[0037] Details of one or more embodiments are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the specification, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings, in which:
[0039] Figure 1 Techniques for positioning a vehicle equipped with a sensor device according to an embodiment are shown.
[0040] Figure 2 A neural network architecture for rendering a descriptor map included in a positioning device according to an embodiment is shown.
[0041] Figure 3 Techniques for training a neural network suitable for a positioning device according to an embodiment are shown.
[0042] Figure 4 Shows the use of a neural network trained using techniques such as those shown in Figure 3 to position a device.
[0043] Figure 5 is a flowchart of a method for positioning a device equipped with a sensor device according to an embodiment.
[0044] Figure 6 A positioning device according to an embodiment is shown. DETAILED DESCRIPTION
[0045] A method for positioning a device (e.g., a vehicle) equipped with a sensor device (e.g., a camera device or a LIDAR device) is provided herein. Specifically, the method may be based on a Neural Radiance Field (NeRF) scene representation. Although the following description of the embodiments relates to NeRF technology, other techniques for implicit representation of an environment / scene based on neural fields and stereo rendering may be suitably used in place of the embodiments. Highly accurate positioning results can be obtained at a relatively low computational cost.
[0046] Figure 1Illustrates the positioning of a vehicle according to an embodiment. The vehicle navigates 11 in a known environment. The vehicle is equipped with a sensor device (e.g., a camera device or a LIDAR device) and captures sensor data representing the environment (e.g., an image or a 3D point cloud). Specifically, a query sensor data set (e.g., a query image or a query 3D point cloud) is captured 12 and used to position the vehicle in the known environment. The query sensor data set is input into a neural network trained for descriptor extraction (e.g., a deep convolutional neural network (CCN)), and a query descriptor map is generated 13 based on the query sensor data set and the extracted descriptors. The descriptors are independent of the viewing direction. Since the descriptors are independent of the viewing direction, positioning failures caused by textureless content obtained according to the viewing direction or significant differences between the reference data and the query sensor data obtained during the positioning process can be significantly reduced.
[0047] In the art, the extracted descriptors or features are compared with a three-dimensional reference descriptor map generated based on a training sensor data set. When using such a three-dimensional reference descriptor map, the memory requirements are large. Contrary to the art, according to Figure 1 the illustrated embodiment, another neural network is used 14 to provide another rendered descriptor map to match the query descriptor map provided by the neural network trained for descriptor extraction. Input data based on the query sensor data captured by the vehicle's sensor device is input into the other neural network. For example, the estimated three-dimensional position data and viewing direction data of the sensor device associated with the captured query image or 3D point cloud are input into the other neural network, which is used to output descriptors of local positions (features). These descriptors are part of a neural field, which can also include color values and volume density values. The neural field including the local descriptors can be referred to as a neural position feature field.
[0048] According to specific implementation manners, other neural networks include or consist of one or more (deep) multilayer perceptrons (MLPs), that is, the other neural networks include or are fully connected feedforward neural networks, and the fully connected feedforward neural networks are trained based on the Neural Radiance Field (NeRF) technology proposed by B. Mildenhall et al. in the paper titled "Nerf: Representing Scenes as Neural Radiance Fields for View Synthesis" ("Computer Vision - ECCV 2020", 16th European Conference, Glasgow, UK, August 23 - 28, 2020, Springer, Cham, 2020) or any improvement thereof (e.g., Nerf-W technology). For the Nerf-W technology, see R. Martin-Brualla et al., "NeRF in the Wild: Neural Radiance Fields for Unconstrained Photo Collections", 2021 IEEE International Conference on Computer Vision and Pattern Recognition (Computer Vision and Pattern Recognition, CVPR), June 19 - 25, 2021, pp. 7210 - 7219, Computer Vision Foundation, IEEE, 2021. The NeRF initially introduced by B. Mildenhall et al. can obtain a neural network representation of the environment based on color values and spatially correlated volume density values (representing the neural field).
[0049] The input data of the neural network represents 3D positions and viewing directions The neural network trained by NeRF outputs view-related color values (e.g., RGB) and volume density value σ. In this way, the MLP realizes F Θ : and the weights Θ that are optimized during the training process.
[0050] Stereo rendering is based on rays passing through the scene (projected from all pixels in the image). The volume density σ(x, y, z) can be interpreted as the differential probability that a ray terminates on an infinitesimal particle at (x, y, z). By collecting all the volume density values along the ray direction, the cumulative transmittance T(s) along the ray direction can be calculated: The cumulative transmittance T(s) along the ray from the origin 0 to s represents the probability that the ray reaches s along its path without hitting any particles.
[0051] Query the implicit representation (neural field) at multiple positions along the ray, and then combine the obtained samples into an image.
[0052] According to Figure 1In the illustrated embodiment, this type of neural network can be employed. However, according to this embodiment, these neural networks are trained not only to obtain volume density values and color values (if the camera device is used as a sensor device), but also to obtain descriptors (modified NeRF-trained neural networks). Stereo rendering of the descriptors output by the neural network can generate a rendered (reference) descriptor map, which is to be matched (compared) with the query descriptor map. Using other neural networks instead of the reference 3D maps used in the art saves memory space (usually about 1000 times), but can achieve very accurate pose estimation, thereby obtaining a localization result.
[0053] For example, the query descriptor map is matched with the reference descriptor map using cosine similarity 15. For example, if the similarity is higher than a predetermined threshold, and the two descriptors represent the best candidates in both directions in the two descriptor maps (matching each other), then the two descriptors are a match (i.e., the two descriptors are similar to each other).
[0054] As Figure 1 shown, according to the correspondence between the query descriptor map and the rendered descriptor map generated by stereo rendering of the descriptors output by other (modified NeRF-trained) neural networks, the pose (position and viewing angle) of the sensor device can be calculated, as is known in the art. For example, starting from the pose prior (initial estimate) of the sensor device, the actual pose of the query sensor data can be iteratively estimated based on the Perspective-N-Points (PnP) method in combination with the Random Sample Consensus (RANSAC) algorithm 16. The view observed from the pose prior should have overlapping content with the query sensor data to make the matching process feasible. The pose prior can be obtained by matching the global image descriptor with an image retrieval database or an implicit graph known in the art. Similar to the process described by A. Moreau et al. in "ImPosing: Implicit Pose Encoding for Visual Localization" (IEEE / CVF Winter Conference on Applications of Computer Vision (WACV 2023), January 2023, Waikoloa Village, USA, pp. 2893-2902), the obtained pose estimate of the sensor device can be used as a new pose prior, the matching iteration, and the result of the PnP process combined with RANSAC to refine the pose estimate of the sensor device, thereby refining the localization result. Although the classical 3D reference model can only access a limited set of reference descriptors, the reference descriptors can be calculated by other neural networks according to any camera pose, which can improve the accuracy of the localization result.
[0055] Figure 2Shows a neural network architecture for rendering a descriptor map of a positioning device equipped with a camera device according to a specific embodiment. The neural network shown in the figure is trained based on the NeRF-W technology (R. Martin-Brualla et al., "NeRF in the Wild: Neural Radiance Fields for Unconstrained Photo Collections", 2021 IEEE International Conference on Computer Vision and Pattern Recognition (Computer Vision and Pattern Recognition, CVPR), June 19 - 25, 2021, pp. 7210 - 7219, Computer Vision Foundation, IEEE, 2021). Due to the appearance embedding included in the implicit function, this NeRF neural network can adapt to dynamic (lighting) changes in outdoor scenes. As Figure 2 shown, similar to the originally introduced NeRF neural network, the NeRF-W neural network is fed with input data based on sensor data, which is in the form of 3D position (x, y, z)21 and viewing direction d. It can be considered that the NeRF-W neural network includes a first Multilayer Perceptron (MLP)22 and a second MLP 23 including a first logical part 23a and a second logical part 23b. The first MLP22 receives the position (x, y, z) as input data and outputs data including information on the position (x, y, z) input to the second MLP 23. In addition, the first MLP 22 outputs a volume density value σ that does not depend on the viewing direction d. The first part 23a of the second MLP 23 is trained to output a color value RGB of the position (x, y, z) and the viewing direction d input to the second MLP 23. The first part 23a of the second MLP23 also receives an appearance embedding application (app) to account for dynamic (lighting) changes in the captured scene. The first part 23a of the second MLP 23 may or may not receive an instantaneous embedding.
[0056] Based on the (RGB) color value and the volume density value σ, an image can be generated by stereo rendering. According to Figure 2The illustrated embodiment, different from the teachings of R. Martin-Brualla et al. in "NeRF in the Wild: Neural Radiance Fields for Unconstrained Photo Collections" (Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), June 19 - 25, 2021, pp. 7210 - 7219, Computer Vision Foundation, IEEE, 2021), trains the second part 23b of the second MLP 23 to output local descriptors that depend neither on the viewing direction (angle) nor on the location (x, y, z) of the appearance embedding. The rendered (reference) descriptor maps can be generated by stereographically rendering the local descriptors, and these rendered descriptor maps are provided for matching with the query descriptor maps (refer to the description in Figure 1 above) for localization purposes.
[0057] Figure 3 Shows a general training process for a neural network (feature extractor) 31 for extracting descriptors / features from images and a neural network (neural renderer) 32 for rendering (reference) descriptor maps. For example, the first neural network 31 can be a fully convolutional neural network with 8 layers, ReLU activation, and max pooling layers, and the second neural network 32 can be the Figure 2 neural network shown. By defining an optimization objective (total loss function) and utilizing scene geometry information, these two neural networks are jointly trained in a self-supervised manner. Thus, dedicated descriptors of the target scene are obtained, which not only describe the visual content of the viewing point but also the 3D position of the viewing point, enabling better discrimination than the general descriptors provided by off-the-shelf feature extractors. The generated descriptors do not depend on the viewing direction or appearance embedding.
[0058] Input the training (e.g., RGB) image I captured by the camera device into the neural network 31, which is used to output the first descriptor map DesM1 based on the input training image. For the same training image, input the input data based on the training image into the neural network 32. In the Figure 3 illustrated embodiment, the input data is the pose of the camera device that captured the training image I (refer to the combination with Figure 2The described 3D position (x, y, z) and orientation d). For example, the pose can be obtained from the training image I by Structure from Motion (SfM) techniques. The neural network 32 can also use appearance embeddings. The neural network 32 outputs a training descriptor, a volume density value σ (i.e., depth information), and color. Stereo rendering can separately generate a second (rendered training) descriptor map DesM2, a rendered training depth map DepM, and a rendered training image IM. Stereo rendering includes the aggregation of descriptor, volume density values, and color values along virtual camera rays, as known in the art.
[0059] Apply the photometric loss (mean squared error loss) L MSE On the rendered training image IM supervised by the training image I to train the (radiance field) neural network 32, as known in the art. Additionally, the structural dissimilarity loss L SSIM . Apply the descriptor loss L pos 、L neg To the first training descriptor map DesM1 and the second training descriptor map DesM2 for joint training of the two neural networks 31 and 32. L pos Can be applied to maximize the similarity between the first training descriptor map DesM1 and the second training descriptor map DesM2, and L neg Can be applied to ensure that pixel pairs (p1, p2) in pixel p1 of the first training descriptor map DesM1 and pixel pairs (p1, p2) in pixel p2 of the second training descriptor map DesM2, which are geometrically far apart from each other, have different descriptors.
[0060] The regularization loss L TV Can be applied to the rendered training depth map DepM to further improve the quality of the geometric structure learned by the second neural network 32 by smoothing artifacts and limiting artifacts. Minimizing the structural dissimilarity loss L SSIM And the regularization loss L TV Can improve the localization accuracy.
[0061] These losses form the overall objective function that needs to be optimized during the training process. Further details about the losses are shown below.
[0062] In Figure 2 And Figure 3 In the illustrated embodiments, the sensor devices for the localization and training processes are camera devices respectively. In alternative embodiments, the sensor device is a LIDAR device that provides 3D point cloud data instead of images, and Figure 2 And Figure 3 The configurations shown in need to be modified directly.
[0063] Figure 4 An embodiment of a method for positioning a device including a sensor device by a first neural network and a second neural network 42 is shown, wherein the first neural network is used for descriptor / feature extraction, and the second neural network is used to provide a neural field including descriptors. For example, the second neural network 42 can be Figure 2 the neural network shown. For example, in a manner similar to the joint training of the neural networks 31 and 32 shown Figure 3 the first neural network 41 and the second neural network 42 are jointly trained.
[0064] During the actual navigation of a mobile device equipped with a sensor device (e.g., a vehicle equipped with a camera device or a LIDAR device), a query sensor data set SD (e.g., a query image or a query 3D point cloud) is input into the first neural network 41 for descriptor extraction to generate a first descriptor map DesM1. Starting from a pose prior (an initial estimate of the pose) (i.e., position and viewing direction), the second neural network outputs descriptors, volume density values, and color values respectively to generate a second descriptor map DesM2, a depth map DepM, and an (RGB) image IM. The two 2D descriptor maps DesM1 and DesM2 are matched with each other to establish 2D-2D local correspondences. The depth map is used for the second descriptor map, and 2D-3D correspondences can be established through depth information in a third direction. Based on matching the descriptor structures, pose estimation can be calculated 43 by the Perspective-N-Points (PnP) method in combination with the Random Sample Consensus (RANSAC) algorithm, as is known in the art. The estimated sensor pose can be used as a new pose prior. Subsequently, based on this new pose prior, a new descriptor map can be rendered to match the first descriptor map DesM2. The process of pose estimation can be iterated multiple times to improve the accuracy of the estimated pose, thereby improving the positioning accuracy of the device.
[0065] Figure 5 The flowchart of shows an embodiment of a method 50 for determining the position of a device including a sensor device. Method 50 includes: the sensor device acquires S51 sensor data representing the environment of the device. For example, the sensor device is a camera that captures a 2D image of the environment, or a LIDAR device that captures a 3D point cloud representing the environment. Method 50 also includes: the first neural network generates S52 a first descriptor map based on the sensor data. The first neural network can be one of the examples of the first neural network for descriptor / feature extraction described above, for example, Figure 3 the neural network 31 shown or Figure 4 the neural network 41 shown.
[0066] Method 50 further includes: inputting input data into a second neural network different from the first neural network based on sensor data; the second neural network outputs a descriptor (specifically, this descriptor can be independent of the viewing direction) based on the input data. For example, the input data represents an initial estimate or a developed estimate of the pose of the sensor device obtained based on sensor data. The second neural network can be one of the examples of the above-mentioned second neural network. For example, Figure 3 the neural network 32 shown or Figure 4 the neural network 42 shown, and may include Figure 2 the MLP 32a and MLP 32b shown. The first neural network and the second neural network can be jointly trained based on the descriptor loss applied to the training descriptor map provided by the first neural network and the training descriptor map rendered based on the descriptor output by the second neural network to maximize the similarity of the corresponding training descriptor maps.
[0067] Method 50 further includes: performing stereo rendering S55 on the descriptor to obtain a second descriptor map; matching the first descriptor map with the second descriptor map S56. Based on the matching S56, the pose of the sensor device is determined S57 (the descriptor maps match each other to a certain predetermined degree). The details of the specific steps of Method 50 can be similar to the above steps.
[0068] Figure 6 The positioning device 60 according to an embodiment is shown. The positioning device 60 can be installed on a vehicle (e.g., an automobile). The positioning device 60 can be used to perform the method steps of Method 50 described above in combination with Figure 5 The positioning device includes a sensor device 61, for example, a camera that captures a 2D image of the environment of the positioning device, or a LIDAR device that captures a 3D point cloud representing the environment. The sensor data captured by the sensor device 61 is processed by a first neural network 62 and a second neural network 63 different from the first neural network 62. The first neural network is used to generate a first descriptor map based on sensor data (specifically, including descriptors that can be independent of the viewing direction). The first neural network can be one of the examples of the above-mentioned first neural network for descriptor / feature extraction. For example, Figure 3 the neural network 31 shown or Figure 4 the neural network 41 shown.
[0069] The second neural network is used to output a descriptor (which can be independent of the viewing direction) based on input data (e.g., pose prior or developed pose estimate), where the input data is based on sensor data. The second neural network can be one of the examples of the above-mentioned second neural network. For example, Figure 3 the neural network 32 shown or Figure 4 the neural network 42 shown, and may includeFigure 2 The MLP 32a and MLP 32b shown. The first neural network 62 and the second neural network 63 can be jointly trained based on the descriptor loss applied to the training descriptor map provided by the first neural network 62 and the training descriptor map rendered based on the descriptors output by the second neural network 63 to maximize the similarity of the corresponding training descriptor maps.
[0070] In addition, the positioning device includes a processing unit 64 that is used to process the outputs provided by the first neural network 62 and the second neural network 63. The processing unit 64 is used to: perform stereoscopic rendering on the descriptors to obtain a second descriptor map; match the first descriptor map with the second descriptor map; and determine the pose of the sensor device based on this match (the descriptor maps match each other to a certain predetermined degree).
[0071] It should be noted that in Figures 1 to 6 the embodiment shown, the data processing can be entirely performed on the device to be positioned (e.g., a vehicle), thus fully ensuring the privacy of user data. In the case of limited computing power, the data processing can be partially or entirely performed on an external processing unit (server). For example, the query sensor data set is transmitted to the external processing unit, which performs the positioning and notifies the device of its location. In this case, since only the location information of the device is provided, the privacy of the server-side data is ensured. According to another embodiment, the extraction of features / descriptors from the query sensor data can be performed at the device site, and the extracted features / descriptors are transmitted to the external processing unit to perform the remaining steps of the positioning process.
[0072] The embodiments of the above methods and devices can be appropriately integrated in vehicles such as automobiles, Automated Guided Vehicles (AGVs), and autonomous mobile robots to facilitate navigation, positioning, and obstacle avoidance. In an automobile, the embodiments of the above methods and devices can be composed of ADAS. Other embodiments of the above methods and devices can be appropriately implemented in augmented reality applications.
[0073] Description of further detailed information of the specific embodiments described above
[0074] The camera repositioning method estimates the position and orientation of an image captured in a given environment. The method to solve this problem is to match the image with a precomputed map constructed based on data previously collected in the target area.
[0075] The present invention describes the following embodiments: The reference figure is represented by a neural field, which can render consistent local descriptors with 3D coordinates from any viewpoint. Compared with classical sparse 3D models, this has many advantages: dense feature matching can be performed, pose estimation can be iteratively improved, and storage requirements are reduced. The proposed system learns local features in the scene in a self-supervised manner, and the performance of this system is better than related methods, and accurate matches can be established even when the pose prior is far from the actual camera pose.
[0076] Visual localization (i.e., the problem of camera pose estimation in a known environment) can use cameras to build a localization system for various applications such as autonomous driving, robotics, or augmented reality. The goal is to predict the 6-DoF camera pose (translation and orientation) based on 2D camera sensor measurements, which is a problem of projective geometry. The best-performing methods in this field are called structure-based methods, which operate by matching image features with a pre-computed 3D model of the environment represented by a map. Usually, these maps are modeled based on the results of Structure-from-Motion (SfM): a sparse 3D point cloud constructed from a reference image, on which key points are triangulated. The storage requirements of such a 3D model are high, but the information stored in it is still sparse.
[0077] Recently, a new method for implicitly representing scenes has emerged, namely the Neural Radiance Field (NeRF). The scene is implicitly represented by a neural network instead of an explicit representation such as a point cloud, mesh, or voxel grid. The neural network learns the mapping from 3D point coordinates to density and radiance. NeRF is trained using a set of sparse pose images representing a given scene and learns the underlying 3D geometric information without supervision. The generated model is continuous, that is, the radiance of all 3D points in the scene can be calculated, so that a photo-realistic view can be rendered from any viewpoint. Subsequent work has shown that additional modalities (e.g., semantics) can be incorporated into the radiance field and accurately rendered. Through the neural field, the dense information of the scene can be stored in a compact manner: the neural network weights only represent a few megabytes.
[0078] We propose to introduce local descriptors in the NeRF implicit scheme and use the resulting model as a localization map. We train both a CNN feature extractor and a neural renderer to provide consistent scene-specific descriptors in a self-supervised manner. Unlike radiance, the rendering of descriptors depends neither on the viewing direction nor on the image appearance. The advantage of this scheme is to learn repeatable features, enabling accurate matching under extreme viewpoint changes. By leveraging different advanced neural rendering techniques, we make the model computationally simpler with the multi-resolution hash encoding of Instant-NGP and adapt the model to dynamic outdoor scenes with the appearance embedding of Nerf-W. During training, we utilize the 3D information learned by the radiance field in a metric learning optimization objective that does not require image pairs or supervised pixel correspondences on pre-computed 3D models.
[0079] Finally, we demonstrate that these features can solve visual relocalization tasks through simple structure-based methods based on sparse feature matching and / or dense feature alignment. The commonly used sparse 3D models obtained from Structure from Motion are replaced by a neural field from which dense reference features can be queried from any camera pose, with very compact storage requirements compared to the state of the art.
[0080] Camera Relocalization: Estimating the 6-DoF camera pose of a query image in a known environment has been addressed in the literature. Structure-based methods compare local image features with existing 3D models. The classical pipeline mainly consists of two steps: First, the nearest reference images are retrieved from the query image using global image descriptors. Then, based on the extracted local features, these priors are refined through geometric reasoning. This pose refinement can be based on sparse feature matching, direct feature alignment, or relative pose regression. Structure-based relocalization methods are the most accurate but require storing and leveraging a large amount of map information, meaning high computational costs and memory footprint. Even though compression methods have been developed, storing dense maps remains a challenging task. An effective alternative is absolute pose regression, which associates the query image and the associated camera pose in a single neural network forward pass but is less accurate. Scene coordinate regression can learn the mapping between image pixels and 3D coordinates, enabling accurate pose calculation through the Perspective-n-Point method but with poor scalability in large environments. Our scheme refines camera pose priors in a structure-based method but replaces the traditional 3D model with a compact implicit representation.
[0081] Localization Using Neural Scene Representations: Recently, NeRF models and related models have been used in various ways to improve localization methods. iNeRF iteratively optimizes camera poses by minimizing the NeRF photometric error. LENS improves the accuracy of absolute pose regression methods by using newly rendered views from NeRFs uniformly distributed on the map as additional training data. iMAP and NICE-SLAM obtain results more competitive than state-of-the-art methods via neural implicit maps for solving RGB-D SLAM problems. ImPosing solves kilometer-scale localization problems by measuring the similarity between global image descriptors and implicit camera pose representations. The feature query network related to our work learns descriptors in implicit functions for relocalization. Its model is trained only on precomputed sparse 3D point clouds, using off-the-shelf feature extractors as supervision, and learns how descriptors change with the viewpoint, enabling iterative pose refinement. We take the opposite approach: enabling the radiance field to learn 3D geometric information and viewpoint-invariant descriptors simultaneously. The latter is learned in a self-supervised manner without any additional training data. To our knowledge, no one has proposed learning visual localization descriptors in neural radiance fields without supervision before.
[0082] Learning-Based Local Feature Descriptors: Local descriptors can provide useful descriptions of regions of interest, enabling accurate correspondences to be established between image pairs depicting the same scene. Although handcrafted descriptors such as SIFT and SURF have been highly successful, the focus in recent years has shifted to learning feature extraction from large amounts of visual data. Many learning-based schemes rely on training siamese convolutional networks using pairs or triplets of images / patches supervised by correspondences. By augmenting two versions of the same image, the feature extractor can be trained without annotated correspondences. SuperPoint uses an isomorphic graph, while Novotny et al. utilize image warping. The method we propose follows a different path to learn repeatable descriptors: we constrain the feature extractor to provide the same descriptor map as the volumetric neural renderer. Since the neural renderer employs a ray marching scheme and is geometrically consistent by design, this approach allows us to learn dense scene-specific descriptors without annotated correspondences.
[0083] Our method estimates the 6-DoF camera pose (i.e., 3D translation and 3D orientation) of a query image in an already visited environment. We first train the module in an offline step using a set of reference images with corresponding poses captured in advance in the region of interest. Since we learn the scene geometry during training, a 3D model of the scene is not necessary.
[0084] 1. Neural Rendering of Local Descriptors:
[0085] NeRF can render views from any camera pose in a given scene while being trained using only a sparse set of observations. Given the camera poses with known intrinsics, 2D pixels are backprojected into the 3D scene by ray marching. The density σ and RGB color c of each point p=(x, y, z) along the ray are evaluated by an MLP. The final pixel color of the pixel is computed by volume rendering along the ray, which is differentiable and enables the training of the whole system by minimizing the photometric error of the rendered image.
[0086] NeRF assumes that the lighting in the scene remains constant over time, but this does not hold in many real-world scenarios. NeRF-W overcomes this limitation by modeling the appearance using appearance embeddings that control the appearance of each rendered view (see also Figure 2 ). Another limitation of this neural scene representation is the computation time: rendering an image requires H×W×N evaluations of an 8-layer MLP, where N is the number of points sampled per ray. Instant-NGP proposes using multi-resolution hash encoding to accelerate the process by storing local features in a hash table, which are then processed by a smaller MLP (compared to the MLP in NeRF), significantly reducing the training time and inference time.
[0087] Descriptor - NeRF: Our neural renderer combines the above three techniques to efficiently render dynamic scenes. However, our main goal is not photo-realistic rendering, but features that match new observations. While it is possible to align a query image with a NeRF model by minimizing the photometric error, this method lacks robustness to changes in lighting. Instead, we propose adding local descriptors, i.e., D-dimensional latent vectors that describe the visual content of regions of interest in the scene, as an additional output of the radiance field function. Unlike the rendered colors, we model these descriptors as being independent of the viewing direction d and appearance vectors, and confirm below that this makes the matching process more robust. Similar to the colors, the 2D descriptors of the camera rays are also aggregated by the common volume rendering formula applied to the descriptors of each point along the ray. The neural renderer architecture is shown in Figure 2 and the training process will be explained in the next section.
[0088] 2. Training: Self-Supervised Feature Extraction in Neural Fields
[0089] Motivation:While the neural renderer described above can represent the map of our relocalization method, we also need to extract features from the query image. The simple solution proposed by FQN is to use an off-the-shelf pre-trained feature extractor (e.g., SuperPoint or D2-Net) and train the neural renderer to memorize the observed descriptors according to the viewing direction. Instead, we propose to jointly train the feature extractor and the neural renderer by defining an optimization objective that exploits the scene geometry information. We obtain dedicated descriptors for the target scene, which not only describe the visual content of the observation point but also the 3D position of the observation point, enabling better discrimination than general descriptors.
[0090] The training process is as Figure 3 shown. A training sample is a reference image with the corresponding camera pose. On the one hand, the image is processed by the feature extractor to obtain the descriptor map \(F_{I}\). On the other hand, we use the camera intrinsics to sample each pixel point along the ray, calculate the density, color, and descriptor of each 3D point, and finally perform stereo rendering to obtain the RGB view \(C_{R}\), the descriptor map \(F_{R}\), and the depth map \(D_{R}\).
[0091] Feature extraction: Our feature extractor is a simple fully convolutional neural network with 8 layers, ReLU activation, and max pooling. The input is an RGB image \(I\) of size \(H\times W\), and it generates a dense descriptor map \(F_{I}\) of size \(H / 4\times W / 4\times D\).
[0092] Learning the radiation field: Similar to NeRF, we use the mean squared error loss \(L\) MSE between \(C_{R}\) and the ground truth image to learn the radiance field. When we render the entire (downscaled) image in a single training step, we can exploit the local 2D image structure and optimize the SSIM (denoted as \(L\) SSIM loss), and we observe that clearer images and better results are generated. The localization process uses the depth map to calculate the camera pose, and a better depth map leads to a more accurate pose. Due to the shape-radiance ambiguity, the NeRF model trained with a limited number of training views may produce incorrect depths. We add a regularization loss \(L\) TV to minimize the total variation of the depth of randomly sampled \(5\times5\) image patches, thereby improving the smoothness of the rendered depth map and restricting the artifacts on the rendered depth map. \(L\) SSIM and \(L\) TV improve the localization accuracy and the image reconstruction quality.
[0093] Learning the descriptor:Our goal is to match the descriptors from the feature extractor and the neural renderer. Our self-supervised objective encourages the two models to generate the same features for a given pixel while preventing high matching scores between points that are far apart in the 3D scene. We define a loss function with two terms, L POS and L NEG applied to a pair of descriptor maps, each containing n pixels. We use the cosine similarity, denoted as to measure the similarity between descriptors.
[0094] The first term maximizes the similarity between the descriptor maps F_{I} and F_{R} in the two models;
[0095]
[0096] The second term samples random pixel pairs and ensures that pixel pairs with a large 3D distance have different descriptors.
[0097]
[0098] where t λ (i,j) = max(0, 1 - λ||xyz(i) - xyz(j)||).
[0099] xyz(i) is the 3D coordinate of the point represented by the i-th pixel in the descriptor map. We calculate this coordinate based on the camera parameters of the rendered view and the predicted depth. Note that we do not backpropagate the gradient of this loss to the depth map. λ is a hyperparameter that controls the linear relationship between descriptor similarity and 3D distance. P is a random permutation of pixel indices from 1 to n.
[0100] The proposed self-supervised objective is close to the classical triplet loss function, but we show in Section 4.3 that injecting 3D coordinates into the formula is crucial for learning meaningful descriptors.
[0101] Finally, we optimize the following loss function at each training step:
[0102]
[0103] 3. Visual Localization by Iterative Dense Feature Matching
[0104] This section presents a localization method for estimating the camera pose of a query image from our learning module. An overview of the process is shown as Figure 4 shown. The proposed scheme combines simple and common techniques as a new method. We demonstrate that the high-quality features we learn enable precise localization accuracy when using simple feature matching and pose estimation strategies.
[0105] I. Localization prior: Similar to the related feature matching method, we assume a localization prior, i.e., the camera pose that is relatively close to the query pose. The views observed from the prior should have overlapping visual content with the query image to make the matching process feasible. This prior can be obtained by matching global image descriptors with an image retrieval database or an implicit graph.
[0106] II. Feature extraction: First, we extract dense descriptors from the query image through a CNN and use a neural renderer to extract descriptors and depth from the localization prior.
[0107] III. Dense feature matching: Match the query descriptors with the reference descriptors using cosine similarity. We consider that if the similarity is higher than the threshold theta and if the match represents the best candidate in both directions in the other graph (mutual match), then the two descriptors are a match. Then, we calculate the predicted 3D coordinates of the rendered pixels that have been matched (depending on the camera parameters and depth) and obtain a set of 2D-3D matches.
[0108] IV. Camera pose estimation: We use the Perspective-n-Point method in combination with RANSAC to obtain a robust estimate by rejecting outlier matches.
[0109] V. Iterative pose refinement: While classical 3D models can only access a limited set of reference descriptors, our neural renderer can calculate reference descriptors for any camera pose. Similar to FQN and ImPosing, we can consider the camera pose estimation as a new localization prior and iterate the previously mentioned steps multiple times to refine the camera pose.
[0110] We compared our results with state-of-the-art visual relocalization methods. We used the top 1 reference pose retrieved using DenseVLAD as the localization prior (while others used the top 10 reference images).
[0111] Table 1 Results for the seven-scene dataset
[0112] cm / ° Fire Chess Stairs Red kitchen Office Pumpkin Head Prior 0.34 / 13.2 0.22 / 12.1 0.26 / 15.8 0.29 / 12.0 0.31 / 10.8 0.16 / 15.0 1st iteration 0.08 / 2.8 0.02 / 0.7 0.21 / 2.9 0.04 / 1.1 0.04 / 1.0 0.05 / 2.8 2nd iteration 0.06 / 2.2 0.01 / 0.5 0.28 / 4.1 0.03 / 0.8 0.03 / 0.8 0.03 / 2.5 3rd iteration 0.05 / 1.9 0.01 / 0.4 0.29 / 4.4 0.02 / 0.8 0.03 / 0.8 0.03 / 2.3
[0113] Table 2 Results for the Cambridge Landmarks dataset
[0114]
[0115]
[0116] We propose to represent visual localization maps in a neural field manner. This enables the representation of dense scenes with a small memory footprint, learning local features for target regions without supervision, and outperforming related visual localization methods. The proposed pipeline should be compatible with future improvements in the field of neural rendering to scale these models to larger scenes.
[0117] All the embodiments discussed above are not intended to be limiting, but rather to illustrate the features and advantages of the present invention. It should be understood that the above-mentioned features, in part or in whole, can also be combined in different ways.
Claims
1. A method (50) for determining the position of a device comprising a sensor device (61), characterized in that, Comprising the following steps: The sensor device (61) acquires (S51) sensor data representing the environment of the device; The first neural network (31, 41) generates (S52) a first descriptor map based on the sensor data; Based on the sensor data, input data is input (S53) into a second neural network (32, 42, 63) different from the first neural network (31, 41, 62); The second neural network (32, 42, 63) outputs (S54) a descriptor based on the input data; The descriptor is stereorendered (S55) to obtain a second descriptor map; The first descriptor map is matched with the second descriptor map (S56); Based on the matching, the pose of the sensor device (61) is determined (S57).
2. The method (50) according to claim 1, wherein The descriptor is independent of the viewing direction of the sensor device (61).
3. The method (50) according to any one of the preceding claims, characterized in that, The descriptor represents the local content of the input data and the three-dimensional positions of the data points of the input data.
4. The method (50) according to any one of the above claims, characterized in that, Further comprising: The second neural network (32, 42, 63) acquires a depth map, wherein the matching of the first descriptor map and the second descriptor map is based on the acquired depth map.
5. The method (50) according to any one of the preceding claims, characterized in that, The pose of the sensor device (61) is determined iteratively starting from a pose prior.
6. The method (50) according to any one of the above claims, characterized in that Based on matching a first training descriptor map obtained by the first neural network (31, 41, 62) with a second training descriptor map obtained by stereorendering the training descriptor output by the second neural network, the first neural network (31, 41, 62) and the second neural network (32, 42, 63) are jointly trained for the environment.
7. The method (50) according to any one of the preceding claims, characterized in that, The device is a vehicle, and the sensor device (61) is a camera device or a Light Detection and Ranging (LIDAR) device.
8. The method (50) according to claim 7, characterized in that, The sensor device (61) is the camera device; further comprising jointly training the first neural network (31, 41, 62) and the second neural network (32, 42, 63) for the environment, wherein the joint training comprises: Training the camera device to acquire training image data of different training poses of the training camera device; Based on the training image data, training input data is input into the first neural network (31, 41, 62), and based on the training image data, training pose data according to the different training poses of the training camera device is input into the second neural network; The first neural network (31, 41, 62) outputs a first training descriptor map based on the training input data; The second neural network (32, 42, 63) outputs training color data, training volume density data, and training descriptor data; Rendering the training color data, the training volume density data, and the training descriptor data to respectively obtain a rendered training image, a rendered training depth map, and a rendered second training descriptor map; Minimize a first objective function representing the difference between the rendered training image and the corresponding pre-stored reference image, or maximize a first objective function representing the similarity between the rendered training image and the pre-stored reference image; Minimize a second objective function representing the difference between the first training descriptor map and the corresponding rendered second training descriptor map, or maximize a second objective function representing the similarity between the first descriptor map and the corresponding rendered second training descriptor map.
9. The method (50) according to claim 7, characterized in that, The sensor device (61) is the LIDAR device; further comprising jointly training the first neural network and the second neural network for the environment, wherein the joint training includes: Training the LIDAR device to obtain training three-dimensional point cloud data of different training poses of the training LIDAR device; Inputting training input data into the first neural network based on the training three-dimensional point cloud data, and inputting training pose data according to the different training poses of the training LIDAR device into the second neural network based on the training three-dimensional point cloud data; The first neural network outputs a first training descriptor map based on the training input data; The second neural network outputs training volume density data and training descriptor data; Rendering the training volume density data and the training descriptor data to respectively obtain a rendered training depth map and a rendered second training descriptor map; Minimize a first objective function representing the difference between the rendered training depth map and the corresponding pre-stored reference depth map, or maximize a first objective function representing the similarity between the rendered training depth map and the pre-stored reference depth map; Minimize a second objective function representing the difference between the first training descriptor map and the corresponding rendered second training descriptor map, or maximize a second objective function representing the similarity between the first descriptor map and the corresponding rendered second training descriptor map.
10. The method (50) according to claim 8 or 9, characterized in that, Further comprising: Applying a loss function based on the rendered training depth map to suppress the minimization or maximization of the second objective function for a data point of a first training descriptor map in the first training descriptor map, wherein the geometric distance between the data point of the first training descriptor map in the first training descriptor map and the data point of the corresponding rendered second training descriptor map in the rendered second training descriptor map is greater than a predetermined threshold.
11. A computer program product comprising computer-readable instructions, characterized in that, When the computer-readable instructions are run on a computer, the computer program product is used to execute or control the steps of the method (50) according to any one of the above claims.
12. A positioning device (60), characterized in that, Comprising: A sensor device (61) for obtaining sensor data representing the environment of the sensor device (61); A first neural network (31, 41, 62) for generating a first descriptor map based on the sensor data; A second neural network (32, 42, 63) different from the first neural network (31, 41, 62) for outputting a descriptor based on input data, wherein the input data is based on the sensor data; A processing unit (64) for performing the following operations: Stereo-render the descriptor to obtain a second descriptor map; Match the first descriptor map with the second descriptor map; Determine the pose of the sensor device (61) based on the matching; 13. The positioning device (60) according to claim 12, characterized in that, The descriptor is independent of the viewing direction of the sensor device (61).
14. The positioning device (60) according to claim 12 or 13, characterized in that, The descriptor represents the local content of the input data and the three-dimensional positions of the data points of the input data.
15. The positioning device (60) according to any one of claims 12 to 14, characterized in that, The second neural network (32, 42, 63) is further configured to obtain a depth map, and the processing unit (64) is further configured to match the first descriptor map with the second descriptor map based on the depth map.
16. The positioning device (60) according to any one of claims 12 to 15, characterized in that, The processing unit (64) is further configured to iteratively determine the pose of the sensor device (61) starting from a pose prior.
17. The positioning device (60) according to any one of claims 12 to 16, characterized in that Based on matching a first training descriptor map obtained by the first neural network (31, 41, 62) with a second training descriptor map obtained by stereo-rendering a training descriptor output by the second neural network, the first neural network (31, 41, 62) and the second neural network (32, 42, 63) are jointly trained for the environment.
18. The positioning device (60) according to any one of claims 12 to 17, characterized in that, The sensor device (61) is a camera device or a Light Detection and Ranging (LIDAR) device.
19. A vehicle, characterized in that, Comprising the positioning device (60) according to any one of claims 12 to 18.
Citation Information
Cited By
Method for evaluating uncertainty of wind tunnel balance calibration system based on weight loading
CN121595155A