Neural network location estimation

A neural network-based method for encoding and matching 3D point clouds across different coordinate systems addresses the inefficiencies in sensor localization, enhancing accuracy and alignment in environmental mapping.

JP7749736B2Active Publication Date: 2025-10-06NOKIA SOLUTIONS & NETWORKS OY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024069433
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2023-04-26
Filing Date
2024-04-23
Publication Date
2025-10-06
Estimated Expiration
2044-04-23

AI Technical Summary

Technical Problem

Existing localization methods for sensors in an environment, such as cameras or RF antennas, struggle to accurately determine the pose and 3D structure of the environment, leading to inefficiencies in connecting the digital and real worlds.

Method used

A neural network-based approach that encodes and matches 3D point clouds across different coordinate systems using encoder and matching descriptor layers, adjusting feature vectors and center points to calculate correlations and transformations, thereby improving alignment accuracy.

Benefits of technology

Enhances the precision of sensor localization by reducing transformation deviation and increasing point correlation, enabling more accurate mapping and localization in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007749736000140
    Figure 0007749736000140
  • Figure 0007749736000141
    Figure 0007749736000141
  • Figure 0007749736000142
    Figure 0007749736000142
Patent Text Reader

Abstract

To provide a device, method, and program for improving localization by a neural network.SOLUTION: A method comprises: encoding a first 3D point cloud of a first coordinate system into a first encoded map comprising first feature center points and first feature vectors; encoding a second 3D point cloud of a second coordinate system into a second encoded map comprising second feature center points and second feature vectors; adapting the first input feature vectors based on the first input feature vectors and the second input feature vectors, to obtain a first joint map; adapting the second input feature vectors based on the first input feature vectors and the second input feature vectors to obtain a second joint map; calculating similarities between the first joint feature vectors and the second joint feature vectors and point correlations based on the similarities; extracting the coordinates of the first joint feature center points and the coordinates of the second joint feature center points if the correlation condition is fulfilled by the similarities; and calculating a transformation between the first coordinate system and the second coordinate system based on the pairs of extracted coordinates.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] FIELD OF THE DISCLOSURE The present disclosure relates to position estimation.

[0002] Abbreviation 3D-3D AP - Access Point RF - Radio Frequency RGB - Red, Green, Blue RGBD - RGB and Depth SLAM - Simultaneous Localization and Mapping RSSI - Received Signal Strength VIO - Visual Inertial Odometry RF - Radio Frequency BoW - Bag of Words SVD - Singular Value Decomposition CV - Computer Vision DL - Deep Learning CSI - Channel State Information AoA - Angle of Arrival ToF - Time of Flight BSSID - Basic Service Set Identifier GNSS - Global Navigation Satellite System DoF - degrees of freedom IMU - Inertial Measurement Unit VPS - Visual Positioning System dB - decibel dBm - decibel milliwatt [Background technology]

[0003] Localization of sensors (such as cameras or RF antennas) in an environment is relevant to the industrial metaverse. The most ubiquitous sensor for creating a localization map is a camera. Given a set of camera images, one can attempt to find the pose of the camera at which the images were taken and find the 3D structure of the environment (e.g., represented as a point cloud). This 3D structure, along with metadata, can serve as a visual localization map. This localization map can find the pose of a query image in a world coordinate frame with six degrees of freedom (6DoF, i.e., translational 3DoF and rotational 3DoF), thus connecting the digital world to the real world. Summary of the Invention

[0004] The aim is to improve upon the prior art.

[0005] According to a first aspect, an apparatus includes one or more processors and a memory storing instructions, the instructions, when executed by the one or more processors, causing the apparatus to input a first 3D point cloud of a first modality and a first grid including one or more first cells into a first encoder layer, where coordinates of points of the first 3D point cloud are expressed in a first coordinate system; and encoding, by the first encoder layer, the first 3D point cloud of the first modality into a first encoded map of the first modality, where the coordinates of points of the first 3D point cloud are expressed in a first coordinate system. a first coded map including, for each of the first cells of the first grid, a respective first feature center point and a respective first feature vector, the coordinates of the first feature center points being expressed in a first coordinate system; and inputting the first coded input map into a first matching descriptor layer of a first hierarchical level, the first coded input map being based on a first coded map of a first modality and including, for each of the first cells, a respective first joint feature center point and a respective first input feature vector, the first joint feature center point and the respective first input feature vector. inputting a second 3D point cloud of the first modality and a second grid including one or more second cells into a second encoder layer, where coordinates of points of the second 3D point cloud are expressed in the second coordinate system; encoding, by the second encoder layer, the second 3D point cloud of the first modality into a second encoded map of the first modality, where the second encoded map of the first modality includes, for each of the second cells of the second grid, a respective second feature center point and a second grid including one or more second cells; and respective second feature vectors, the coordinates of the second feature center points being expressed in a second coordinate system; inputting the second coded input map into a first matching descriptor layer, the second coded input map being based on the second coded map of the first modality and comprising, for each of the second cells, respective second joint feature center points and respective second input feature vectors, the coordinates of the second joint feature center points being expressed in the second coordinate system; and inputting the first input feature vectors,adapting each of the first input feature vectors optionally based on their first feature center points and second input feature vectors, optionally based on their second feature center points, to obtain a first joint map, the first joint map including, for each of the first joint feature center points, a respective first joint feature vector; adapting each of the second input feature vectors, optionally based on the first input feature vectors, optionally based on their first feature center points and second input feature vectors, optionally based on their second feature center points, to obtain a second joint map, the second joint map including, for each of the second joint feature center points, a respective second joint feature vector; and adapting each of the first input feature vectors, optionally based on their first feature center points and second input feature vectors, optionally based on their second feature center points, to obtain a second joint map, the second joint map including, for each of the second joint feature center points, a first joint feature vector of each first cell, by a first optimal matching layer. calculating, by a first optimal matching layer, for each of the first cells and each of the second cells, a first point correlation between each of the first cells and each of the second cells based on the similarity between the first joint feature vector of each of the first cells and some second joint feature vectors of some second cells and based on the similarity between the second joint feature vector of each of the second cells and some first joint feature vectors of some of the first cells; and for each of the first cells and each of the second cells, checking whether at least one of one or more first correlation conditions is satisfied for each of the first cells and each of the second cells, wherein the one or more first correlation conditions include: a first point correlation between each first cell and each second cell is greater than a first correlation threshold for the first tier; or the first point correlation between each first cell and each second cell is among the largest k values ​​of first point correlations, where k is a fixed value; An apparatus is provided that performs the steps of: extracting, for each first cell and each second cell, coordinates of a first joint feature center point of the respective first cell in a first coordinate system and coordinates of a second joint feature center point of the respective second cell in a second coordinate system, and obtaining each pair of extracted coordinates of the first hierarchical layer in response to checking whether at least one of one or more first correlation conditions is satisfied for each first cell and each second cell; and calculating an estimated transformation between the first coordinate system and the second coordinate system based on the pair of extracted coordinates of the first hierarchical layer.

[0006] The first encoder layer, the second encoder layer, and the first matching descriptor layer of the first hierarchy are configured with respective first parameters, and the instructions, when executed by the one or more processors, may further cause the device to perform a step of training the device to determine the first parameters so that a deviation between an estimated transformation and a known transformation between the first coordinate system and the second coordinate system is less than a transformation estimation threshold and so that a first point correlation is increased.

[0007] The first modality may include one of a photographic image, a LIDAR image, and an ultrasound image.

[0008] The first encoder layer may be the same as the second encoder layer.

[0009] The instructions, when executed by the one or more processors, may further cause the apparatus to perform at least the steps of generating a first 3D point cloud based on a sequence of first images of a first modality, wherein a respective position in a first coordinate system is associated with each of the first images, or generating a second 3D point cloud based on a sequence of second images of the first modality, wherein a respective position in a second coordinate system is associated with each of the second images.

[0010] The first coded input map may be the first coded map, the first joint feature center points may be the first feature center points, the second coded input map may be the second coded map, and the second joint feature center points may be the second feature center points.

[0011] The instructions, when executed by the one or more processors, further cause the apparatus to input a third 3D point cloud of a second modality and the first grid into a third encoder layer, where coordinates of points of the third 3D point cloud are expressed in the first coordinate system, the second modality being different from the first modality; and encoding, by the third encoder layer, the third 3D point cloud of the second modality into a third encoded map of the second modality, where the third encoded map of the second modality includes, for each of the first cells of the first grid, , each of the first cells has a respective third feature center point and a respective third feature vector, the coordinates of the third feature center points being indicated in a first coordinate system; inputting the third encoded map to a first neural fusion layer; inputting the first encoded map to the first neural fusion layer; and generating a first encoded input map by the first neural fusion layer based on the first encoded map and the third encoded map, wherein for each of the first cells, the respective first joint feature center points are respectively represented by a respective first Based on the first feature center points of the cells and the third feature center points of the respective first cells, each first input feature vector is calculated based on the first feature vector and the first feature vector, optionally the first feature center points and the third feature vector of each first cell adapted by those first feature center points, optionally the third feature center points, based on the third feature vector and the first feature vector, optionally the third feature center points and the third feature vector of each first cell adapted by those first feature center points, optionally the third feature center points and the third feature vector of each first cell adapted by those first feature center points, optionally the third feature center points inputting a fourth 3D point cloud of the second modality and the second grid into a fourth encoder layer, wherein coordinates of points of the fourth 3D point cloud are represented in a second coordinate system; and encoding, by the fourth encoder layer, the fourth 3D point cloud of the second modality into a fourth encoded map of the second modality, wherein the fourth encoded map of the second modality includes, for each second cell of the second grid, a respective fourth feature center point and a respective fourth feature vector, and the coordinates of the fourth feature center point are represented byThe method performs the steps shown in the second coordinate system, inputting the fourth coded map to a second neural fusion layer, inputting the second coded map to the second neural fusion layer, and generating a second coded input map by the second neural fusion layer based on the second coded map and the fourth coded map, wherein for each of the second cells, the respective second joint feature center points are based on the second feature center points of the respective second cells and the fourth feature center points of the respective second cells, and the respective second input feature vectors are based on the second feature vector and the second feature vector, optionally the second feature center points and the fourth feature vector of the respective second cells adapted by the second feature center points, optionally the fourth feature center points, and the fourth feature vector and the second feature vector, optionally the fourth feature center points and the fourth feature vector of the respective second cells adapted by the second feature center points, optionally the fourth feature center points.

[0012] The second modality may include at least one of a photographic image, an RF fingerprint, a LIDAR image, or an ultrasound image.

[0013] The third encoder layer, the fourth encoder layer, the first neural fusion layer, and the second neural fusion layer are configured with respective second parameters, and the instructions, when executed by the one or more processors, may further cause the apparatus to perform the step of training the apparatus to determine the first parameters and the second parameters so that a deviation between the estimated transformation and the known transformation between the first coordinate system and the second coordinate system is less than a transformation estimation threshold and so that the first point correlation is increased.

[0014] The first image may be taken by a first sensor of the first sensor device, and the second image may be taken by a second sensor of the second sensor device, and the instructions, when executed by the one or more processors, may further cause the apparatus to perform the steps of: taking a plurality of first measurements of a second modality using a third sensor of the first sensor device; taking a plurality of second measurements of the second modality using a fourth sensor of the second sensor device; generating a third 3D point cloud of the second modality based on the plurality of first measurements, positions associated with the first image, and a calibration between the first sensor of the first sensor device and the third sensor of the first sensor device; and generating a fourth 3D point cloud of the second modality based on the plurality of second measurements, positions associated with the second image, and a calibration between the second sensor of the second sensor device and the fourth sensor of the second sensor device.

[0015] The third encoder layer may be the same as the fourth encoder layer.

[0016] The first neural fusion layer may be the same as the second neural fusion layer.

[0017] The instructions, when executed by the one or more processors, further cause the apparatus to input the first encoded map of the first modality and a third grid including one or more third cells into a fifth encoder layer, each of the third cells including two or more of the first cells, and the first grid including two or more of the first cells; and encoding, by the fifth encoder layer, the first encoded map of the first modality into a fifth encoded map of the first modality, the third coded input map comprises, for each of the third cells of the third grid, a respective fifth feature center point and a respective fifth feature vector, the coordinates of the fifth feature center points being represented in the first coordinate system; and inputting the third coded input map into a second matching descriptor layer of a second hierarchical level, the third coded input map being based on the fifth coded map of the first modality and comprising, for each of the third cells, a respective third joint feature center point and a respective third input feature vector, the coordinates of the third joint feature center points being represented in the first coordinate system. a step of inputting a second encoded map of the first modality and a fourth grid including one or more fourth cells into a sixth encoder layer, the fourth cells each including two or more of the second cells, the second grid including two or more of the second cells, the sixth encoder layer encoding the second encoded map of the first modality into a sixth encoded map of the first modality, the sixth encoded map of the first modality being a fourth grid including one or more fourth cells, the sixth encoded map of the first modality being a fourth grid including two or more of the second cells, the sixth encoded map of the first modality being a fourth grid including one or more fourth cells, the sixth encoder layer encoding the second encoded map of the first modality into a sixth encoded map of the first modality, the sixth encoded map of the first modality being a fourth grid including one or more fourth cells, the sixth encoded map of the first modality being a fourth grid including two or more fourth cells, the sixth encoder layer encoding the second encoded map of the first modality into a sixth encoded map of the first modality, the sixth encoded map of the first modality being a fourth grid including two or more fourth cells, the sixth encoder layer encoding the sixth ... sixth encoded map of the first modality, the sixth encoded map of the first modality being a fourth grid including two or more fourth cells inputting a fourth coded input map into a second matching descriptor layer, the fourth coded input map being based on the sixth coded map of the first modality, and the fourth coded input map being based on the sixth coded map of the first modality, the fourth coded input map being based on the sixth coded map of the first modality, the fourth coded input map being based on the sixth coded map of the first modality, the fourth coded input map being based on the sixth coded map of the first modality, the fourth coded input map being based on the sixth coded map of the first modality, the fourth coded input map being based on the fourth coded map of the fourth cell, the fourth coded input map being based on the fourth coded input map,Adapting each of the third input feature vectors based on the third input feature vectors, optionally their third joint feature center points, and the fourth input feature vectors, optionally their fourth joint feature center points, by a second matching descriptor layer; obtaining a third joint map, the third joint map including a respective third joint feature vector for each of the third joint feature center points; adapting each of the fourth input feature vectors based on the third input feature vectors, optionally their third joint feature center points, and the fourth input feature vectors, optionally their fourth joint feature center points, by a second matching descriptor layer; obtaining a fourth joint map, the fourth joint map including a respective fourth joint feature vector for each of the fourth joint feature center points; calculating, for each of the fourth cells, a similarity between the third joint feature vector of the respective third cell and the fourth joint feature vector of the respective fourth cell; calculating, by the second optimal matching layer, for each of the third cells and each of the fourth cells, a second point correlation between each of the third cells and each of the fourth cells based on the similarity between the third joint feature vector of the respective third cell and the fourth joint feature vector of the fourth cell and based on the similarity between the fourth joint feature vector of the respective fourth cell and the third joint feature vector of the third cell; and checking, for each of the third cells and each of the fourth cells, whether at least one of one or more second correlation conditions for each of the third cells and each of the fourth cells is satisfied, wherein the one or more second correlation conditions are: a second point correlation between each third cell and each fourth cell is greater than a second correlation threshold for the second tier; or The second point correlation between each third cell and each fourth cell is among the largest l values ​​of the second point correlation, where l is a fixed value.

[0018] The instructions, when executed by the one or more processors, may further cause the apparatus to perform the step of prohibiting the first matching descriptor layer from checking, for each of the first cells and each of the second cells, whether a first point correlation between the first joint feature vector of the respective first cell and the second joint feature vector of the respective second cell is greater than a first correlation threshold of the first tier in response to checking that at least one of the one or more second correlation conditions for each of the third cells and each of the fourth cells is satisfied.

[0019] When executed by the one or more processors, the instructions may further cause the apparatus to perform the steps of: extracting, for each of the third cells and each of the fourth cells, coordinates of a third joint feature center point of the respective third cell in the first coordinate system and coordinates of a fourth joint feature center point of the respective fourth cell in the second coordinate system, and obtaining respective pairs of extracted coordinates of the second hierarchy in response to checking that at least one of the one or more second correlation conditions for the respective third cell and each fourth cell is satisfied; and calculating an estimated transformation between the first coordinate system and the second coordinate system based on the pairs of extracted coordinates of the first hierarchy and the pairs of extracted coordinates of the second hierarchy.

[0020] The instructions, when executed by the one or more processors, may further cause the apparatus to perform a step of calculating an estimated transformation such that, in the calculating step, extracted coordinate pairs of the first hierarchical level have different weights than extracted coordinate pairs of the second hierarchical level.

[0021] The second matching descriptor layer of the second hierarchy is configured by respective third parameters, and the instructions, when executed by the one or more processors, may further cause the device to perform the step of training the device to determine the first and third parameters so that a deviation between the estimated transformation and the known transformation between the first coordinate system and the second coordinate system is less than a transformation estimation threshold, and so that the first point correlation and the second point correlation are increased.

[0022] The fifth encoder layer may be the same as the sixth encoder layer.

[0023] The fifth encoder layer may be the same as the first encoder layer.

[0024] The sixth encoder layer may be the same as the second encoder layer.

[0025] The second matching descriptor layer may be the same as the first matching descriptor layer.

[0026] The second optimal matching layer may be the same as the first optimal matching layer.

[0027] The third coded input map may be a third coded map of the first modality, the third joint feature center points may be third feature center points, the fourth coded input map may be a fourth coded map of the first modality, and the fourth joint feature center points may be fourth feature center points.

[0028] The instructions, when executed by the one or more processors, further cause the apparatus to input the third coded map of the second modality and the third grid to a seventh encoder layer; and encoding, by the seventh encoder layer, the third coded map of the second modality into a seventh coded map of the second modality, the seventh coded map of the second modality comprising, for each of the third cells of the third grid, a respective seventh feature center point and a respective seventh feature vector, the coordinates of the seventh feature center point being determined by the first the steps of: inputting the seventh coded map to a third neural fusion layer; inputting the fifth coded map to the third neural fusion layer; generating a third coded input map by the third neural fusion layer based on the seventh coded map and the fifth coded map, wherein for each of the third cells, a respective third joint feature center point is based on the fifth feature center point of the respective third cell and the seventh feature center point of the respective third cell; a third input feature vector based on the fifth feature vector of each third cell adapted by the fifth feature vector and the seventh feature vector, and based on the seventh feature vector of each third cell adapted by the fifth feature vector and the seventh feature vector, inputting the fourth encoded map of the second modality and the fourth grid into an eighth encoder layer; and encoding the fourth encoded map of the second modality into an eighth encoded map of the second modality by the eighth encoder layer, an eighth coded map of the second modality comprising, for each of the fourth cells of the fourth grid, a respective eighth feature center point and a respective eighth feature vector, the coordinates of the eighth feature center points being represented in the second coordinate system; inputting the eighth coded map to a fourth neural fusion layer; inputting the sixth coded map to the fourth neural fusion layer; and generating, by the fourth neural fusion layer, a fourth coded input map based on the sixth coded map and the eighth coded map;For each fourth cell, the respective fourth joint feature center point is based on the sixth feature center point of the respective fourth cell and the eighth feature center point of the respective fourth cell, and the respective fourth input feature vector is based on the sixth feature vector of the respective fourth cell adapted by the sixth feature vector and the eighth feature vector, and based on the eighth feature vector of the respective fourth cell adapted by the sixth feature vector and the eighth feature vector.

[0029] The instructions, when executed by the one or more processors, further cause the apparatus to at least: input the first coded input map to a fifth encoder layer, and encode, by the fifth encoder layer, the first coded input map and the first coded map of the first modality into a fifth coded map of the first modality; input the first coded input map to a seventh encoder layer, and encode, by the seventh encoder layer, the first coded input map and the third coded map of the second modality into a seventh coded map of the second modality; inputting the second coded input map to a sixth encoder layer and encoding, by the sixth encoder layer, the second coded input map and the second coded map of the first modality into a sixth coded map of the first modality; inputting the second coded input map to an eighth encoder layer and encoding, by the eighth encoder layer, the second coded input map and the fourth coded map of the second modality into an eighth coded map of the second modality.

[0030] The second matching descriptor layer of the second hierarchy is configured by respective third parameters, and the instructions, when executed by the one or more processors, may further cause the device to perform the step of training the device to determine the first, second, and third parameters so that the deviation between the estimated transformation and the known transformation between the first coordinate system and the second coordinate system is less than a transformation estimation threshold, and so that the first point correlation and the second point correlation increase.

[0031] The seventh encoder layer may be the same as the eighth encoder layer.

[0032] The seventh encoder layer may be the same as the third encoder layer.

[0033] The eighth encoder layer may be the same as the fourth encoder layer.

[0034] The third neural fusion layer may be the same as the fourth neural fusion layer.

[0035] The third neural fusion layer may be the same as the first neural fusion layer.

[0036] The fourth neural fusion layer may be the same as the second neural fusion layer.

[0037] The instructions, when executed by the one or more processors, may further cause the apparatus to perform at least one of: re-locating a second sequence of images in a second coordinate system in a first sequence of images in a first coordinate system based on the estimated transformation; or co-locating a fourth sequence of images in the second coordinate system with a third sequence of images in the first coordinate system based on the estimated transformation; or looping a portion of the image sequence with the entire image sequence based on the estimated transformation, wherein the positions of the images of the portion of the image sequence are shown in the second coordinate system and the images of the entire image sequence are shown in the first coordinate system.

[0038] According to a second aspect of the present invention, a method includes inputting a first 3D point cloud of a first modality and a first grid including one or more first cells into a first encoder layer, where coordinates of points of the first 3D point cloud are expressed in a first coordinate system; and encoding, by the first encoder layer, the first 3D point cloud of the first modality into a first encoded map of the first modality, where the first encoded map of the first modality includes, for each first cell of the first grid, a respective first feature center point and its associated coordinates. a step of inputting a first coded input map to a first matching descriptor layer of a first hierarchical level, the first coded input map being based on a first coded map of a first modality and comprising, for each of the first cells, a respective first joint feature center point and a respective first input feature vector, the coordinates of the first joint feature center point being expressed in the first coordinate system; , and a second grid including one or more second cells to a second encoder layer, where coordinates of points of the second 3D point cloud are represented in a second coordinate system; and encoding, by the second encoder layer, the second 3D point cloud of the first modality into a second encoded map of the first modality, where the second encoded map of the first modality includes, for each of the second cells of the second grid, a respective second feature center point and a respective second feature vector, where the coordinates of the second feature center point are represented in the second coordinate system. a step of inputting a second coded input map into a first matching descriptor layer, the second coded input map being based on the second coded map of the first modality and comprising, for each of the second cells, a respective second joint feature center point and a respective second input feature vector, the coordinates of the second joint feature center points being represented in a second coordinate system; and a step of inputting the first input feature vector, optionally the first feature center points and the second input feature vector, by the first matching descriptor layer.and optionally their second feature center points, by a first matching descriptor layer. The method further comprises: adapting each of the first input feature vectors based on the first input feature vectors, optionally their first feature center points, and the second input feature vectors, optionally their second feature center points, by a first matching descriptor layer, by obtaining a second joint map, the second joint map including a respective second joint feature vector for each of the second joint feature center points; and adapting each of the second input feature vectors based on the first input feature vectors, optionally their first feature center points, and the second input feature vectors, optionally their second feature center points, by a first optimal matching layer, for each of the first cells and each of the second cells. calculating, by a first optimal matching layer, for each of the first cells and each of the second cells, a first point correlation between each of the first cells and each of the second cells based on the similarity between the first joint feature vector of each of the first cells and some second joint feature vectors of some second cells and based on the similarity between the second joint feature vector of each of the second cells and some first joint feature vectors of some of the first cells; and for each of the first cells and each of the second cells, checking whether at least one of one or more first correlation conditions is satisfied for each of the first cells and each of the second cells, wherein the one or more first correlation conditions include: a first point correlation between each first cell and each second cell is greater than a first correlation threshold for the first tier; or the first point correlation between each first cell and each second cell is among the largest k values ​​of first point correlations, where k is a fixed value; An apparatus is provided that performs the steps of: extracting, for each first cell and each second cell, coordinates of a first joint feature center point of the respective first cell in a first coordinate system and coordinates of a second joint feature center point of the respective second cell in a second coordinate system, and obtaining each pair of extracted coordinates of the first hierarchical layer in response to checking whether at least one of one or more first correlation conditions is satisfied for each first cell and each second cell; and calculating an estimated transformation between the first coordinate system and the second coordinate system based on the pair of extracted coordinates of the first hierarchical layer.

[0039] The first encoder layer, the second encoder layer, and the first matching descriptor layer of the first hierarchy are configured by respective first parameters, and the method may further include a step of training the device to determine the first parameters so that a deviation between the estimated transformation and the known transformation between the first coordinate system and the second coordinate system is less than a transformation estimation threshold and so that the first point correlation is increased.

[0040] The first modality may include one of a photographic image, a LIDAR image, and an ultrasound image.

[0041] The first encoder layer may be the same as the second encoder layer.

[0042] The method may further include at least generating a first 3D point cloud based on a sequence of first images of a first modality, wherein a respective position in the first coordinate system is associated with each of the first images, or generating a second 3D point cloud based on a sequence of second images of the first modality, wherein a respective position in the second coordinate system is associated with each of the second images.

[0043] The first coded input map may be a first coded map, the first joint feature center points may be first feature center points, the second coded input map may be a second coded map, and the second joint feature center points may be second feature center points.

[0044] The method further includes inputting a third 3D point cloud of a second modality and the first grid into a third encoder layer, where coordinates of points of the third 3D point cloud are expressed in a first coordinate system, the second modality being different from the first modality; and encoding, by the third encoder layer, the third 3D point cloud of the second modality into a third encoded map of the second modality, where the third encoded map of the second modality includes, for each first cell of the first grid, a respective third feature center point and a respective third feature center point. a feature vector, wherein the coordinates of the third feature center points are represented in a first coordinate system; inputting the third coded map to a first neural fusion layer; inputting the first coded map to the first neural fusion layer; and generating a first coded input map by the first neural fusion layer based on the first coded map and the third coded map, wherein for each of the first cells, the respective first joint feature center points are represented by the first feature center points of the respective first cells and the third joint feature center points of the respective first cells. a step based on third feature center points, and each first input feature vector is based on the first feature vector and the first feature vector, optionally the first feature center points and third feature vector of each first cell adapted by those first feature center points, optionally their third feature center points, and a step based on the third feature vector and the first feature vector, optionally the third feature center points and third feature vector of each first cell adapted by those first feature center points, optionally their third feature center points; inputting the fourth 3D point cloud of the second modality and the second grid into a fourth encoder layer, where coordinates of points of the fourth 3D point cloud are represented in the second coordinate system; and encoding, by the fourth encoder layer, the fourth 3D point cloud of the second modality into a fourth encoded map of the second modality, where the fourth encoded map of the second modality includes, for each second cell of the second grid, a respective fourth feature center point and a respective fourth feature vector, where coordinates of the fourth feature center point are represented in the second coordinate system;The method may include inputting the fourth coded map to a second neural fusion layer, inputting the second coded map to the second neural fusion layer, and generating a second coded input map by the second neural fusion layer based on the second coded map and the fourth coded map, wherein for each of the second cells, the respective second joint feature center points are based on the second feature center points of the respective second cells and the fourth feature center points of the respective second cells, and the respective second input feature vectors are based on the second feature vector and the second feature vector, optionally the second feature center points and the fourth feature vector of the respective second cells adapted by the second feature center points, optionally the fourth feature center points, and the fourth feature vector and the second feature vector, optionally the fourth feature center points and the fourth feature vector of the respective second cells adapted by the second feature center points, optionally the fourth feature center points.

[0045] The second modality may include at least one of a photographic image, an RF fingerprint, a LIDAR image, or an ultrasound image.

[0046] The third encoder layer, the fourth encoder layer, the first neural fusion layer, and the second neural fusion layer are configured with respective second parameters, and the method may further include training the apparatus to determine the first parameters and the second parameters such that a deviation between the estimated transformation and the known transformation between the first coordinate system and the second coordinate system is less than a transformation estimation threshold and such that the first point correlation is increased.

[0047] The first image may be taken by a first sensor of the first sensor device, and the second image may be taken by a second sensor of the second sensor device, and the method may further include taking a plurality of first measurements of a second modality using a third sensor of the first sensor device; taking a plurality of second measurements of the second modality using a fourth sensor of the second sensor device; generating a third 3D point cloud of the second modality based on the plurality of first measurements, positions associated with the first image, and a calibration between the first sensor of the first sensor device and the third sensor of the first sensor device; and generating a fourth 3D point cloud of the second modality based on the plurality of second measurements, positions associated with the second image, and a calibration between the second sensor of the second sensor device and the fourth sensor of the second sensor device.

[0048] The third encoder layer may be the same as the fourth encoder layer.

[0049] The first neural fusion layer may be the same as the second neural fusion layer.

[0050] The method further includes inputting the first encoded map of the first modality and a third grid including one or more third cells into a fifth encoder layer, wherein each of the third cells includes two or more of the first cells, and the first grid includes two or more of the first cells; and encoding, by the fifth encoder layer, the first encoded map of the first modality into a fifth encoded map of the first modality, wherein the fifth encoded map of the first modality includes, for each of the third cells of the third grid: a step of inputting a third coded input map to a second matching descriptor layer of a second hierarchical level, the third coded input map being based on the fifth coded map of the first modality and comprising, for each of the third cells, a respective third joint feature center point and a respective third input feature vector, the coordinates of the third joint feature center point being expressed in the first coordinate system; inputting the second coded map of the first modality and a fourth grid including one or more fourth cells into a sixth encoder layer, wherein each of the fourth cells includes two or more of the second cells, and the second grid includes two or more of the second cells; encoding the second coded map of the first modality into a sixth coded map of the first modality by the sixth encoder layer, wherein the sixth coded map of the first modality includes, for each of the fourth cells of the fourth grid, a respective sixth feature center point and a corresponding and a respective sixth feature vector, the coordinates of the sixth feature center points being expressed in a second coordinate system; inputting the fourth coded input map into a second matching descriptor layer, the fourth coded input map being based on the sixth coded map of the first modality and comprising, for each of the fourth cells, a respective fourth joint feature center point and a respective fourth input feature vector, the coordinates of the fourth joint feature center points being expressed in the second coordinate system; and inputting the third input feature vector,and optionally adapting each of the third input feature vectors based on the third joint feature center points and the fourth input feature vector, and optionally the fourth joint feature center points; obtaining a third joint map, the third joint map including a respective third joint feature vector for each of the third joint feature center points; adapting each of the fourth input feature vectors based on the third input feature vector, and optionally the third joint feature center points and the fourth input feature vector, and optionally the fourth joint feature center points, by a second matching descriptor layer; obtaining a fourth joint map, the fourth joint map including a respective fourth joint feature vector for each of the fourth joint feature center points; and adapting each of the fourth input feature vectors based on the third input feature vector, and optionally the third joint feature center points and the fourth input feature vector, and optionally the fourth joint feature center points, by a second optimal matching layer. calculating a similarity between a third joint feature vector of each third cell and a fourth joint feature vector of each fourth cell by the second optimal matching layer; calculating, for each of the third cells and each of the fourth cells, a second point correlation between each of the third cells and each of the fourth cells based on the similarity between the third joint feature vector of each of the third cells and the fourth joint feature vector of the fourth cell and based on the similarity between the fourth joint feature vector of each of the fourth cells and the third joint feature vector of the third cell; and checking, for each of the third cells and each of the fourth cells, whether at least one of one or more second correlation conditions for each of the third cells and each of the fourth cells is satisfied, wherein the one or more second correlation conditions include: a second point correlation between each third cell and each fourth cell is greater than a second correlation threshold for the second tier; or The second point correlation between each third cell and each fourth cell may include a step of being among the largest l values ​​of the second point correlation, where l is a fixed value.

[0051] The method may further include a step of prohibiting the first matching descriptor layer from checking, for each of the first cells and each of the second cells, whether a first point correlation between the first joint feature vector of the respective first cell and the second joint feature vector of the respective second cell is greater than a first correlation threshold of the first tier in response to checking that at least one of the one or more second correlation conditions for each of the third cells and each of the fourth cells is satisfied.

[0052] The method may further include steps of: extracting, for each of the third cells and each of the fourth cells, coordinates of a third joint feature center point of the respective third cell in the first coordinate system and coordinates of a fourth joint feature center point of the respective fourth cell in the second coordinate system; and obtaining each pair of extracted coordinates of the second hierarchy in response to checking that at least one of one or more second correlation conditions for the respective third cell and each fourth cell is satisfied; and calculating an estimated transformation between the first coordinate system and the second coordinate system based on the pair of extracted coordinates of the first hierarchy and the pair of extracted coordinates of the second hierarchy.

[0053] The method further includes calculating the estimated transformation such that in the calculating step, extracted coordinate pairs in the first hierarchical layer have different weights than extracted coordinate pairs in the second hierarchical layer.

[0054] The second matching descriptor layer of the second hierarchy is configured by respective third parameters, and the method may further include a step of training the apparatus to determine the first and third parameters so that the deviation between the estimated transformation and the known transformation between the first coordinate system and the second coordinate system is less than a transformation estimation threshold, and so that the first point correlation and the second point correlation increase.

[0055] The fifth encoder layer may be the same as the sixth encoder layer.

[0056] The fifth encoder layer may be the same as the first encoder layer.

[0057] The sixth encoder layer may be the same as the second encoder layer.

[0058] The second matching descriptor layer may be the same as the first matching descriptor layer.

[0059] The second optimal matching layer may be the same as the first optimal matching layer.

[0060] The third coded input map may be a third coded map of the first modality, the third joint feature center points may be third feature center points, the fourth coded input map may be a fourth coded map of the first modality, and the fourth joint feature center points may be fourth feature center points.

[0061] The method further includes inputting the third coded map of the second modality and the third grid into a seventh encoder layer; encoding, by the seventh encoder layer, the third coded map of the second modality into a seventh coded map of the second modality, the seventh coded map of the second modality comprising, for each third cell of the third grid, a respective seventh feature center point and a respective seventh feature vector, the coordinates of the seventh feature center point being expressed in the first coordinate system; and encoding the seventh coded map by the seventh encoder layer. inputting the fifth encoded map to a third neural fusion layer; and generating a third encoded input map by the third neural fusion layer based on the seventh encoded map and the fifth encoded map, wherein for each of the third cells, a respective third joint feature center point is based on the fifth feature center point of the respective third cell and the seventh feature center point of the respective third cell, and a respective third input feature vector is based on the fifth feature vector and the seventh feature vector. and based on a seventh feature vector of each third cell adapted by the fifth feature vector and the seventh feature vector, inputting the fourth encoded map of the second modality and the fourth grid to an eighth encoder layer, and encoding the fourth encoded map of the second modality into an eighth encoded map of the second modality by the eighth encoder layer, wherein the eighth encoded map of the second modality is based on the fifth feature vector of each third cell adapted by the fifth feature vector and the seventh feature vector, and generating a fourth coded input map by the fourth neural fusion layer based on the sixth coded map and the eighth coded map, wherein for each of the fourth cells, the respective fourth joint feature center points and the respective eighth feature vectors are represented in a second coordinate system; inputting the eighth coded map to a fourth neural fusion layer; inputting the sixth coded map to the fourth neural fusion layer; and generating a fourth coded input map by the fourth neural fusion layer based on the sixth coded map and the eighth coded map, wherein for each of the fourth cells, the respective fourth joint feature center points are represented in a second coordinate system.The method may include steps of: based on the sixth feature center point of each fourth cell and the eighth feature center point of each fourth cell; and based on the sixth feature vector of each fourth cell adapted by the sixth feature vector and the eighth feature vector, and based on the eighth feature vector of each fourth cell adapted by the sixth feature vector and the eighth feature vector.

[0062] The method further includes inputting the first coded input map to a fifth encoder layer and encoding, by the fifth encoder layer, the first coded input map and the first coded map of the first modality into a fifth coded map of the first modality; inputting the first coded input map to a seventh encoder layer and encoding, by the seventh encoder layer, the first coded input map and the third coded map of the second modality into a seventh coded map of the second modality. , inputting the second coded input map to a sixth encoder layer and encoding, by the sixth encoder layer, the second coded input map and the second coded map of the first modality into a sixth coded map of the first modality; and inputting the second coded input map to an eighth encoder layer and encoding, by the eighth encoder layer, the second coded input map and the fourth coded map of the second modality into an eighth coded map of the second modality.

[0063] The second matching descriptor layer of the second hierarchy is configured by respective third parameters, and the method may further include a step of training the apparatus to determine the first, second, and third parameters so that the deviation between the estimated transformation and the known transformation between the first coordinate system and the second coordinate system is less than a transformation estimation threshold, and so that the first point correlation and the second point correlation increase.

[0064] The seventh encoder layer may be the same as the eighth encoder layer.

[0065] The seventh encoder layer may be the same as the third encoder layer.

[0066] The eighth encoder layer may be the same as the fourth encoder layer.

[0067] The third neural fusion layer may be the same as the fourth neural fusion layer.

[0068] The third neural fusion layer may be the same as the first neural fusion layer.

[0069] The fourth neural fusion layer may be the same as the second neural fusion layer.

[0070] The method may further include at least re-localizing a second sequence of images in a second coordinate system in a first sequence of images in a first coordinate system based on the estimated transformation, or co-localizing a fourth sequence of images in the second coordinate system with a third sequence of images in the first coordinate system based on the estimated transformation, or loop-closing a portion of the image sequence with the entire image sequence based on the estimated transformation, wherein the positions of the images of the portion of the image sequence are shown in the second coordinate system and the positions of the images of the entire image sequence are shown in the first coordinate system.

[0071] The method of the second aspect may be a method of location estimation.

[0072] According to a third aspect of the present invention there is provided a computer program product comprising a set of instructions arranged, when executed on an apparatus, to cause the apparatus to carry out a method according to the second aspect.

[0073] The computer program product may be embodied as a computer-readable medium or may be directly loadable into a computer.

[0074] According to some exemplary embodiments, at least one of the following advantages may be achieved: Robustness in location estimation can be improved because the location estimation is based on a point cloud (which can be generated based on a sequence of images) instead of on a single image. Robust location estimation can be performed in real time. Large areas can be covered. Ambiguities (e.g., in repetitive regions) can be reduced especially by fusion of different modalities. The accuracy of the location estimation can be improved, especially by scaling it through hierarchical methods. RF maps can be predicted from implicit propagation models contained in the data, and thus sensor poses can be predicted even in areas where no measurements have been made before. The process can be camera independent and less susceptible to image distortions. · Training data collection does not require manual effort. In the encoding layer, the 3D positions of feature vectors are regressed in a data-driven, differentiable, and learnable manner to achieve high accuracy in rigid transformation estimation even on the coarsest layers.

[0075] It is to be understood that any of the above modifications may be applied singly or in combination to the respective aspects to which they refer, unless expressly stated as excluding alternatives.

[0076] Further details, features, objects and advantages will be apparent from the following detailed description of preferred exemplary embodiments, taken in conjunction with the accompanying drawings. [Brief explanation of the drawings]

[0077] [Figure 1] FIG. 1 illustrates SLAM building blocks. [Figure 2] FIG. 1 illustrates a functional structure of an exemplary embodiment. [Figure 3] FIG. 1 illustrates a functional structure of an exemplary embodiment. [Figure 4] FIG. 1 illustrates a functional structure of an exemplary embodiment. [Figure 5] FIG. 5 shows a cutout of the functional structure of the exemplary embodiment of FIG. [Figure 6] FIG. 1 illustrates an apparatus in accordance with an exemplary embodiment. [Figure 7] FIG. 1 illustrates a method according to an exemplary embodiment. [Figure 8] 1 illustrates an apparatus in accordance with an exemplary embodiment; DETAILED DESCRIPTION OF THE INVENTION

[0078] The present invention has been described in detail below, and the cases where the features of the present embodiment can be freely combined will be described with reference to the accompanying drawings. However, it should be clearly understood that the description of the specific embodiments is given by way of example only, and that it is in no way intended to limit the invention to the details disclosed.

[0079] Furthermore, in some cases, only an apparatus or only a method is described, but it should be understood that the apparatus is configured to perform the corresponding method.

[0080] Some example embodiments aim to provide a transformation between a first map and a second map so that the first map matches the second map. For example, the first map may be a global map of, for example, a building, floor, etc., and the second map may be a local map acquired based on images from a sensor, such as a robot. Each of the first map and the second map may be acquired from a respective sequence of images of one or more modalities, such as visual images, LIDAR images, RF measurements, or ultrasound images. At least one of the modalities allows for the creation of a map based on the respective images, such as the visual or LIDAR image modality.

[0081] The second map has its own coordinate system, which is typically different from the coordinate system of the first map. An SE(3) transformation is defined between the coordinate frames. The transformation between the coordinates of the second map and the coordinates of the first map may include translation (represented by a translation vector t with three independent components) and rotation (represented by a rotation matrix R with three independent components (a rotation matrix is ​​a 3x3 matrix with three degrees of freedom)). Thus, in general, the pose has six independent components (6DoF, also called "6D pose," corresponding to position and orientation). In some exemplary embodiments, the pose may have less than 6DoF if respective restrictions are imposed.

[0082] In particular, one, two, or three of the following operations may be performed by some example embodiments. Relocalization: Matching the local map (second map) with the global map (first map). Co-localization: Matching a second local map (second map) with a first local map (first map). Loop Closure: A partial local map (second map), obtained from a subset of the images from which the local map is obtained, is aligned with the (global) local map (first map). This procedure is useful for correcting drift in the local map, which can result from the accumulation of small odometry errors.

[0083] Below, we describe several exemplary embodiments for relocation where the first map is a global map and the second map is a local map. The transformation indicates the position and orientation (posture) of the sensor used to acquire the image on which the local map is based in the coordinate system of the global map. The described procedures can be applied correspondingly to co-localization and loop closure.

[0084] The 3D structure of the environment may be represented by a point cloud. Here, we describe the generation of a point cloud from camera images as an example. Each camera image is associated with the camera's position at the time of capture. Typically, visual inertial odometry can provide the relative transformation between adjacent camera frames, which serves as a good initial value for SLAM state estimation. The output of the SLAM algorithm will further refine the initial relative transformation of the frames and provide depth estimates for a subset of image pixels. Based on the estimated depth information and relative transformation, a 3D point cloud is constructed that represents the structure of the environment. Relocation can be used when tracking is lost between frames and / or for drift error correction.

[0085] In the process of creating location estimation maps and location estimation algorithms, robustness can be achieved through multimodal sensor fusion, such as that used by visual inertial odometry. Here, an inertial measurement unit (IMU) can measure the angular velocity and acceleration of an agent without being subject to significant light fluctuations, flare, or dynamic environments such as a camera, which can significantly improve robustness. In a tightly coupled sensor fusion system, information from camera images can facilitate estimation of gravity vectors, gyroscope, and accelerometer biases, resulting in significant drift reduction. In some exemplary embodiments, the position of the camera at the time of image capture may be obtained solely by visual or inertial odometry, rather than by visual inertial odometry.

[0086] The mathematical framework for creating a 3D reconstruction of the environment while determining the 6DoF track of the camera is the Simultaneous Localization and Mapping (SLAM) algorithm or the Structure for Motion (SFM) algorithm. There are several versions of such algorithms.

[0087] Typically, SLAM algorithms include two stages. In the first stage (global mapping stage), the SLAM algorithm provides a map suitable for localization purposes. Depending on the number of images used to obtain the map, this stage may be slow or fast. In the next stage, the localization may be used to identify the pose of a new image or image sequence within the acquired map.

[0088] In the global map construction stage, a complex optimization problem called full bundle adjustment is performed to reduce any pose errors and 3D reconstruction errors. An element of map optimization is loop closure detection. Detecting loops in the map construction process can significantly reduce accumulated drift and 3D reconstruction errors. This non-convex optimization involves a large number of state variables, making this stage very computationally intensive and therefore typically performed offline. Local map construction and localization are often performed online but can also be performed offline.

[0089] [Generating point clouds from a sequence of images] With reference to FIG. 1, a SLAM algorithm (SLAM pipeline) is described as an example of how to generate a point cloud based on a sequence of images.

[0090] The main building blocks of an exemplary SLAM pipeline are tracking, loop closure, local map construction, and global image search.

[0091] In the tracking phase, the goal is to determine the relative odometry between two consecutive frames and determine whether the new image frame is a keyframe. Different implementations define different heuristics for keyframe detection, so there is no exact definition of a keyframe. Generally, a keyframe contains new information about the environment that is not available in the previous keyframe. To calculate the odometry between two consecutive frames, feature points (and corresponding feature vectors) are extracted and tracked. If feature tracking is successful, the relative pose of the two images can be calculated, and depending on the sensor settings, depth information of the feature points can be estimated.

[0092] If feature tracking fails, a global relocalization procedure is initiated, aiming to find the current pose in the existing map. This task is particularly challenging because the problem complexity (search space) increases as new keyframes are inserted into the keyframe database. To attempt global image retrieval from the keyframe database, a global descriptor for the query image is created. In many conventional solutions, a bag-of-words (BoW) algorithm quantizes the image's feature descriptors and creates a frequency histogram of the quantized feature vectors. Alternatively, deep learning models can be trained to generate state-of-the-art global image descriptors that can be searched very efficiently. After creating the global descriptor, a similarity matrix is ​​calculated between the query image's global descriptor and the global descriptors of all stored keyframes. Images with the highest similarity scores are searched as candidates for the image matching process, followed by different filtering algorithms to further refine the candidate set. If the search process is successful, the global pose of the query image is calculated and a local map of the environment is searched based on the co-visibility graph. The local map contains keyframes, 3D feature points with local feature descriptors, and depth information. The feature points in the 3D map are searched on the 2D query image, and from these 3D-2D matches, the camera pose can then be recovered in a Perspective-n-Point (PnP) manner. Tracking continues by analyzing new images.

[0093] When a new keyframe is detected, a loop closure procedure can be initiated. Here, the goal is to find similar keyframes from the past to find if the camera has already visited the same location. The search procedure is similar to the one presented earlier. The output of this block is a loop constraint that is added to the bundle adjustment optimization. The optimization can take advantage of the fact that along a loop in the pose graph, the overall pose change should be the same.

[0094] The local map construction module is responsible for refining the local map by minimizing the discrepancy between the 3D structure and the camera pose in a bundle adjustment procedure. The discrepancy is formulated as a differential residual function. This non-convex optimization is formulated as a least-squares minimization problem. During optimization, a significant number of state variables are updated: pose, depth values ​​(of points), and sensor extrinsic and intrinsic parameters. A reprojection error function connects the pose, depth, and calibration state variables to optical flow measurements. If available, an IMU-based residual function may connect neighboring poses and IMU-internal state variables with pre-integrated IMU measurements. Finally, depending on the purpose of the constructed map, the number of state variables can be reduced, resulting in a local map that represents only a limited region of the environment. Conversely, if every keyframe is retained, the local map represents a larger region of the environment.

[0095] In the global map optimization component (which can be performed in the background or at an offline stage), the goal is to highly accurately estimate (i.) a 3D reconstruction of the environment, (ii.) a camera pose from where the images were captured, and (iii.) descriptors attached to the 3D points that help solve the localization problem. Marginalization or keyframe culling can be performed to limit the computational complexity to the given processing hardware.

[0096] Thus, from a sequence of images, a point cloud is generated, where each point in the point cloud is defined by its coordinate in the respective coordinate system and a feature vector, which describes the properties of the point, such as RGB, reflectance, etc.

[0097] The use of point clouds as input for relocalization (or colocalization or loop closure) has advantages over the use of single images: existing solutions attempt to find the pose of a single query image. However, there are environments where the lack of descriptive features or recurring patterns leads to high ambiguity. This ambiguity can be resolved (at least in part) by the use of point clouds based on a sequence of captured images and / or the fusion of different modalities.

[0098] Image-based localization problems are notoriously difficult to solve because lighting conditions encountered during the day or seasonal changes in the environment introduce different artifacts. Creating visual descriptors and models that can generalize to all conditions is particularly challenging. By creating a local map for structural representation and registering RF measurements with the camera pose, a better representation can be constructed for localization purposes. After a global map is constructed, different deep learning models can be trained to register features from the local maps created under different conditions. Nevertheless, it is also possible to train a single model that generalizes to all conditions by creating a set of local maps of the same area at different times of the day and / or season.

[0099] Location estimation, image retrieval algorithms, and data-driven training models can be significantly affected by camera lens distortion. When query images are created with a different camera than the keyframe images, only suboptimal accuracy is achieved. Some exemplary embodiments are camera-independent because the camera image sequence is converted into a point cloud.

[0100] As another option, the 3D point cloud may be pre-defined based on planning data that describes, for example, how an environment (such as a building) was planned. However, if the 3D point cloud is based on planning data, there is a risk that the reality of the environment as captured by a sensor (such as a camera) will not match the planning data.

[0101] Basic System According to Some Example Embodiments With reference to FIG. 2 , this section describes a basic system according to some exemplary embodiments. The basic system comprises two pipelines 1 a and 1 b: pipeline 1 a for a global map represented by a global 3D point cloud, and pipeline 1 b for a local map represented by a local 3D point cloud. Pipelines 1 a and 1 b and their parameters may be the same for both point clouds or may be different for the two point clouds. Pipelines 1 a and 1 b operate independently of each other. The parameters of the layers included in the pipelines are learned in a training process, which is described further below. The same applies to the pipelines described for further embodiments below.

[0102] For estimation (query phase), 3D point clouds (coordinates and respective feature vectors) are input to each pipeline. In addition, respective 3D grids defining cells (voxels) are input to each pipeline. The 3D grids may be the same or different for Pipeline 1a and Pipeline 1b. The 3D grid for the local map may be shifted and / or rotated relative to the 3D grid for the global map. Furthermore, the number of (associated) cells may be the same or different for the global and local maps, since these maps may cover volumes of different sizes. Typically, the cells of the grid are cubic, but they may have any shape (e.g., a general rectangular prism or pyramid) that can cover the entire volume. The size of the cells of the 3D grid can be selected depending on the required resolution and available processing power. The number of cells of the 3D grid may be one, but is typically two or more. The cells of the 3D grid may extend to infinity in one direction (e.g., the height direction) so that the cell can be represented by only two coordinates. That is, a 3D grid can be equivalent to a 2D grid.

[0103] The functions of the pipelines 1a and 1b (collectively referred to as 1) will be described below as representative examples.

[0104] Pipeline 1 includes structural encoder layers ("first and second encoder layers"). The structural encoder layers encode the 3D point cloud in a transformation-invariant manner. This task is difficult because the points in the 3D point cloud are not regularly spaced. Some cells of the 3D grid may not contain any points, while other cells may contain one or more points with their respective feature vectors. Several suitable encoding techniques exist, including (i) projection networks, in which points are projected onto regular 2D and 3D grid structures followed by regular 2D or 3D convolution operators; (ii) graph convolution networks; (iii) point-wise multilayer perceptron networks; or (iv) point convolution networks.

[0105] The 3D grid can be used for data compression. By applying data compression, the number of points is reduced and they become better organized structurally. When data compression is performed, the goal is to define (i.) a structural descriptor encoded as a feature vector for each grid cell, and (ii.) regress feature (center) point coordinates of the structural descriptor and RF descriptor used in computing the rigid body transformation.

[0106] The point compression operation for one 3D grid configuration can be formulated as follows: (i.) Define a 3D grid with a given grid cell size. (ii.) Grouping input points based on grid cells. (iii.) Regress the locations of the output feature centroids and any additional parameters of the selected convolution type. (iv.) For every feature center point based on the feature vectors of the support points (the set of support points is the set of neighbors of the feature center point, typically defined by a convolution radius or based on a neighbor selection algorithm, and further, the convolution radius may be equal to, greater than, or less than the size of the grid cell), a feature vector is created for the output feature point (i.e., feature center point). The feature vector is calculated by performing several groups of convolution, normalization, and activation operations.

[0107] Thus, for each cell, the structure encoder layer obtains the coordinates of the feature center point, i.e., the point representing the cell that has a feature vector representing the cell's features.

[0108] Note that due to the coding, the features in the coded global map may not have the usual meaning of RGB or reflectance, etc. They must be considered as abstract features. The same is true for the output of the other coding layers described below.

[0109] The first and second encoder layers of pipelines 1a and 1b may have the same or different architectures (e.g., number of convolutional layers, etc.) and / or the same or different parameters. In some exemplary embodiments, the first and second encoder layers may be the same.

[0110] The output of the structure encoder layer of both pipelines (i.e., their respective encoded maps) is input to the matching descriptor layer. In other words, in this exemplary embodiment, the output of the structure encoder layer can be considered as an encoded input map to the matching descriptor layer. The matching descriptor layer (and other layers following the matching descriptor layer) are common to the global map and the local map and do not belong to either the global map or local map-specific pipelines 1a and 1b.

[0111] In the matching descriptor layer, a function with learnable parameters is applied to each feature vector of the feature center points of each encoded map. The function describes how the feature vector and the positions of other feature center points of each encoded map affect the feature vector of a selected feature center point of the same map. The function may depend on the distance between the respective feature center points. An example of such a mechanism is self-attention. In addition, a further function is applied that describes how the feature vector of each feature center point of one of the encoded maps is affected by the feature vector and the positions of feature center points of another of the encoded maps. The further function may depend on the distance between the respective feature center points. An example of such a mechanism is cross-attention. The parameters of these functions are learned through a training process, which is described further below.

[0112] The result of the matching descriptor layer is a joint global map and a joint local map (the "joint map"), which contains an encoded feature vector for each feature center point of the two encoded maps.

[0113] The joint map is input to an optimal matching layer, where the similarity between the feature vectors of the feature center points in the joint global map and the feature vectors of the feature center points in the joint local map is calculated. Based on the similarity values ​​as a second step, a soft assignment matrix containing the matching probability for each pair of features is calculated. The soft assignments are found by solving an optimal transportation problem formulated as a sum maximization problem. We call the values ​​of this soft assignment matrix point correspondences.

[0114] The soft-assignment matrix has as many columns (or rows) as the number of cells in the encoded global map, as many rows (or columns) as the number of cells in the encoded local map, and one or more dust bin rows and one or more dust bin columns for indicating that a first feature vector does not correspond to any second feature vector. The initial values ​​of the dust bin rows and dust bin columns are parameters learned during training. The Sinkhorn algorithm can be applied to this initial similarity matrix to obtain the soft-assignment matrix. The top k values ​​(k is a predetermined constant) and / or values ​​greater than a threshold (correlation threshold) can be selected from the soft-assignment matrix to identify cells whose feature vectors correspond to each other. The output of the optimal matching layer is a set of corresponding point pairs.

[0115] The transformation layer extracts the coordinates of the feature centers of cells in the encoded global map and the coordinates of the feature centers of corresponding cells in the encoded local map (i.e., for cells assumed to correspond to each other according to the soft assignment matrix) and calculates the transformation between these feature centers (i.e., position t and orientation R). If there are multiple correspondences, they are jointly considered by an algorithm such as a least-squares algorithm. Through the transformation, each pose in the local coordinate system can be transformed into the corresponding pose in the global coordinate system (and vice versa).

[0116] Modality Fusion According to Some Exemplary Embodiments An exemplary embodiment in which maps obtained by different modalities are fused, such as images in the visible range (photographic images), RF fingerprints, and LIDAR images, is described with reference to Fig. 3. In the exemplary embodiment of Fig. 3, maps obtained from photographic images and RF measurements are fused.

[0117] In an exemplary embodiment, the RF fingerprint is registered to the 6DoF camera pose at the time of image capture to generate a respective global or local map (point cloud). That is, because the calibration (i.e., distance and relative orientation) between the camera and RF antenna for receiving RF signals is known for each sensor (device), the pose of the RF fingerprint can be expressed in the camera's coordinate system. If the RF measurement is taken at a time when no photographic image was captured, and if the sensor may have moved since the last image capture, the pose of the RF measurement can be obtained by interpolation between the poses at the image captures before and after the RF measurement (or by extrapolation from the poses at the image captures before (or after) the RF measurement).

[0118] From these joint measurements, a joint feature vector is generated by fusing spatial information from the point cloud and the RF fingerprint. If registration fails due to non-overlapping maps or if there is not enough information to disambiguate the match, the local map can be expanded by acquiring more images until a match is successful. The RF fingerprint may comprise, for example, the BSSID and / or RSSI values ​​of one or more RF transmitters, such as WiFi APs.

[0119] During the data collection phase, image-based position estimation (in the respective coordinate system) determines the pose of the RF fingerprint measurements. Therefore, an RF map is obtained, each containing RF fingerprints with a specific pose. The RF map can be considered as a 3D point cloud of modality RF measurements. The 3D coordinates are the locations where the measurements were made, and the feature vector is composed of the measured RF fingerprints and the orientation of the measuring device. The RF map can be coded in the same way as a structural map.

[0120] As shown in FIG. 3, there are a global map pipeline 2a and a local map pipeline 2b. These pipelines include first and second (structural) encoder layers, as described with respect to the exemplary embodiment of FIG. 2. In addition, each of the pipelines includes a respective RF map encoder layer (the "third and fourth encoder layers") and a respective neural fusion layer. The RF map encoder layer and / or the neural fusion layer and their parameters may be the same or different for the two pipelines. Unlike FIG. 2, the encoded global map output from the structural encoder layer is input to the neural fusion layer of the respective pipeline, rather than to the matching descriptor layer. Pipelines 2a and 2b operate independently of each other.

[0121] The RF map encoder layer corresponds functionally to the structure encoder layer, but instead of features of the points of the point cloud encoded in the structure encoder layer, the RF map encoder layer encodes RF fingerprints with their poses. The same 3D grid as the structure encoder layer is used. The resulting encoded RF map includes feature centers for each cell with its respective encoded RF fingerprint. Here, the feature centers of the encoded global / local RF map may be the same as or different from the feature centers of the encoded global / local map. The parameters of the RF map encoder layer are learned through a training process, which is further described below.

[0122] Both the encoded global / local RF maps and the encoded global / local structural maps of each of pipelines 2a and 2b are input to the respective neural fusion layers. In the neural fusion layers, similar to the matching descriptor layer described with respect to the exemplary embodiment of FIG. 2, a function is applied to each feature vector of all feature centers of each encoded map. The function describes how other points of each encoded map affect a selected point of a map of the same modality. That is, the structural feature vector is updated based on the structural feature vector, and the RF feature vector is updated based on the RF feature vector. The function may depend on the distance between the respective feature centers. An example of such a mechanism is self-attention. Furthermore, a further function (which may be the same as or different from the further function used in the matching descriptor layer) is applied that describes how each point of the encoded map of one modality is affected by a point of the encoded map of the other modality, and vice versa. That is, the RF feature vector is updated based on the structural feature vector, and vice versa. The further function may depend on the distance between the respective feature centers. An example of such a mechanism is mutual attention. The parameters of these functions are learned through a training process described further below.

[0123] Furthermore, the neural fusion layer calculates a new feature center point ("joint feature center point") for each cell. For example, the new feature center point of a cell can be calculated as the arithmetic mean of the feature center points of the cell in the two encoded maps input to the neural fusion layer. Alternatively, one of the feature center points of one of the modalities can be selected as the new feature center point.

[0124] The result of processing by each neural fusion layer in the pipeline is a respective encoded fused global map, which includes a feature center point and an encoded feature vector for each cell of the 3D grid (where the cell contains at least one point). Thus, the encoded fused global / local map of FIG. 3 corresponds to the encoded global / local map of FIG. 2, but with less ambiguity due to the fusion of different modalities. The encoded fused global map and the encoded fused local map are input to a matching descriptor layer. In other words, in this exemplary embodiment, the output of the neural fusion layer can be considered as an encoded input map to the matching descriptor layer. The processing of further layers is the same as that described with reference to FIG. 2.

[0125] In some exemplary embodiments, not only two modalities but three or more modalities may be fused.

[0126] Hierarchical Model According to Some Example Embodiments Figure 4 shows an example embodiment including modality fusion (as described with respect to Figure 3) and a hierarchical setup with a 3D grid of increasing cell size. While Figure 4 includes fusion of different modalities (photographic images and RF measurements), in some example embodiments, point clouds of only one modality may be used (as in the example embodiment of Figure 2). While Figure 4 shows an example with three layers, the number of layers may be two, four, or more.

[0127] At each tier, the process is the same as that described with respect to FIG. 3. For simplicity, each map shown in FIG. 3 is not shown in FIG. 4. The 3D grids become coarser from tier 1 to tier 3, i.e., the cells of 3D Grid 1 (i.e., tier 1) are smaller than the cells of 3D Grid 2 (i.e., tier 2), which are smaller than the cells of 3D Grid 3 (i.e., tier 3). In the exemplary embodiment of FIG. 4, for the global map, Grid 3 at tier 3 has four cells (Cell 1, Cell 2, Cell 3, Cell 4), while Grid 2 at tier 2 has 16 cells, with each four cells in Grid 2 corresponding to one of the cells in Grid 3. Correspondingly, Grid 1 at tier 1 has 64 cells, with each four cells in Grid 1 corresponding to one of the cells in Grid 2. For the local map, each of the two cells A and B in Grid 3 of Layer 3 corresponds to each of the four cells in Grid 2 of Layer 2, and each of the eight cells in Grid 2 of Layer 2 corresponds to each of the four cells in Grid 1 of Layer 1. In other exemplary embodiments, the correspondence ratio between the number of corresponding cells in different layers (1:4 in FIG. 4 ) may differ from this value, such as, but not limited to, 1:2, 1:3, 1:5, 1:8, or 1:16.

[0128] The global map pipeline and the local map pipeline may operate independently of the other map pipelines. For each map, the pipelines (RF map encoder layer, structure encoder layer, neural fusion layer) operate in order from the lowest layer (finest 3D grid) to the highest layer (coarsest 3D grid). That is, the input to the RF map encoder layer in layer 2 is the output of layer 1 (encoded global / local RF maps), the input to the structure encoder layer in layer 2 is the output of layer 1 (encoded global / local maps), and so on up to higher layers. Furthermore, in some exemplary embodiments, in one or both pipelines, the output of the neural fusion layer in the lower layer may additionally be input to the RF map encoder layer and structure encoder layer in the higher layer. This option is indicated by dashed lines in FIG. 4.

[0129] In some exemplary embodiments, information about point correspondences in higher layers is input to a matching descriptor layer in a lower layer. Based on this information, the matching descriptor layer in a lower layer can exclude those cells from the calculation of mutual influence (e.g., self-attention mechanism and / or mutual attention mechanism). Therefore, although computational effort can be saved, the calculated matching descriptor may miss important information that was present in the feature vectors of the filtered points.

[0130] Alternatively, every feature point can be used to compute a matching descriptor, but the output feature points are filtered out based on a coarser level of point correspondence. In this case, the matching descriptor may be of higher quality, and some computational gain can still be obtained from the optimal matching layer, since the soft-assignment matrix is ​​computed among fewer matching descriptors.

[0131] For example, 3D Grid 3 in Layer 3 contains four coarse cells. Layer 3 identifies that two coarse cells in the global map are similar to two coarse cells in the local map, and that another coarse cell in the global map is similar to none of the coarse cells in the local map, and vice versa. Each of the coarse cells in 3D Grid 3 corresponds to four finer cells in 3D Grid 2 in Layer 2, for a total of 16 finer cells. With knowledge from Layer 3, Layer 2 must find similarities between eight finer cells in the global map and eight finer cells in the local map (64 calculations). Without input from Layer 3, Layer 2 must find similarities between 16 finer cells in the global map and 16 finer cells in the local map, resulting in 256 calculations.

[0132] Correspondingly, the point correspondences obtained in Layer 2 are input to the Layer 1 matching descriptor layer, which has finer cells, thereby saving computational effort in Layer 1 as well.

[0133] In some exemplary embodiments, soft-assignment matrices of a single layer (typically the lowest layer (Layer 1)) are used by the transform layer to extract each coordinate pair and estimate a transformation based on the extracted coordinate pair. In other exemplary embodiments, soft-assignment matrices of multiple layers (typically including the lowest layer (Layer 1)) are used by the transform layer to extract each coordinate pair and estimate a transformation based on the extracted coordinate pair. Each coordinate pair may have a different weight in estimating the transformation. The weight may depend, for example, on the coarseness of the grid (the coarser the grid, the lower the weight) and / or the respective values ​​of the soft-assignment matrix (the higher the value, the higher the weight).

[0134] In some exemplary embodiments, one or more of the layers may not include a matching descriptor layer and an optimal matching layer, in which case only the processing in the pipeline is performed.

[0135] Figure 5 shows a cutout of the functional structure of the exemplary embodiment of Figure 4, including only layers 1 and 2. The output layer (including the transformation layer) is omitted. Figure 5 illustrates the terms used in the claims. These terms apply correspondingly to the claims relating to the exemplary embodiments of Figures 2 and 3.

[0136] [Training stage] Each of the exemplary embodiments of Figures 2-4 must be trained prior to estimation (i.e., the query phase). For training, a global map is compared to a local map for which the ground truth is known. For example, such a local map may be generated from a subset of the images from which the global map was generated, using a mechanism such as SLAM to generate the respective point clouds of the local map. However, the local maps of the training dataset are not limited to point clouds obtained from a subset of images.

[0137] During the training phase, the point clouds of the global and local maps (potentially according to different modalities, depending on the implementation) and the 3D grid are input to the system. Two loss terms can be calculated: The first loss function, L_corr, attempts to maximize the point correlation of corresponding point pairs. A second loss function L_geo between the computed transformation and the known transformation.

[0138] These loss terms can be weighted and combined for a loss function L (e.g., L=α*L_corr m +β*L_geo nwhere α, β, m, and n are predetermined constant values, e.g., m=n=1; other combinations of L_corr and L_geo are also possible. Through iterations, the parameters of the different layers (the structural encoder layer, the neural fusion layer (if present), and the matching descriptor layer) are trained so that the predicted transformations adapt sufficiently well to the actual transformations (ground truth). Typically, a number of epochs of iterations are performed until both the training loss function (the error of the model based on the training set used to initially train the model) and the validation loss function decrease and become smaller than their respective thresholds. Optionally, training may be stopped after a predetermined number of iterations have been performed.

[0139] In the training phase of the exemplary embodiment of Figure 4, point correspondences in higher layers may not be considered in the processing by the matching descriptor layer and the optimal matching layer in lower layers, i.e., the computational gain (saved computational effort) may be obtained only in the estimation (query) phase.

[0140] Aspects of Some Exemplary Embodiments First, some aspects of some exemplary embodiments will be described from different perspectives below, and it should be noted that these aspects are not limiting.

[0141] Some example embodiments may be used to solve problems such as loop closure, relocalization, and co-localization in a SLAM pipeline. Some example embodiments may have one or more of the following highlights: The problems of relocalization, colocalization, and loop closure are solved by using a sequence of images (instead of a single image) to create a 3D point cloud representation of the environment using 3D reconstruction algorithms. Alternatively, one can utilize a lidar point cloud, for example. To increase robustness and make the use of large maps feasible, six DoF poses, called RF fingerprints, are registered with the RSSI values ​​of different APs. Based on the structural information of the environment and the registered RF fingerprints, highly discriminative spatial matching descriptors (previously referred to as joint feature vectors) of different granularities are created in a data-driven manner using a fully distinguishable and trainable model. These new descriptors can help solve the correspondence problem between point clouds. By obtaining accurate correspondences, a rigid transformation can be calculated to align the point clouds. To create the spatial matching descriptor, a fusion layer is used, which creates a fused feature vector (previously called a joint feature vector) by an attention graph neural network architecture, e.g., using self-attention and cross-attention to exchange information between the structural coding and the RF coding. Fading and geometric embedding are defined for the attention stage, which implicitly encodes RF propagation characteristics locally, resulting in better RF signal interpolation and extrapolation. In the encoding layer, the 3D positions of feature vectors are similarly regressed in a data-driven, differentiable, and learnable manner to achieve high accuracy of rigid transformation estimation even on the coarsest layers.

[0142] Some example embodiments accumulate localization information from a sequence of images instead of using a single query image. By capturing the 3D structure of the environment and determining the camera pose, the RF fingerprint can be registered to the 6DoF camera pose. From these joint measurements, a spatial matching descriptor can be created by training a deep learning model that fuses spatial information from the point cloud and the RF fingerprint. Such a spatial matching descriptor can be used to efficiently solve the map-to-map registration problem. If registration fails due to non-overlapping maps or if there is not enough information to disambiguate the match, the local map can be expanded by acquiring more images until the match is successful.

[0143] Some exemplary embodiments significantly reduce the computational effort by creating discriminative hierarchical spatial matching descriptors that reduce the search space by first registering matching descriptors at the coarsest level, and continuing the registration process only in regions that belong to successfully registered coarse spatial matching descriptors.

[0144] To create robust and discriminative spatial matching descriptors for feature points, the system of some exemplary embodiments is trainable end-to-end via backpropagation, meaning that all formulas for all building blocks are distinguishable. Based on the positions of these matching descriptors for different point clouds, a rigid body transformation for point cloud alignment can be estimated. The deep learning model receives as input, for example, a (previously optimized) global map including the registered RF fingerprint map, a 3D point cloud describing the structure of the environment, and a bundle-adjusted local map with actual RF fingerprint measurements. These maps are encoded through a series of layers: a structure and RF map encoder layer, a fusion layer, a spatial matching descriptor layer, an optimal matching layer, and a transformation layer.

[0145] In the encoder layer, the input map data passes through a series of convolution, normalization, activation, and data compression layers. In some exemplary embodiments, point convolution may be used to encode the point cloud features and RF map features, although other types of convolution may be used in different embodiments. A mathematical definition of point convolution is provided further below. For normalization, several approaches can be applied. For example, data compression may be achieved through a grid division mechanism. A series of grids are defined with increasing grid cell sizes. In the data compression layer, the space is divided by the selected grid, and the location of the output feature vector is defined based on the data points belonging to each grid cell. These points are called feature centroids. In some embodiments, the output location (feature centroid) for each grid cell may be calculated as the arithmetic mean of the grid cell data point locations. The location of the feature vector is very important because the rigid transformation (which aligns the two point clouds) is estimated based on the location of the feature vector. Therefore, several convolution layers are used to regress the location of the output feature vector. In the case of point convolution, once the output position (feature center point) is determined, supporting data points are selected based on some neighborhood criteria (optional). Typically, a radius is defined around the output position, or the kth point closest to the output position is selected. Based on the feature vectors of the supporting points, a feature vector for the output feature point is created. It is important to note that several point convolution iterations can occur between two compression layers. In this case, the position of the output feature point is the same as the position of the input feature point.

[0146] After every compression layer, input / output data point connections can be stored. During estimation, points at the highest / coarsest level are registered first. After coarse registration, an initial rigid transformation is estimated based on the 3D positions of the feature points, and is iteratively refined by further extending the aligned points and computing more refined transformations. This hierarchical registration approach eliminates the need to load the entire set of feature points in the global localization map. Point registration drives the loading of the relevant parts of the global map. Furthermore, only the relevant parts of the global map are fed into the matching layer of the deep learning model, significantly reducing computational complexity.

[0147] A hierarchical approach is particularly effective when the spatial matching descriptors of feature points are highly discriminative. This can be achieved by converting RF fingerprints into RF spatial encodings, converting 3D point clouds into structural encodings, and then fusing these encodings through distinct layers. Although 3D structures can be repetitive in industrial environments, fingerprint maps can reduce location ambiguity. Furthermore, by encoding structural representations using RF signal measurements, RF signal propagation can be better described compared to existing models (e.g., Rayleigh fading, Rician fading), because fading characteristics are implicitly encoded in the network. If the pose of the access point (AP) is known, geometric encoding of the pose of the RF measurements along with the AP pose is also possible.

[0148] The fusion layer can be built on an attention graph neural network architecture, including a self-attention stage and a mutual attention stage. In the self-attention stage, feature points of the same type share information with each other through the attention mechanism, while in the mutual attention stage, structural coding shares information with RF spatial coding. In the self-attention stage, an embedding term is introduced that encodes channel characteristics in the form of a path loss exponent, the relative distance and orientation of current measurements from their corresponding APs. In the mutual attention stage, the positions of the RF spatial coding and structural coding can be different, so a distance-based embedding is defined. The output of the fusion layer is a feature vector with 3D positions that encodes structural and RF signal information. The exact mathematical formulation is presented further below.

[0149] The fused spatial feature vector is fed to the spatial matching descriptor layer. This is the first layer where the fused spatial feature vectors of the two point clouds to be matched can exchange information. An attention graph neural network architecture is used to drive the information exchange. The output of this layer is called the spatial matching descriptor. Feature points that describe the same region in the global and local maps are expected to have similar descriptors.

[0150] In the optimal mashing layer, a pairwise score matrix is ​​calculated as the similarity of the spatial matching descriptors. The pairwise score can be calculated as the dot product or Gaussian correlation between feature vectors. The point matching problem can then be formulated as an optimal transport problem that maximizes the soft-assignment matrix via a differentiable sinkhorn algorithm. Local point correspondences are extracted by reciprocal top-k selection on the values ​​of the soft-assignment matrix.

[0151] In the transform layer, based on point correspondences, rigid transformation estimation is formulated as a least-squares problem that can be solved in closed form using differentiable weighted SVD. This layer has no learnable parameters, but is differentiable, allowing us to define a loss function between the estimated transformation and the ground truth transformation.

[0152] For soft assignment matrices, a different loss function (loss term) may be used, where the basis assignment matrices can be derived from the ground truth alignment transformation between point clouds (which is known during the training phase).

[0153] A deep learning model built based on the above layers can solve the problems of relocalization, loop closure, and colocalization. For relocalization, a global map and a local map are given as inputs to the deep learning model. Since the global map can be the result of a SLAM pipeline, hierarchical fusion features can be pre-computed. Therefore, at the time of estimation, only the local map can be encoded, and structural and RF features can be fused. For loop closure detection and colocalization, two local maps are inputs to the deep learning model. In this case, feature points are always extracted from the local map because it is continuously changing, and the changes can propagate to the highest level (global map).

[0154] Image-based localization problems are known to be difficult to solve because lighting conditions encountered during the day or seasonal changes in the environment introduce different artifacts. Creating visual descriptors and models that can generalize to all conditions is particularly challenging. Image retrieval approaches to solving localization problems are assumed to be non-robust due to the aforementioned issues. By creating a local map for structural representation and registering RF measurements with the camera pose, a better representation can be constructed for localization purposes. After a global map is constructed, different deep learning models can be trained to register features from the local maps created under different conditions. Nevertheless, it is also possible to train a single model that generalizes to all conditions by creating a set of local maps of the same area at different times of the day and / or seasons.

[0155] Location estimation, image retrieval algorithms, and data-driven training models can be significantly affected by camera lens distortion. When query images are created with a different camera than the keyframe images, only suboptimal accuracy is achieved. Image distortion removal techniques can help mitigate this problem, but results may still be suboptimal. For data-driven algorithms, the data distribution is defined by the images in the training set. To create robust algorithms, a good approach is to collect training images from a wide set of cameras with different distortion effects. This is useful because global maps are typically created using cameras with similar distortion parameters, but a wide variety of user devices must be supported during location estimation. Some exemplary embodiments are camera-independent because camera image sequences are converted into point clouds.

[0156] DETAILED DESCRIPTION OF ILLUSTRATIVE EMBODIMENTS An exemplary embodiment corresponding to Fig. 4 will be described in detail below. Note that this embodiment is not limiting.

[0157] Map definition The map produced by the SLAM pipeline consists of (i.) 3D map points

number

number

number

number

[0158] Detailed RF measurements are captured by an agent containing an RF module and a camera. A SLAM map is constructed based on the image sequence captured by the camera and registers the RF measurements to the camera's 6DoF pose. The radio fingerprint includes the average received signal strength (RSS) from surrounding WiFi access points (APs), base stations, and Bluetooth devices. If the number of access points in the mapped area is K, the radio map

number

number

number

number

number

number

number

[0159] Problem definition The goal is to align a local SLAM map within a global SLAM map, or to align two different local SLAM maps, but for simplicity, we may consider two point clouds.

number

number

number

number

number

number

number

[0160] Neural Network Building Blocks Map Encoder Layer Encoding point clouds is challenging because the data is not well-structured like image data. Several encoding techniques exist: (i.) projection networks, in which points are projected onto regular 2D and 3D grid structures followed by regular 2D or 3D convolution operators; (ii.) graph convolution networks; (iii.) point-wise multilayer perceptron networks; or (iv.) point convolution networks. The point cloud encoder layer in different embodiments may implement different point cloud encoding strategies. The main goal is to create a representation that encodes the structure of the environment and the measured RF signals in a hierarchical and transformation-invariant manner.

[0161] In this section, the details of the encoder layer are formalized. The encoder layer is applied to the structure map and the RF map in parallel. The encoder layer is defined by a set of convolution operations, normalization operations, activation operations, and data compression operations.

[0162] Convolution: Because point cloud data is unstructured, the convolution operator is redefined. In this embodiment, the following approach to encoding point cloud features and RF map features is adopted, but other embodiments may use different types of convolution.

[0163] point

number

number

number

number

number

[0164] N x is the neighborhood of a central point which can be defined in several ways. The most common approach is to take the k nearest neighbors of the central point x, or to define a sphere with radius r around the central point, and thus the neighborhood set is

number

number

number

number

number

number

number

number

number

[0165] In embodiments where the encoder layer is formulated by successive point convolution operations, the output position x is the same for all input point positions, so the number of evaluated convolutions is equal to N.

[0166] Regularization: Normalization layers play an important role in increasing the convergence of training. Several different techniques can be applied.

[0167] Data Compression: For large point clouds, it is impractical to create a model using only point convolution, since the number of convolution evaluations is equal to the number of input points. By applying data compression, the number of points is reduced and becomes better organized structurally. Data compression is achieved through a grid division mechanism. A set of 3D grids is defined with increasing cell size. When data compression is performed, the goal is to define: (i.) structural and RF descriptors encoded as feature vectors per grid cell, and (ii.) regress feature (center) point coordinates on the structural and RF descriptors to be used in computing the rigid body transformation T. At lower levels, the grid cell size decreases, and as the coarsest resolution is approached, the grid cell size increases.

[0168] The point compression operation for one grid configuration is formulated as follows: (i.) define a 3D grid with a given grid cell size; (ii.) group input points based on the grid cells; (iii.) regress the positions of output feature center points and any additional parameters of the selected convolution type; (iv.) define support points for all feature center points (the set of support points is the set of neighbors of the feature center points, typically defined by the radius of the convolution or based on a neighbor selection algorithm; further, the radius of the convolution may be equal to, greater than, or smaller than the size of the grid cell); and (vi.) create a structural or RF descriptor for each grid cell based on the points within the cell by performing several groups of convolution, normalization, and activation operations.

[0169] For further data compression, the grid cell size is increased and point cloud encoding continues as described above. Structural descriptors and centroid locations are created in a data-driven manner to produce good feature vectors for point registration and pose estimation.

[0170] Using the kernel point convolution presented above, the desired structural and RF descriptors and output feature center points can be created using a two-stage kernel point convolution. In the first stage, grid cell C j Center point of

number

number

number

number

[0171] Encoding of the RF map (i.e., RF measurements are recorded at 3D locations) is performed similarly. Because fingerprint measurement locations can be structurally sparser than the reconstructed point cloud (and note that RF measurement 3D locations are not the same as structural 3D points), it may be beneficial if hierarchical encoding occurs in fewer stages, although this is a design choice that may vary in different exemplary embodiments. It is recommended to use the same grid configuration for the layer in which RF features are fused with structural features. Feature center points in the RF map and structural map are typically disjoint locations. Therefore, this must be addressed in the fusion process. Furthermore, when applying a data compression layer to the RF map, measurement points are selected as feature center points instead of regressing feature center points. An obvious choice would be the measurement point with the most valid RSS measurement in the fingerprint feature vector or the measurement point closest to the arithmetic mean of the measurement locations within the grid cell. The motivation for selecting measurement points as feature center points is the availability of the RF fingerprint feature vector used to create the RF embedding in the attention mechanism. Estimating the RF fingerprint vector at new locations introduces inaccuracies. Center point selected measurement point

number

[0172] Hierarchical features are created by iteratively applying data compression. After defining feature centers for a new grid configuration, point-to-node grouping is performed. Point-to-node grouping can be implemented based on several heuristics. All points

number

number

number

number

[0173] The point registration process starts at the coarsest level, and the rigid pose transformation is refined in an iterative manner by expanding the registered nodes and recomputing the refined pose. More formally, the feature center points from the point cloud P are

number

number

number

number

[0174] ·Fusion layer The RF and structural feature fusion architecture can be designed in many ways. In one embodiment, features at one level with the same grid cell size are concatenated, followed by several convolutional layers.

[0175] In this exemplary embodiment, a graph neural network architecture with an attention mechanism is used. The attention mechanism is estimated to be able to better capture the propagation characteristics of RF signals in a large environment compared to a group of convolutional layers because information sharing occurs among all points where valid RSS measurements are made. By fusing sparse RF coding with structural coding, better prediction of RF signal propagation and a more specific description of any 3D location in the world are expected.

[0176] A series of self-attention and mutual-attention modules are employed to share structural information between feature centers. By aggregating this information, a feature vector is created that is expected to have better global context information, which eases the registration problem.

[0177] Structural Feature Self-Attention: In the following, we formalize a single-head self-attention mechanism for structural feature vectors (in another exemplary embodiment, a multi-head attention mechanism can be used).

[0178] Structural feature vector of layer l

number

number

number

number

number

number

number

number

number

number

number

number

[0179] The relative distance embedding is done by applying sinusoidal coding to the coordinates

number

number

number

number

number

number

number

[0180] RF self-attention: An RF embedding mechanism is used to generate a transformation-invariant RF propagation embedding. The RF feature vector at layer l

number

number

number

number

number

number

[0181] Feature Vector

number

number

number

number

[0182] If the AP pose is registered, the RF embedding consists of three terms: distance (relative position) encoding

number

number

number

[0183] RF embedding term

number

number

number

[0184] Measurement point

number

number

number

[0185] The sinusoidal encoding of the relative orientation between the ith measurement point and the jth AP is encoded as follows:

number

number

[0186] In fading embedding, the goal is to implicitly encode the path loss exponent. The log-distance path loss model is formally defined as follows:

number

number

number

number

number

number

number

number

number

number

[0187] The fading embedded sinusoidal coding is formulated as follows:

number

[0188] Mutual attention of RF and structural features: Output RF feature vectors from the self-attention encoding stage

number

number

[0189] The update of the RF feature vector in the first step is formulated as follows:

number

number

[0190] The update of the structural feature vector in the second step is formulated as follows:

number

number

[0191] The mutual attention stage is where the structural and RF information is shared and fused. After several iterations of the self-attention and mutual attention steps, the final fused feature vector

number

[0192] In other exemplary embodiments, the RF center point can be similarly regressed and used as the output of the fusion layer.

[0193] For every iteration in which the attention mechanism is computed, a new trainable projection matrix is ​​used.

[0194] Matching Descriptor Layer The matching descriptor layer is responsible for exchanging information between the fused feature vectors of two different point clouds P and Q and creating matching descriptors that are used to calculate similarity scores for different regions in the point clouds.

[0195] The fused feature vector of point cloud P at layer l

number

number

number

number

number

number

[0196] In the self-attention step, the feature vector is updated as follows:

number

number

number

number

number

number

[0197] In the mutual attention stage, the attention mechanism is applied as follows:

number

[0198] In the mutual attention stage, weighting factors are calculated based on the mutual attention scores.

number

[0199] Mutual attention features

number

[0200] Finally, the attention features are transformed into a projection matrix

number

number

number

number

number

[0201] In other exemplary embodiments, a multi-head architecture may be feasible, in which case the number of projection matrices increases.

[0202] For every iteration, a new trainable projection matrix is ​​used when the attention mechanism is computed.

[0203] Optimal matching layer In the optimal matching layer, the matching descriptor is

number

number

[0204] To be able to mark certain feature points as unregistered, we use the matrix

number

number

number

number

number

number

[0205] Soft allocation matrix A l can be computed in a differentiable way by the Sinkhorn algorithm. The local point correspondences are calculated using the soft assignment matrix A l are extracted by mutual top-k selection on the values ​​of

[0206] Conversion layer By defining correspondences between matching feature points, rigid transformation estimation is formulated as a least-squares problem that can be solved in closed form using differentiable weighted SVD. This layer has no learnable parameters, but because it is differentiable, a loss function can be defined between the estimated transformation and the ground truth transformation.

[0207] The accuracy of the estimated pose can be further improved by expanding the feature points in layer l and computing matching between the expanded features in lower layers. The accuracy improves because the number of feature points is large and the quantization effect of the center selection process results in less error. Registration in lower layers is only performed between points whose parent nodes have been previously registered. This mechanism allows fast registration to be achieved even for large maps.

[0208] Loss function Since every layer presented previously is fully differentiable, training a deep learning model involves predicting the rigid transformation at layer l.

number

number

[0209] A possible geometric loss function can be formulated for layer l as follows:

number

[0210] Soft allocation matrix A l In the training phase, the log-likelihood loss for each point correspondence is defined.

number

number

number

[0211] The total loss can be formulated as follows:

number

number

[0212] [General embodiment] FIG. 6 illustrates an apparatus according to an exemplary embodiment. The apparatus may be a location estimation device or an element thereof. FIG. 7 illustrates a method according to an exemplary embodiment. The apparatus according to FIG. 6 may execute the method of FIG. 7, but is not limited to this method. The method of FIG. 7 may be executed by the apparatus of FIG. 6, but is not limited to being executed by this apparatus.

[0213] The apparatus comprises first to fourth input means 110, 130, 140, 160, first and second encoding means 120, 150, first and second adapting means 170, 180, first, second and third calculating means 190, 200, 230, checking means 210, and extracting means 220. The first to fourth input means 110, 130, 140, 160, first and second encoding means 120, 150, first and second adapting means 170, 180, first, second and third calculating means 190, 200, 230, checking means 210, and extracting means 220 may be the first to fourth input means, first and second encoding means, first and second adapting means, first, second and third calculating means, checking means, and extracting means, respectively. The first to fourth input means 110, 130, 140, 160, the first and second encoding means 120, 150, the first and second adapting means 170, 180, the first, second and third calculating means 190, 200, 230, the checking means 210 and the extracting means 220 may be first to fourth input units, first and second encoders, first and second adapters, first, second and third calculators, checkers and extractors, respectively. The first to fourth input means 110, 130, 140, 160, the first and second encoding means 120, 150, the first and second adaptation means 170, 180, the first, second and third calculation means 190, 200, 230, the check means 200 and the extraction means 210 may be first to fourth input processors, first and second encoding processors, first and second adaptation processors, first, second and third calculation processors, check processors and extraction processors, respectively.

[0214] The first input means 110, the first encoding means 120, and the second input means 130 constitute a first pipeline. The third input means 140, the second encoding means 150, and the fourth input means 160 constitute a second pipeline. The work of the first pipeline may be performed independently of the work of the second pipeline, and the work of the second pipeline may be performed independently of the work of the first pipeline. The work of the first pipeline may be performed in any time sequence relative to the work of the second pipeline. The work of the first pipeline may be performed fully or partially in parallel with the work of the second pipeline.

[0215] A first input means 110 inputs a first 3D point cloud of a first modality and a first grid to a first encoder layer (S110). The first grid includes one or more first cells. Coordinates of points in the first 3D point cloud are expressed in a first coordinate system.

[0216] The first encoding means 120 encodes the first 3D point cloud of the first modality into a first encoded map of the first modality using a first encoder layer (S120). The first encoded map of the first modality includes, for each first cell of the first grid, a respective first feature center point and a respective first feature vector. The coordinates of the first feature center point are represented in a first coordinate system.

[0217] The second input means 130 inputs a first coded input map to the first matching descriptor layer of the first hierarchy (S130). The first coded input map is based on the first coded map of the first modality. The first coded input map includes, for each of the first cells, a respective first joint feature center point and a respective first input feature vector. The coordinates of the first joint feature center point are represented in a first coordinate system.

[0218] A third input means 140 inputs a second 3D point cloud of the first modality and a second grid including one or more second cells to the second encoder layer (S140). Coordinates of points in the second 3D point cloud are represented in a second coordinate system. The second coordinate system may or may not be different from the first coordinate system.

[0219] The second encoding means 150 encodes the second 3D point cloud of the first modality into a second encoded map of the first modality (S150) using a second encoder layer. The second encoded map of the first modality comprises, for each second cell of the second grid, a respective second feature center point and a respective second feature vector. The coordinates of the second feature center points are represented in a second coordinate system.

[0220] A fourth input means 160 inputs a second coded input map to the first matching descriptor layer (S160). The second coded input map is based on the second coded map of the first modality. The second coded input map includes, for each second cell, a respective second joint feature center point and a respective second input feature vector. The coordinates of the second joint feature center point are represented in a second coordinate system.

[0221] After the two pipeline activities S110 to S130 and S140 to S160 have been performed, activities S170 to S230 may follow.

[0222] The first adaptation means 170 adapts each of the first input feature vectors based on the first input feature vector and the second input feature vector by the first matching descriptor layer to obtain a first joint map (S170). Optionally, the first matching descriptor layer may adapt each of the first input feature vectors based on the first input feature vectors having their first feature center points and / or based on the second input feature vectors having their second feature center points. The first joint map includes a respective first joint feature vector for each of the first joint feature center points.

[0223] The second adaptation means 180 adapts each of the second input feature vectors based on the first input feature vector and the second input feature vector by the first matching descriptor layer to obtain a second joint map (S180). Optionally, the first matching descriptor layer may adapt each of the second input feature vectors based on the first input feature vectors having their first feature center points and / or based on the second input feature vectors having their second feature center points. The second joint map includes a respective second joint feature vector for each of the second joint feature center points.

[0224] The first calculation means 190 calculates, for each of the first cells and each of the second cells, the similarity between the first joint feature vector of each of the first cells and the second joint feature vector of each of the second cells by the first optimal matching layer of the first hierarchy (S190).

[0225] The second calculation means 200 calculates a first point correlation between each of the first cells and each of the second cells by the first optimal matching layer (S200). The first point correlation is calculated based on the similarity between the first joint feature vector of each of the first cells and the second joint feature vector of the second cell, and based on the similarity between the second joint feature vector of each of the second cells and the first joint feature vector of the first cell.

[0226] The checking means 210 checks whether at least one of one or more first correlation conditions is satisfied for each of the first cells and each of the second cells (S210). The one or more correlation conditions include: a first point correlation between each first cell and each second cell is greater than a first correlation threshold for the first tier; or The first point correlation between each first cell and each second cell is among the largest k values ​​of first point correlations, where k is a fixed value.

[0227] When the extraction means 220 checks that at least one of the one or more first correlation conditions for each of the first cells and each of the second cells is satisfied (S210=yes), it extracts, for each of the first cells and each of the second cells, the coordinates of the first joint feature center points of each of the first cells in the first coordinate system and the coordinates of the second joint feature center points of each of the second cells in the second coordinate system, to obtain each pair of extracted coordinates (S220).When the extraction means 220 checks that at least one of the one or more first correlation conditions for each of the first cells and each of the second cells is not satisfied (S210=no), it may not extract, for each of the first cells and each of the second cells, the coordinates of the first joint feature center points of each of the first cells in the first coordinate system and the coordinates of the second joint feature center points of each of the second cells in the second coordinate system.

[0228] Based on the coordinate pairs extracted in S220, the second calculation means 230 calculates an estimated transformation between the first coordinate system and the second coordinate system (S230).

[0229] 8 illustrates an apparatus according to an exemplary embodiment, comprising at least one processor 810 and at least one memory 820 storing instructions that, when executed by the at least one processor 810, cause the apparatus to perform at least the method according to FIG.

[0230] Unless otherwise stated or clear from the context, a statement that two entities are different means that they perform different functions. It does not necessarily mean that they are based on different hardware. That is, each of the entities described herein may be based on different hardware, or some or all of the entities may be based on the same hardware. It does not necessarily mean that they are based on different software. That is, each of the entities described herein may be based on different software, or some or all of the entities may be based on the same software. Each of the entities described herein may be deployed in the cloud.

[0231] Thus, in accordance with the above description, it will be apparent that examples provide, for example, a neural network or components thereof, apparatus embodying the same, methods for controlling and / or operating the same, computer programs for controlling and / or operating the same, and media carrying such computer programs and forming computer program products.

[0232] Implementations of any of the blocks, apparatus, systems, techniques, or methods described above include, by way of non-limiting example, implementations as hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controllers or other computing devices, or any combination thereof. Each of the entities described herein may be embodied within a cloud.

[0233] It should be understood that the foregoing are what are presently considered to be preferred exemplary embodiments. It should be noted, however, that the description of the preferred exemplary embodiments is given by way of example only, and that various modifications may be made without departing from the scope of the present invention as defined by the appended claims.

[0234] The terms "first X" and "second X," unless otherwise specified, include the option where the "first X" is the same as the "second X," and the option where the "first X" is different from the "second X." As used herein, "at least one of" is similar to phrases such as "a list of two or more elements" and "at least one of ," where a list of two or more elements is joined by "and" or "or," means at least any one of the elements, or at least any two or more of the elements, or at least all of the elements.

Claims

1. 1. An apparatus comprising one or more processors and a memory storing instructions that, when executed by the one or more processors, cause the apparatus to: inputting a first 3D point cloud of a first modality and a first grid including one or more first cells into a first encoder layer, wherein coordinates of points of the first 3D point cloud are expressed in a first coordinate system; encoding, by the first encoder layer, the first 3D point cloud of the first modality into a first encoded map of the first modality, the first encoded map of the first modality including, for each of the first cells of the first grid, a respective first feature center point and a respective first feature vector, the coordinates of the first feature center points being expressed in the first coordinate system; inputting a first coded input map into a first matching descriptor layer of a first hierarchy, the first coded input map being based on the first coded map of the first modality and comprising, for each of the first cells, a respective first joint feature center point and a respective first input feature vector, the coordinates of the first joint feature center points being represented in the first coordinate system; inputting a second 3D point cloud of the first modality and a second grid including one or more second cells into a second encoder layer, wherein the coordinates of points of the second 3D point cloud are expressed in a second coordinate system; encoding, by the second encoder layer, the second 3D point cloud of the first modality into a second encoded map of the first modality, the second encoded map of the first modality including, for each of the second cells of the second grid, a respective second feature center point and a respective second feature vector, the coordinates of the second feature center points being expressed in the second coordinate system; inputting a second coded input map into the first matching descriptor layer, the second coded input map being based on the second coded map of the first modality and comprising, for each of the second cells, a respective second joint feature center point and a respective second input feature vector, the coordinates of the second joint feature center points being represented in the second coordinate system; adapting, by the first matching descriptor layer, each of the first input feature vectors based on the first input feature vectors, optionally their first feature center points, and the second input feature vectors, optionally their second feature center points, to obtain a first joint map, the first joint map including, for each of the first joint feature center points, a respective first joint feature vector; adapting, by the first matching descriptor layer, each of the second input feature vectors based on the first input feature vectors, optionally their first feature center points, and the second input feature vectors, optionally their second feature center points, to obtain a second joint map, the second joint map including, for each of the second joint feature center points, a respective second joint feature vector; calculating, for each of the first cells and each of the second cells, a similarity between the first joint feature vector of each of the first cells and the second joint feature vector of each of the second cells by a first optimal matching layer; calculating, by the first optimal matching layer, for each of the first cells and each of the second cells, a first point correlation between each of the first cells and each of the second cells based on the similarity between the first joint feature vector of each of the first cells and some of the second joint feature vectors of some of the second cells, and based on the similarity between the second joint feature vector of each of the second cells and some of the first joint feature vectors of some of the first cells; for each of the first cells and each of the second cells, checking whether at least one of one or more first correlation conditions is satisfied for each of the first cells and each of the second cells, wherein the one or more first correlation conditions include: the first point correlation between each of the first cells and each of the second cells is greater than a first correlation threshold for the first tier; or the first point correlation between each of the first cells and each of the second cells is among the highest k values ​​of the first point correlation, where k is a fixed value; extracting, for each of the first cells and each of the second cells, the coordinates of the first joint feature center points of the respective first cells in the first coordinate system and the coordinates of the second joint feature center points of the respective second cells in the second coordinate system, and obtaining, for each of the first cells and each of the second cells, each pair of extracted coordinates of the first hierarchy in response to checking whether at least one of the one or more first correlation conditions is satisfied; calculating an estimated transformation between the first coordinate system and the second coordinate system based on the extracted coordinate pairs of the first layer; An apparatus characterized by executing the above.

2. the first encoder layer, the second encoder layer, and the first matching descriptor layer of the first hierarchy are configured with respective first parameters, and the instructions, when executed by the one or more processors, further cause the apparatus to: training the device to determine the first parameters such that a deviation between the estimated transformation and a known transformation between the first coordinate system and the second coordinate system is less than a transformation estimation threshold and such that the first point correlation is increased; 2. The apparatus according to claim 1, wherein the apparatus executes the following steps:

3. the first modality includes one of a photographic image, a LIDAR image, and an ultrasound image; 3. The device according to claim 1 or 2.

4. 3. The apparatus of claim 1, wherein the first encoder layer is the same as the second encoder layer.

5. The instructions, when executed by the one or more processors, further cause the device to: generating the first 3D point cloud based on a sequence of first images of the first modality, wherein a respective position in the first coordinate system is associated with each of the first images; or generating the second 3D point cloud based on a sequence of second images of the first modality, wherein a respective position in the second coordinate system is associated with each of the second images; 3. The apparatus according to claim 1, wherein the apparatus executes the following steps:

6. 2. The apparatus of claim 1, wherein the first coded input map is the first coded map, the first joint feature center points are the first feature center points, the second coded input map is the second coded map, and the second joint feature center points are the second feature center points.

7. The instructions, when executed by the one or more processors, further cause the apparatus to: inputting a third 3D point cloud of a second modality and the first grid into a third encoder layer, wherein the coordinates of points of the third 3D point cloud are expressed in the first coordinate system, and the second modality is different from the first modality; encoding, by the third encoder layer, the third 3D point cloud of the second modality into a third encoded map of the second modality, the third encoded map of the second modality comprising, for each of the first cells of the first grid, a respective third feature center point and a respective third feature vector, the coordinates of the third feature center points being expressed in the first coordinate system; inputting the third encoded map into a first neural fusion layer; inputting the first encoded map into the first neural fusion layer; generating, by the first neural fusion layer, the first coded input map based on the first coded map and the third coded map, wherein for each of the first cells, each of the first joint feature center points is based on the first feature center points of each of the first cells and the third feature center points of each of the first cells, each of the first input feature vectors is based on the first feature vector and the first feature vector, optionally the first feature center points and the third feature vector of each of the first cells adapted by their first feature center points, optionally their third feature center points, and based on the third feature vector and the first feature vector, optionally the third feature center points and the third feature vector of each of the first cells adapted by their first feature center points, optionally their third feature center points; inputting a fourth 3D point cloud of the second modality and the second grid into a fourth encoder layer, wherein the coordinates of points of the fourth 3D point cloud are represented in the second coordinate system; encoding, by the fourth encoder layer, the fourth 3D point cloud of the second modality into a fourth encoded map of the second modality, the fourth encoded map of the second modality including, for each of the second cells of the second grid, a respective fourth feature center point and a respective fourth feature vector, the coordinates of the fourth feature center points being expressed in the second coordinate system; inputting the fourth encoded map into a second neural fusion layer; inputting the second encoded map into the second neural fusion layer; generating, by the second neural fusion layer, the second coded input map based on the second coded map and the fourth coded map, wherein for each of the second cells, each of the second joint feature center points is based on the second feature center points of each of the second cells and the fourth feature center points of each of the second cells, each of the second input feature vectors is based on the second feature vector and the second feature vector, optionally the second feature center points and the fourth feature vector of each of the second cells adapted by their second feature center points, optionally their fourth feature center points, and based on the fourth feature vector and the second feature vector, optionally the fourth feature center points and the fourth feature vector of each of the second cells adapted by their second feature center points, optionally their fourth feature center points; 10. The apparatus of claim 1, 2, or 6, further comprising:

8. the second modality includes at least one of a photographic image, an RF fingerprint, a LIDAR image, or an ultrasound image; 8. The device according to claim 7.

9. inputting a first 3D point cloud of a first modality and a first grid including one or more first cells into a first encoder layer, wherein coordinates of points of the first 3D point cloud are expressed in a first coordinate system; encoding, by the first encoder layer, the first 3D point cloud of the first modality into a first encoded map of the first modality, the first encoded map of the first modality including, for each of the first cells of the first grid, a respective first feature center point and a respective first feature vector, the coordinates of the first feature center points being expressed in the first coordinate system; inputting a first coded input map into a first matching descriptor layer of a first hierarchy, the first coded input map being based on the first coded map of the first modality and comprising, for each of the first cells, a respective first joint feature center point and a respective first input feature vector, the coordinates of the first joint feature center points being represented in the first coordinate system; inputting a second 3D point cloud of the first modality and a second grid including one or more second cells into a second encoder layer, wherein the coordinates of points of the second 3D point cloud are expressed in a second coordinate system; encoding, by the second encoder layer, the second 3D point cloud of the first modality into a second encoded map of the first modality, the second encoded map of the first modality including, for each of the second cells of the second grid, a respective second feature center point and a respective second feature vector, the coordinates of the second feature center points being expressed in the second coordinate system; inputting a second coded input map into the first matching descriptor layer, the second coded input map being based on the second coded map of the first modality and comprising, for each of the second cells, a respective second joint feature center point and a respective second input feature vector, the coordinates of the second joint feature center points being represented in the second coordinate system; adapting, by the first matching descriptor layer, each of the first input feature vectors based on the first input feature vectors, optionally their first feature center points, and the second input feature vectors, optionally their second feature center points, to obtain a first joint map, the first joint map including, for each of the first joint feature center points, a respective first joint feature vector; adapting, by the first matching descriptor layer, each of the second input feature vectors based on the first input feature vectors, optionally their first feature center points, and the second input feature vectors, optionally their second feature center points, to obtain a second joint map, the second joint map including, for each of the second joint feature center points, a respective second joint feature vector; calculating, for each of the first cells and each of the second cells, a similarity between the first joint feature vector of each of the first cells and the second joint feature vector of each of the second cells by a first optimal matching layer; calculating, by the first optimal matching layer, for each of the first cells and each of the second cells, a first point correlation between each of the first cells and each of the second cells based on the similarity between the first joint feature vector of each of the first cells and some of the second joint feature vectors of some of the second cells, and based on the similarity between the second joint feature vector of each of the second cells and some of the first joint feature vectors of some of the first cells; for each of the first cells and each of the second cells, checking whether at least one of one or more first correlation conditions is satisfied for each of the first cells and each of the second cells, wherein the one or more first correlation conditions include: the first point correlation between each of the first cells and each of the second cells is greater than a first correlation threshold for the first tier; or the first point correlation between each of the first cells and each of the second cells is among the highest k values ​​of the first point correlation, where k is a fixed value; extracting, for each of the first cells and each of the second cells, the coordinates of the first joint feature center points of the respective first cells in the first coordinate system and the coordinates of the second joint feature center points of the respective second cells in the second coordinate system, and obtaining, for each of the first cells and each of the second cells, each pair of extracted coordinates of the first hierarchy in response to checking whether at least one of the one or more first correlation conditions is satisfied; calculating an estimated transformation between the first coordinate system and the second coordinate system based on the extracted coordinate pairs of the first layer; A method comprising:

10. A computer program comprising a set of instructions configured, when executed on an apparatus, to cause said apparatus to carry out the method of claim 9.

11. A computer-readable medium having recorded thereon a computer program comprising a set of instructions configured, when executed on a device, to cause the device to perform the method of claim 9.

Citation Information

Patent Citations

  • Method and device for classifying objects

    JP2021536634A

  • Depth estimation using a neural network

    US20220335638A1