Apparatus and method for visual positioning

By training a neural network model based on confidence-based perceptual learning, a 3D model rendering training sample set is generated, which solves the problems of insufficient adaptability and coverage of traditional visual positioning methods in dynamic environments, and realizes flexible and reliable pose information prediction and equipment deployment.

CN116324884BActive Publication Date: 2025-11-21HUAWEI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080105976.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-03
Publication Date
2025-11-21
Estimated Expiration
2040-12-03

AI Technical Summary

Technical Problem

Traditional learning-based visual localization methods cannot adapt to dynamic changes in the target environment, have limited coverage, cannot provide reliable confidence estimates, and rely on specific sensor configurations, resulting in inflexible deployment and lack of portability.

Method used

A confidence-based perceptual learning approach is adopted, which trains a neural network model and generates a training sample set using 3D model rendering to enhance the dynamic reflection and diversity of the environment. Combined with domain transformation and data augmentation techniques, the robustness and flexibility of the model are achieved.

Benefits of technology

It improves the predictive reliability and flexibility of visual positioning, can adapt to changes in environmental parameters, provides reliable pose information prediction, and supports the deployment of devices with different sensor configurations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116324884B_ABST
    Figure CN116324884B_ABST
Patent Text Reader

Abstract

The present invention relates to a computing device for supporting a mobile device in performing visual localization in a target environment. The computing device is configured to obtain and modify one or more 3D models of the target environment to generate a set of 3D models. Then, the computing device is configured to determine a training sample set from the set of 3D models. Each training sample comprises an image derived from the set of 3D models and corresponding pose information. Then, the computing device is configured to train a neural network model from the training sample set such that the trained neural network model is configured to predict pose information for an image captured in the target environment and output a confidence value associated with the predicted pose information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an apparatus and method for performing visual positioning in a target environment, and more particularly to a mobile device for performing visual positioning and a computing device for supporting the mobile device in performing visual positioning. Background Technology

[0002] Visual localization (also known as image-based localization or camera relocalization) is widely used in robotics and computer vision. Visual localization refers to the process of determining the pose information of an object or camera within an environment. Pose information includes at least one of position and orientation information. Determining pose information through visual localization can be used in applications such as navigation, advanced driver assistance systems (ADAS), autonomous driving, augmented reality, virtual reality, structure from motion (SfM), and simultaneous localization and mapping (SLAM).

[0003] Using artificial intelligence (AI) methods, neural networks can be applied to perform visual localization. Traditional learning-based visual localization typically involves two phases: an offline phase (also known as the training phase) and an online phase (also known as the prediction phase). In the offline phase, one or more sensors in the environment collect images and associated DoF poses, labeling them as training samples. It's important to note that DoF stands for degrees of freedom. An object or camera in three-dimensional (3D) space can move in six ways: three orientation axes and three rotation axes. Each of these axes represents the DoF. Depending on the hardware capabilities of one or more sensors, one or more of these six axes can be tracked. Specifically, one or more of the three orientation axes can be tracked to determine positional information, and one or more of the three rotation axes can be tracked to determine orientational information. In the offline phase, the training samples are used to train a neural network model to obtain a trained (or learned) neural network model. Then, in the online phase, the trained neural network model is applied to perform visual localization in the target environment or to support visual localization in the target environment. Summary of the Invention

[0004] Traditional learning-based visual localization uses models trained based on fixed representations of the target environment. Specifically, the training samples are static and acquired within one or more specific time periods. Therefore, these training samples cannot reflect the dynamic evolution of the target environment.

[0005] Furthermore, the training samples only provide partial coverage of the target environment. Specifically, the training samples comprise a finite set of observations that do not completely cover the target environment spatially.

[0006] Furthermore, traditional learning-based visual localization cannot provide reliable confidence estimates relative to prediction results. Therefore, the practical application of traditional learning-based visual localization is limited.

[0007] Furthermore, traditional learning-based visual localization requires a dependency between the collected data and the sensors. Specifically, the training samples in the offline phase and the input data in the online phase need to be similar in terms of modality (e.g., radiometrics and geometry) and sensor configuration (e.g., camera configuration). Therefore, the deployment of traditional learning-based visual localization is limited and not portable.

[0008] In view of the above problems, embodiments of the present invention aim to provide a robust learning-based visual localization. Furthermore, embodiments of the present invention aim to provide efficient and flexible deployment for the learning-based visual localization and to provide reliable pose confidence prediction.

[0009] These or other objectives can be achieved by embodiments of the invention described in the appended independent claims. Other implementations of the invention are further defined in the dependent claims.

[0010] The basic concept of this invention is to implement a visual localization method based on confidence-based perceptual learning. This method employs a neural network model trained on visual data, which is rendered from a set of 3D models. Optionally, the set of 3D models can be continuously improved over time by transferring crowdsourced feedback from various platforms performing visual localization using the trained neural network model in the environment.

[0011] A first aspect of the present invention provides a computing device for supporting a mobile device to perform visual positioning in a target environment.

[0012] The computing device is used to: acquire one or more 3D models of the target environment.

[0013] Then, the computing device is used to generate a set of 3D models of the target environment by modifying one or more 3D models of the acquired target environment.

[0014] Then, the computing device is used to: determine a training sample set, wherein each training sample includes an image derived from the set of 3D models and corresponding pose information in the target environment.

[0015] In this way, the dynamic evolution of the target environment can be enhanced and reflected in both the set of 3D models and the training sample set. Specifically, by deriving the image of each training sample, each aspect of the target environment can be reflected.

[0016] Finally, the computing device is used to train a model based on the training sample set, wherein the trained model is used to predict the pose information of the image acquired in the target environment and output a confidence value associated with the predicted pose information.

[0017] In this way, the trained model can be robust to changes in parameters (such as light intensity, shadows, reflections, and texture variations) in the target environment. Furthermore, the diversity of the representation of the target environment can be increased, which can improve prediction reliability.

[0018] It is worth noting that the pose information can be similar to the pose of a sensor or camera (e.g., a sensor or camera of the mobile device) in the target environment corresponding to the image; that is, the content presented by the image of the target environment when captured by the sensor or camera having the pose. The pose information includes at least one of position information and orientation information. The position information can represent the position of the sensor or camera in the target environment, which can be defined relative to three orientation axes (e.g., axes of a Cartesian coordinate system). The orientation information can represent the orientation of the sensor or camera in the target environment, which can be defined relative to three rotation axes (e.g., the yaw axis, pitch axis, and roll axis in the Cartesian coordinate system).

[0019] Optionally, each of the acquired one or more 3D models may be based on one or more sample images and associated sample pose information. The one or more sample images and the associated sample pose information can be acquired by scanning the target environment at one or more locations and from one or more directions using a set of mapping platforms. Each of the set of mapping platforms may be equipped with one or more sensors. The one or more sensors may include at least one of the following: one or more visual cameras, one or more depth cameras, one or more time-of-flight (ToF) cameras, one or more monocular cameras, a light detection and ranging (LiDAR) system, and one or more ranging sensors. The one or more visual cameras may include at least one of a color sensor and a monochrome sensor. The one or more ranging sensors may include at least one of an inertial measurement unit (IMU), a magnetic compass, and a wheeled odometer.

[0020] In one implementation of the first aspect, the one or more 3D models may include multiple different 3D models of the target environment.

[0021] Specifically, the multiple different 3D models can be based on different sample images and different sample pose information located at different positions and in different directions in the target environment.

[0022] This can increase the diversity of the one or more 3D models.

[0023] In another implementation of the first aspect, the image of the target environment may include at least one of the following:

[0024] - Images of the target environment;

[0025] - A depth map of the target environment;

[0026] - A scanned image of the target environment.

[0027] It is worth noting that the image of the target environment can be an image derived from the set of 3D models, or it can be an image obtained in the target environment for predicting the pose information.

[0028] In another implementation of the first aspect, the computing device may also be used for:

[0029] - Model an artificial mobile device, which includes one or more visual sensors;

[0030] - Further determine the training sample set based on the artificial mobile device.

[0031] Optionally, the artificial mobile device may have a general sensor configuration as one of the mapping platforms.

[0032] Alternatively, the sensor configuration of the artificial mobile device can be modified to determine the training sample set.

[0033] It is worth noting that the general sensor arrangement can be understood as a substantially identical sensor arrangement.

[0034] In another implementation of the first aspect, the computing device may also be used for:

[0035] - Calculate a set of trajectories, where each trajectory models the movement of the mobile device in the target environment;

[0036] - Further determine the training sample set based on the set of trajectories.

[0037] Optionally, the set of trajectories can traverse the locations of the target environment.

[0038] This allows for near-complete coverage of the target environment. Furthermore, it enhances the flexibility of the trained model. In other words, the trained model can also be used on mobile devices with a wide variety of sensor configurations.

[0039] In another implementation of the first aspect, in order to modify the one or more 3D models of the target environment, the computing device is configured to perform at least one of the following:

[0040] - Enhance the 3D model of the target environment using one or more artificial lighting sources;

[0041] - Add at least one of static objects and dynamic objects to the 3D model of the target environment.

[0042] Specifically, the static object can refer to an object that is stationary in the target environment. The dynamic object can refer to an object that is moving in the target environment.

[0043] In this way, the trained model can be robust to occlusion and visual changes in the target environment.

[0044] In another implementation of the first aspect, in order to modify the one or more 3D models of the target environment, the computing device may be used to: apply a domain transformation operation to the one or more 3D models of the target environment.

[0045] In another implementation of the first aspect, the computing device may also be used to: apply domain transformation operations or other data augmentation operations to one or more training samples in the training sample set before training the model based on the training sample set.

[0046] Alternatively, the domain transformation operation can be performed using a multi-domain image transformation network architecture.

[0047] Specifically, the multi-domain image transformation network architecture can be based on a generative adversarial network (GAN). The GAN can be used to transform one or more images from a first domain to a second domain.

[0048] In another implementation of the first aspect, the computing device may also be used for:

[0049] - Provide the trained model to the mobile device;

[0050] - Receive feedback information from the mobile device, the feedback information including images acquired in the target environment and corresponding pose information;

[0051] - Based on the feedback information, retrain the trained model.

[0052] This improves the accuracy of the trained model. Furthermore, it enables the sustainable long-term deployment of the mobile device used to perform the visual positioning.

[0053] In another implementation of the first aspect, the computing device may also be used for:

[0054] - Provide the trained model to one or more other mobile devices;

[0055] - Receive feedback information from each of the one or more other mobile devices;

[0056] - Based on all the received feedback information, retrain the trained model.

[0057] In another implementation of the first aspect, the training sample set may include:

[0058] - A first subset of training samples is used to train the model to predict pose information of images acquired in the target environment;

[0059] - A second subset of training samples, which does not intersect with the first subset of training samples, is used to train the model to provide confidence values ​​for pose information prediction.

[0060] In this way, the accuracy of the predicted pose information can be quantified. The mobile device can determine a trust threshold based on the confidence value, wherein the trust threshold is used to determine whether to trust the predicted pose information.

[0061] In another implementation of the first aspect, the computing device may be a mobile device and may be used to collaboratively train and / or retrain the model with one or more other mobile devices.

[0062] Specifically, the computing device may be the mobile device that performs the visual positioning in the target environment. Alternatively, the computing device may be a component of the mobile device.

[0063] This can have the following advantages, for example, the offline and online phases can involve only one device, thus saving hardware resources.

[0064] Furthermore, the computing device can be a node in a distributed computing network. The computing device can be used to collaborate (re)train the model with one or more other mobile devices (computing devices).

[0065] This can have the following advantages, for example, it can further improve the performance (i.e. speed and accuracy) of the (re)trained model used for the visual localization.

[0066] In another implementation of the first aspect, the computing device may be used to request the one or more 3D models of the target environment from one or more mapping platforms to generate a 3D model.

[0067] A second aspect of the invention provides a mobile device for performing visual positioning in a target environment. The mobile device is used for:

[0068] - Receive the trained model from the computing device;

[0069] - Capture images in the target environment;

[0070] - Input the image into the trained model, wherein the trained model is used to predict pose information corresponding to the image and output a confidence value associated with the predicted pose information;

[0071] If the confidence value is lower than a threshold, feedback information is provided to the computing device, the feedback information including the image and the predicted pose information.

[0072] It is worth noting that the pose information may include at least one of position information and orientation information.

[0073] In this way, the accuracy of the trained model can be further improved based on the feedback information.

[0074] In one implementation of the second aspect, the image in the target environment may include at least one of the following:

[0075] - Images of the target environment;

[0076] - A depth map of the target environment;

[0077] - A scanned image of the target environment.

[0078] A third aspect of the present invention provides a method for supporting a mobile device to perform visual positioning in a target environment, the method comprising:

[0079] - Obtain one or more 3D models of the target environment;

[0080] - By modifying one or more 3D models of the acquired target environment, a set of 3D models of the target environment is generated;

[0081] - Determine a training sample set, wherein each training sample includes an image derived from the set of 3D models and the corresponding pose information in the target environment;

[0082] - A model is trained based on the training sample set, wherein the trained model is used to predict the pose information of an image acquired in the target environment and outputs a confidence value associated with the predicted pose information.

[0083] It is worth noting that the pose information may include at least one of position information and orientation information.

[0084] In one implementation of the third aspect, the one or more 3D models may include multiple different 3D models of the target environment.

[0085] In another implementation of the third aspect, the image of the target environment may include at least one of the following:

[0086] - Images of the target environment;

[0087] - A depth map of the target environment;

[0088] - A scanned image of the target environment.

[0089] It is worth noting that the image of the target environment can be an image derived from the set of 3D models, or it can be an image obtained in the target environment for predicting the pose information.

[0090] In another implementation of the third aspect, the method may further include:

[0091] - Model an artificial mobile device, which includes one or more visual sensors;

[0092] - Further determine the training sample set based on the artificial mobile device.

[0093] In another implementation of the third aspect, the method may further include:

[0094] - Calculate a set of trajectories, where each trajectory models the movement of the mobile device in the target environment;

[0095] - Further determine the training sample set based on the set of trajectories.

[0096] In another implementation of the third aspect, the step of modifying the one or more 3D models of the target environment may include at least one of the following:

[0097] - Enhance the 3D model of the target environment using one or more artificial lighting sources;

[0098] - Add at least one of static objects and dynamic objects to the 3D model of the target environment.

[0099] In another implementation of the third aspect, the step of modifying the one or more 3D models of the target environment may include: applying a domain transformation operation to the one or more 3D models of the target environment.

[0100] In another implementation of the third aspect, the method may further include: applying a domain transformation operation or other data augmentation operation to one or more training samples in the training sample set before training the model based on the training sample set.

[0101] In another implementation of the third aspect, the method may further include:

[0102] - Provide the trained model to the mobile device;

[0103] - Receive feedback information from the mobile device, the feedback information including images acquired in the target environment and corresponding pose information;

[0104] - Based on the feedback information, retrain the trained model.

[0105] In another implementation of the third aspect, the method may further include:

[0106] - Provide the trained model to one or more other mobile devices;

[0107] - Receive feedback information from each of the one or more other mobile devices;

[0108] - Based on all the received feedback information, retrain the trained model.

[0109] In another implementation of the third aspect, the training sample set may include:

[0110] - A first subset of training samples is used to train the model to predict pose information of images acquired in the target environment;

[0111] - A second subset of training samples, which does not intersect with the first subset of training samples, is used to train the model to provide confidence values ​​for pose information prediction.

[0112] In another implementation of the third aspect, the computing device may be a mobile device, and the method may further include: training and / or retraining the model in collaboration with one or more other mobile devices.

[0113] In another implementation of the third aspect, the method may include: requesting the one or more 3D models of the target environment from one or more mapping platforms to generate a 3D model.

[0114] A fourth aspect of the present invention provides a method for a mobile device to perform visual positioning in a target environment. The method includes:

[0115] - Receive the trained model from the computing device;

[0116] - Capture images in the target environment;

[0117] - Input the image into the trained model, wherein the trained model is used to predict pose information corresponding to the image and output a confidence value associated with the predicted pose information;

[0118] If the confidence value is lower than a threshold, feedback information is provided to the computing device, the feedback information including the image and the predicted pose information.

[0119] It is worth noting that the pose information may include at least one of position information and orientation information.

[0120] In one implementation of the fourth aspect, the image in the target environment may include at least one of the following:

[0121] - Images of the target environment;

[0122] - A depth map of the target environment;

[0123] - A scanned image of the target environment.

[0124] A fifth aspect of the invention provides program code for executing, when executed on a computer, the method described in accordance with the third aspect or any implementation thereof.

[0125] A sixth aspect of the invention provides program code for executing, when executed on a computer, the method described in accordance with the fourth aspect or any implementation thereof.

[0126] A seventh aspect of the invention provides a non-transitory storage medium for storing executable program code, which, when executed by a processor, causes the method described according to the third aspect or any implementation thereof to be performed.

[0127] The eighth aspect of the invention provides a non-transitory storage medium for storing executable program code, which, when executed by a processor, causes the method described according to the fourth aspect or any implementation thereof to be performed.

[0128] A ninth aspect of the present invention provides a computer product including a memory and a processor, the memory and the processor being used to store and execute program code to perform the method according to the third aspect or any implementation thereof.

[0129] A tenth aspect of the present invention provides a computer product including a memory and a processor for storing and executing program code to perform the method according to the fourth aspect or any implementation thereof.

[0130] Specifically, the memory can be distributed across multiple physical devices. The multiple processors that collaboratively execute the program code can be referred to as processors.

[0131] It should be noted that all devices, elements, units, and components described in this application can be implemented by software or hardware elements or any combination thereof. All steps performed by the various entities described in this application and the functions performed by the various entities are intended to indicate that the corresponding entities are used to perform the corresponding steps and functions. Although the specific functions or steps performed by external entities are not reflected in the description of the specific elements of the entities performing the specific steps or functions in the following detailed description of embodiments, those skilled in the art will understand that these methods and functions can be implemented by corresponding software or hardware elements or any combination thereof. Attached Figure Description

[0132] The following description of specific embodiments, in conjunction with the accompanying drawings, will illustrate the above aspects and their implementations, wherein:

[0133] Figure 1 A schematic diagram of the visual positioning method provided in an embodiment of the present invention is shown;

[0134] Figure 2 The mapping operation provided by an embodiment of the present invention is illustrated;

[0135] Figure 3 The rendering operation provided in the embodiment of the present invention is shown;

[0136] Figure 4 The training operation provided in the embodiment of the present invention is shown;

[0137] Figure 5 The deployment operation provided by an embodiment of the present invention is illustrated;

[0138] Figure 6 The optional improved operations provided by the embodiments of the present invention are shown;

[0139] Figure 7 The present invention illustrates a computing device and a mobile device provided in an embodiment of the present invention;

[0140] Figure 8 Detailed illustrations show the interaction between the mapping platform, computing device, and mobile device provided in embodiments of the present invention;

[0141] Figure 9 The present invention illustrates a computing device provided in an embodiment of the present invention;

[0142] Figure 10 A computing device according to another embodiment of the present invention is shown;

[0143] Figure 11 A mobile device provided in an embodiment of the present invention is shown;

[0144] Figure 12 A mobile device provided according to another embodiment of the present invention is shown;

[0145] Figure 13 This invention illustrates a method for supporting visual positioning of a mobile device in a target environment, provided by an embodiment of the present invention.

[0146] Figure 14 This invention illustrates a method for a mobile device to perform visual positioning in a target environment, provided by an embodiment of the present invention. Detailed Implementation

[0147] exist Figures 1 to 14In the figures, the corresponding components are marked with the same reference numerals and have the same characteristics and functions.

[0148] Figure 1 A schematic diagram of a visual positioning method 100 provided in an embodiment of the present invention is shown. The method 100 includes the following steps (also referred to as operations): mapping (101), rendering (102), training (103), and deployment (104). Optionally, the method 100 may also include improvements (105). Details regarding each operation are shown below.

[0149] Figure 2 The mapping (101) operation provided in an embodiment of the present invention is illustrated.

[0150] In the mapping (101) step, one or more mapping platforms 202 may be deployed in the target environment 201.

[0151] It is worth noting that the target environment 201 can be a mapped area used for performing visual positioning. The target environment 201 can be an indoor environment or an outdoor environment. For example, the target environment can be a single room, a private residence, an office building, a warehouse, a shopping mall, a parking lot, a street, a park, etc. This is merely an example. Figure 2 The target environment 201 is shown as a single room with four walls, viewed from above. However, the target environment 201 should not be limited to... Figure 2 The illustration in the image.

[0152] Specifically, the target environment 201 may refer to a physical space. Optionally, the physical space may include one or more objects. The one or more objects may be stationary or moving.

[0153] In some embodiments, each mapping platform 202 may be used to collect mapping data 203 to generate one or more 3D models of the target environment 201. The mapping data 203 may include at least one of the following: images located at different positions and orientations of the target environment, associated pose information of the images, geometric measurements of the target environment, illumination of the target environment, and timestamps of the mapping data. To collect the mapping data 203, the mapping platform 202 may be equipped with at least one of the following sensors:

[0154] - One or more visual cameras;

[0155] - One or more depth cameras;

[0156] - One or more ToF cameras;

[0157] - One or more monocular cameras;

[0158] -LiDAR system;

[0159] - Photometer;

[0160] - One or more ranging sensors.

[0161] Optionally, the one or more vision cameras may include at least one of a color sensor and a monochrome sensor. The one or more ranging sensors may include at least one of an IMU, a magnetic compass, and a wheel odometer.

[0162] Optionally, the image of the target environment 201 may include at least one of a color image, a monochrome image, a depth image, and a scan image of the target environment 201. Specifically, the scan image of the target environment 201 may be a scan image acquired using the LiDAR system.

[0163] In addition, the mapping platform 202 can send the mapping data 203 to a computing device that supports the mobile device in performing the visual positioning in the target environment 201.

[0164] Based on the mapping data 203, the computing device can be used to acquire one or more 3D models 205. For this purpose, the one or more 3D models 205 can be generated using a 3D reconstruction algorithm 204.

[0165] Optionally, the 3D reconstruction algorithm 204 can be a dense reconstruction algorithm that stitches images together based on the corresponding pose information. Specifically, adjacent pose information can be selected, and the adjacent pose information can be used to project into the common frame of the 3D model 205 of the target environment 201; the corresponding image can be projected accordingly, and the corresponding image can be stitched onto the common frame according to the pose information.

[0166] In some cases, the sensor can provide direct geometric measurement data of the target environment 201, such as the depth camera or the LiDAR system. Therefore, the image can be directly projected onto the common frame based on the direct geometric measurement data.

[0167] In some other cases, the sensor can provide indirect geometric measurement data of the target environment 201, such as the monocular camera. Multiple observations are accumulated and projected into a common frame. Furthermore, external calibration of other sensors can be used.

[0168] In some other cases, the direct geometric measurement data and the indirect geometric measurement data can be combined in the dense reconstruction algorithm to obtain an optimal 3D model.

[0169] Optionally, the 3D model can be optimized using robust mapping algorithms such as SLAM, SfM, bundle adjustment, and factor graph optimization. Furthermore, the 3D model can be post-processed using outlier removal algorithms, mesh reconstruction algorithms, and visual equalization algorithms to obtain a dense, coherent representation of the target environment 201.

[0170] In some embodiments, the computing device may directly request the one or more 3D models of the target environment 201 from one or more mapping platforms. In this case, each of the one or more 3D models can be generated in each mapping platform. Similarly, each mapping platform may use the same 3D reconstruction algorithm as described above.

[0171] Optionally, multiple versions of the target environment can be mapped by running multiple mapping sessions at different times of the day, or by collecting multiple mapping data from different mapping platforms. Therefore, multiple 3D models (205a, 205b, 205c) can be obtained. In this case, each 3D model (205a, 205b, 205c) can be labeled as a different version (206a, 206b, 206c). It should be noted that... Figure 2 Three 3D models are depicted for illustrative purposes only. The number of the one or more 3D models is not limited. Figure 2 The illustration.

[0172] In some embodiments, each of the one or more 3D models 205 may be represented as a dense point cloud or as a water-tight textured mesh.

[0173] Figure 3 The rendering (102) operation provided in an embodiment of the present invention is shown.

[0174] In the rendering (102) step, the computing device can modify one or more 3D models 205 to generate a set of 3D models 305, i.e., a set of modified 3D models. Then, the computing device uses the set of modified 3D models 305 to determine a training sample set 306. Therefore, the diversity of the training samples 306 can be increased.

[0175] Specifically, the computing device can be used to model an artificial (or virtual) mobile device 302. The artificial mobile device 302 can be equipped with one or more virtual sensors that match the one or more sensors of the mapping platform 202. For example, if the one or more sensors of the mapping platform 202 include a visual camera, then the one or more virtual sensors of the artificial mobile device 302 can also include a virtual visual camera, which has the same type of lens as the visual camera of the mapping platform 202 (e.g., the same field of view, the same fisheye, and / or the same distortion).

[0176] Furthermore, the arrangement of the one or more virtual sensors on the artificial mobile device 302 can be the same as the arrangement of the one or more sensors on the mapping platform 202. This arrangement may include the position and orientation of the sensors, and may also include sensor occlusion.

[0177] Optionally, in order to model the artificial mobile device 302, the computing device can be used to acquire sensor information of the mobile device performing the visual positioning in the target environment 201. Then, the computing device can be used to configure the artificial mobile device 302 based on the acquired sensor information. It is worth noting that the sensor information may include the number, type, and arrangement of sensors on the mobile device performing the visual positioning in the target environment 201.

[0178] This can have advantages such as improving the similarity between the training sample set generated by the artificial mobile device 302 and the images acquired by the mobile device during the online phase. Therefore, the accuracy of the visual localization performed by the mobile device can be further improved.

[0179] In some embodiments, to modify the one or more 3D models 205 of the target environment 201, the computing device may be used to enhance the one or more 3D models 205 of the target environment 201 using one or more artificial lighting sources 303. Furthermore, the computing device may be used to add one or more static and / or dynamic objects to the one or more 3D models 205 of the target environment 201. Additionally, the computing device may be used to apply domain transformation operations to the one or more 3D models 205 of the target environment 201.

[0180] It is worth noting that the domain transformation operation performed on the one or more 3D models 205 is to transform the one or more 3D models 205 from one domain to another. A domain is characterized by a set of attributes. Each attribute is a meaningful property (or parameter) of the one or more 3D models. For example, but not exhaustively, the meaningful attribute may be a reference point, geometric center, color space, or texture of the 3D model.

[0181] Alternatively, the domain transformation operation can be performed using a GAN. GANs are well known in the art and will not be described further herein.

[0182] After generating the set of modified 3D models 305, the computing device can be used to calculate a set of trajectories. Each trajectory 307 can model the movement of the mobile device in the target environment 201, so that the artificial mobile device can simulate the movement of the mobile device performing the visual positioning.

[0183] Furthermore, the computing device can be used to determine a training sample set 306. Each training sample may include an image derived from the set of modified 3D models and may include corresponding pose information in the target environment 201.

[0184] To determine the training sample set 306, the computing device can be used to synthesize images based on the trajectories. Specifically, the image of each training sample can be "captured" or rendered by the virtual vision camera of the artificial mobile device 302 at a specific location on the set of trajectories. Furthermore, the corresponding pose information for each training sample can be collected by associated virtual sensors of the artificial mobile device 302.

[0185] Optionally, the number of training samples 306 can be adjusted by setting the sampling frequency of the artificial mobile device 302.

[0186] In this way, the size of the training sample set can be controlled and adjusted according to hardware configuration, computing power, and the accuracy and latency requirements of the visual positioning.

[0187] Alternatively, similar to the domain transformation operation performed on the one or more 3D models 205, the computing device can also be used to perform application domain transformation on one or more training samples 306.

[0188] Alternatively, other data augmentation operations, such as image processing methods, can be applied to one or more of the training samples.

[0189] In addition, the training sample set 306 may include depth images and scan maps of the LiDAR system in order to introduce surrogate modalities for the training step.

[0190] It is worth noting that the term "artificial" in this article can refer to an object simulated by the computing device, or it can refer to an object implemented as a computer program running on the computing device. In other words, the object is virtual. That is, the object does not actually exist.

[0191] In some embodiments, the set of modified 3D models 305 may also be realistic hand-designed 3D models based on a detailed map of the target environment, such as building information models (also known as BIM).

[0192] Figure 4 The training 103 operation provided in an embodiment of the present invention is illustrated.

[0193] In the training (103) step, the computing device trains a model based on the training sample set 306 to obtain a trained model. The trained model is used to predict pose information 403 of an image acquired in the target environment 201 and outputs a confidence value 404 associated with the predicted pose information. The confidence value 404 can be a number used to quantify the reliability of the predicted pose information of the trained model. For example, the confidence value can be in the range of 0 to 1 ([0,1]), where 0 represents the least reliable predicted pose information and 1 represents the most reliable predicted pose information.

[0194] It is important to note that the model described is unrelated to any of the aforementioned 3D models. Specifically, the model can be an artificial neural network (ANN) model, which is essentially a mathematical model. The ANN can also be a neural network. In other words, the model can be a neural network model. Specifically, the model can be a convolutional neural network (CNN) model.

[0195] Specifically, the model prior to training (i.e., the untrained model) may include initial weights (i.e., parameters) that are not optimized for performing the visual localization. The purpose of training 103 is to adjust or fine-tune the weights of the model to perform the visual localization. For example, a supervised learning method can be used to train the model. Furthermore, the trained model may be a set of weights for a machine learning algorithm capable of predicting pose information of an image in the target environment and confidence values ​​associated with that pose information.

[0196] In some embodiments, the training sample set 306 may include a first subset 401 and a second subset 402 of training samples. The first subset 401 of training samples may be used to train the model to predict pose information 403 of one or more images acquired in the target environment 201. The second subset 402 of training samples may be disjoint from the first subset 401 of training samples and may be used to train the model to provide one or more confidence values ​​404 corresponding to the predicted pose information 403.

[0197] Specifically, a portion of the adjusted weights of the trained model can be frozen. For example, Figure 4 The pose training and confidence training performed on the model are illustrated. After the pose training, the adjusted parameters (white neurons in half-trained model 405) can be frozen. The model with frozen weights can then be used for confidence training. Therefore, in another half-trained model 406, the corresponding frozen neurons... Figure 4 The white neurons in the middle are gray. Conversely, the adjusted neurons in the other half of the training model 406 (white neurons in 406) can be frozen in the half-training model 405 (gray neurons in 405). Furthermore, the confidence training can be based on network pose prediction and ground truth prediction. The model can be trained to predict the confidence of the predicted pose, rather than using the difference between the network pose prediction and the ground truth prediction as the training signal, where the confidence is inversely proportional to any uncertainty. Then, after the pose training and the confidence training, the combination of the adjusted weights (e.g., ...) Figure 4 The combination of white neurons in models 405 and 406 shown can form the trained model.

[0198] Optionally, for the confidence training, the model can be trained to output a corresponding confidence value for each DoF pose information. For example, in one training round, when three DoF poses are used together as input to the model (i.e., 3-DoF pose estimation), the model can output three confidence values ​​corresponding to each estimated pose information.

[0199] Optionally, when two or more DoF poses are used together as input, the model can output a covariance matrix associated with the two or more DoF pose information. The covariance matrix can be used to indicate the statistical relationship between any two of the pose information estimated by the model. For example, when six DoF poses are used to train the model in one training round (i.e., 6-DoF pose estimation), the covariance matrix can be a 6×6 matrix. The covariance matrix can be integrated in the online phase to further improve the performance of the visual localization.

[0200] Furthermore, the computing device can be used to send the trained model to the mobile device performing the visual positioning in the target environment 201.

[0201] Other aspects of training the model (i.e., the neural network model) are well known in the art. For example, model initialization, weight adjustment, and model iteration are well known to those skilled in the art. Therefore, detailed information regarding the training of the model will not be elaborated herein.

[0202] Figure 5 The deployment (104) operation provided by an embodiment of the present invention is illustrated.

[0203] In the deployment (104) step, the mobile device 501, used to perform the visual localization in the target environment 201, receives a trained model 503 from the computing device. The trained model 503 can be transferred to the mobile device 501 for actual deployment. The mobile device 501 then captures an image 502 in the target environment 201 and inputs the captured image 502 into the trained model 503. The trained model 503 then predicts pose information 504 corresponding to the image 502 and outputs a confidence value 505 associated with the predicted pose information 504. The predicted pose information can be used to reposition the camera of the mobile device 501 in the target environment.

[0204] Specifically, the mobile device 501 may also include a visual camera for capturing the images.

[0205] Figure 6 An optional improvement (105) operation provided by an embodiment of the present invention is shown.

[0206] In the improvement (105) step, if the confidence value is lower than a threshold, for example, the threshold is 0.7, the mobile device 501 is used to provide feedback information to the computing device, wherein the feedback information includes the captured image 502 and the predicted pose information 504.

[0207] In this way, the deployed mobile device 501 can be used to continuously improve the trained model 503, thereby continuously improving the accuracy of the trained model 503.

[0208] In some embodiments, the mobile device 501 can send the feedback information to the computing device in real time. Alternatively, the feedback information can be stored in the mobile device 501 and sent to the computing device at a later stage.

[0209] After receiving the feedback information, the computing device can be used to construct one or more updated 3D models 205' of the target environment 201. Similar to the rendering (102) step described above, the computing device can be used to modify the one or more updated 3D models 205' to obtain a set of modified 3D models 305'.

[0210] Furthermore, the computing device can also be used to retrain the trained model based on the acquired set of modified 3D models 305' to obtain an updated training model. The updated training model can then be sent to the mobile device 501 for updating. The retraining operation described herein can have the same characteristics and functions as the training (103) step.

[0211] Alternatively, the computing device can be used to train a new model based on the acquired set of modified 3D models 305' to obtain the updated training model.

[0212] It is worth noting that the improvement (105), rendering (102), (re)training (103) and deployment (104) steps can be iterated to continuously improve the trained model.

[0213] In some embodiments, the computing device can be used to receive feedback information from multiple mobile devices. Therefore, the accuracy of the trained model can be further improved based on crowdsourced feedback from the mobile devices.

[0214] Figure 7 The present invention illustrates a computing device 701 and a mobile device 501 provided in an embodiment of the present invention.

[0215] The computing device 701 can be used to receive mapping data from the mapping platform 202. After the mapping (101), rendering (102), and training (103) steps, the computing device 701 can be used to send the trained model to the mobile device 501. The mobile device 501 can be deployed according to the trained model and can provide the feedback information to the computing device 701.

[0216] It should be noted that, Figure 7 The elements shown are only used to illustrate the interaction between the mapping platform 202, the computing device 701, and the mobile device 501. The structural relationship between the computing device 701 and the mobile device 501 is not limited to this. Figure 7 The illustration is shown in the figure. In some embodiments, the computing device 701 may be the mobile device 501. That is, the device used in the offline phase may be referred to as the computing device 701; the same device used in the online phase may be referred to as the mobile device 501.

[0217] Furthermore, the computing device 701 may be a component of the mobile device 501. For example, the computing device 701 may include at least one of the central processing unit, graphics processing unit, tensor processing unit, and neural processing unit of the mobile device 501.

[0218] Figure 8 This illustration shows a detailed diagram of the interaction between the mapping device 202, the computing device 701, and a plurality of mobile devices provided in an embodiment of the present invention.

[0219] exist Figure 8 In this process, the computing device 701 can provide the trained model to multiple mobile devices. The computing device can then receive feedback information from one or more of the mobile devices and retrain the trained model based on the received feedback information.

[0220] Figure 9 A computing device 900 according to an embodiment of the present invention is illustrated. The computing device 900 may include one or more processors 901 and a computer memory 902 that can be read by the one or more processors 901. By way of example only and not limitation, the computer memory 902 may include computer storage media in the form of volatile and / or non-volatile memory, such as read-only memory (ROM) and random access memory (RAM).

[0221] In some embodiments, the computer program may be stored in the computer memory 902. When the computer program is executed by the one or more processors 901, it can cause the computing device 900 to perform, respectively, the following: Figures 2 to 4 The mapping (101), rendering (102), and training (103) steps are shown. Optionally, steps such as... can also be performed. Figure 6 The improved (105) steps are shown.

[0222] Figure 10 A computing device 1000 according to another embodiment of the present invention is shown.

[0223] The computing device 1000 may include:

[0224] -Mapping unit 1001, used to perform, for example Figure 2 The mapping (101) steps are shown;

[0225] -Rendering unit 1002, used to perform tasks such as... Figure 3 The rendering (102) steps are shown below;

[0226] - Training unit 1003, used to perform, for example Figure 4 The training steps (103) shown are as follows.

[0227] Optionally, the computing device 1000 may further include an improvement unit 1005 for performing tasks such as... Figure 6 The improved (105) steps are shown.

[0228] It is worth noting that the computing device 1000 can be a single electronic device capable of performing calculations, or it can be a group of connected electronic devices capable of performing calculations using shared system memory. As is well known in the art, such computing capabilities can be incorporated into many different devices; therefore, the term "computing device" can include PCs, servers, mobile phones, tablets, game consoles, graphics cards, etc.

[0229] Figure 11 A mobile device 1100 according to an embodiment of the present invention is illustrated. The mobile device 1100 may include one or more processors 1101 and a computer memory 1102 that can be read by the one or more processors 1101. By way of example only and not limitation, the computer memory 1102 may include computer storage media in the form of volatile and / or non-volatile memory, such as read-only memory (ROM) and random access memory (RAM).

[0230] In some embodiments, the computer program may be stored in the computer memory 1102. When the computer program is executed by the one or more processors 1101, it may cause the mobile device 1100 to perform actions such as... Figure 5 The deployment steps (104) shown are illustrated.

[0231] Figure 12 A mobile device 1200 according to another embodiment of the present invention is shown.

[0232] The computing device 100 may include:

[0233] - Sensor unit 1201;

[0234] - Deployment unit 1202, used to perform, for example Figure 5 The deployment steps (104) shown are illustrated.

[0235] Specifically, the sensor unit 1201 may include a vision camera for capturing one or more images of the target environment 201. Optionally, the sensor unit 1201 may also include a depth camera or a LiDAR system to improve the accuracy of the predicted pose information.

[0236] Figure 13 This invention illustrates a method 1300 for supporting visual positioning of a mobile device in a target environment, as provided in an embodiment of the present invention.

[0237] The method 1300 is executed by a computing device and includes the following steps:

[0238] - Step 1301: Obtain one or more 3D models of the target environment;

[0239] - Step 1302: Generate a set of 3D models of the target environment by modifying one or more 3D models of the acquired target environment;

[0240] - Step 1303: Determine the training sample set, wherein each training sample includes an image derived from the set of 3D models and the corresponding pose information in the target environment;

[0241] - Step 1304: Train a model based on the training sample set, wherein the trained model is used to predict the pose information of the image acquired in the target environment and output a confidence value associated with the predicted pose information.

[0242] It should be noted that, from the above Figures 2 to 4 From this perspective, the steps of method 1300 can have the same functionality and details. Therefore, the corresponding method implementation will not be described in detail here.

[0243] Figure 14 This invention illustrates a method 1400 for a mobile device to perform visual positioning in a target environment, according to an embodiment of the present invention.

[0244] The method 1400 is performed by a mobile device and includes the following steps:

[0245] - Step 1401: Receive the trained model from the computing device;

[0246] -Step 1402: Capture an image in the target environment;

[0247] - Step 1403: Input the image into the trained model, wherein the trained model is used to predict pose information corresponding to the image and output a confidence value associated with the predicted pose information;

[0248] - Step 1404: If the confidence value is lower than the threshold, feedback information is provided to the computing device, the feedback information including the image and the predicted pose information.

[0249] It should be noted that, from the above Figures 5 to 6 From this perspective, the steps of method 1400 can have the same functionality and details. Therefore, the corresponding method implementation will not be described in detail here.

[0250] One application scenario for this invention is to establish a positioning model for a set of robots that evolve and manage goods in a warehouse.

[0251] In this scenario, a handheld mapping platform consisting of RGB-D cameras is operated by an employee to initially map the warehouse in which the robots will operate. This mapping step can be repeated several times at different times and on different dates before the robots are deployed, allowing for the capture of different representations of the warehouse. The mapping data can be collected and stored on a high-computing-power server for calculating one or more 3D models of the warehouse. The server is also used to modify the one or more 3D models, run rendering algorithms to generate a training sample set, and train the localization model. The trained localization model can be uploaded to the group of robots via a Wi-Fi network. During the group of robots' merchandise management operations, the trained localization module can be used to locate each robot and determine the relative position between the robot and the merchandise. After completing the task, the robot can return to its charging compartment and upload a low-confidence location image to the server. Thus, the server can continuously update the one or more 3D models, render new images (considering changes returned from the group of robots), and update the trained localization model. The updated localization model can then be sent to robots in charging mode.

[0252] The embodiments of this invention are advantageous because they can be directly applied to indoor and outdoor visual positioning to construct realistic 3D models. Furthermore, due to the mapping and rendering steps, a large number of 3D models that might not be physically captured in the mapping step can be generated. Moreover, the set of 3D models can be improved in terms of diversity and can be ported to different mapping platforms with different types of sensors and different sensor arrangements. Furthermore, reliable predictions can be achieved due to the output confidence value associated with the predicted pose information and the improvement steps. Additionally, the trained model can be continuously updated and improved based on the feedback information.

[0253] Furthermore, low-cost sensors can be used, thereby reducing the overall cost of performing the visual localization. Additionally, since the mapping platform does not need to physically traverse every location in the target environment to capture image and depth data at every different rotation, the trained model can be acquired quickly. Moreover, the trained model is robust to partial occlusion and visual variations.

[0254] The invention has been described in conjunction with various embodiments and implementations as examples. However, based on a study of the drawings, the invention, and the independent claims, those skilled in the art will be able to understand and implement other variations in practicing the claimed invention. In the claims and the description, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" does not exclude a plurality. A single element or other unit may fulfill the function of several entities or items listed in the claims. Listing measures in dissimilar dependent claims does not imply that combinations of these measures cannot be used in advantageous implementations.

Claims

1. A computing device (701) for supporting a mobile device (501) to perform visual positioning in a target environment (201), characterized in that, The computing device (701) is used for: Obtain one or more three-dimensional (3D) models (205) of the target environment (201); By modifying one or more 3D models (205) of the acquired target environment (201), a set of 3D models (305) of the target environment (201) is generated. A training sample set (306) is determined, wherein each training sample includes an image derived from the set of 3D models (305) and the corresponding pose information in the target environment (201); The model is trained according to the training sample set (306), wherein the trained model (503) is used to predict the pose information (504) of the image (502) acquired in the target environment (201) and outputs a confidence value (505) associated with the predicted pose information (504).

2. The computing device (701) according to claim 1, characterized in that: The one or more 3D models (205) include multiple different 3D models (205a, 205b, 205c) of the target environment (201).

3. The computing device (701) according to claim 1 or 2, characterized in that, The image of the target environment includes at least one of the following: A picture of the target environment (201); Depth map of the target environment (201); A scan of the target environment (201).

4. The computing device (701) according to claim 1 or 2, characterized in that, Also used for: A model is created of an artificial mobile device (302), which includes one or more visual sensors; The training sample set (306) is further determined based on the artificial mobile device (302).

5. The computing device (701) according to claim 1 or 2, characterized in that, Also used for: A set of trajectories (307) is calculated, wherein each trajectory models the movement of the mobile device (501) in the target environment (201); The training sample set (306) is further determined based on the set of trajectories (307).

6. The computing device (701) according to claim 1 or 2, characterized in that, Modifying one or more 3D models (205) of the target environment (201) includes at least one of the following: The 3D model (205) of the target environment (201) is enhanced by one or more artificial lighting sources (303); Add one or more static and / or dynamic objects to the 3D model of the target environment (201).

7. The computing device (701) according to claim 1 or 2, characterized in that, Modifying one or more 3D models (205) of the target environment (201) includes: Apply a domain transformation operation (304) to one or more 3D models (205) of the target environment (201).

8. The computing device (701) according to claim 1 or 2, characterized in that, Also used for: Before training the model based on the training sample set (306), a domain transformation operation or other data augmentation operation is applied to one or more training samples in the training sample set (306).

9. The computing device (701) according to claim 1 or 2, characterized in that, Also used for: The trained model (503) is provided to the mobile device (501); Feedback information is received from the mobile device (501), the feedback information including an image acquired in the target environment (201) and the corresponding pose information; Based on the feedback information, the trained model is retrained (503).

10. The computing device (701) according to claim 9, characterized in that, Also used for: Provide the trained model to one or more other mobile devices (503); Receive feedback information from each of the one or more other mobile devices; Based on all the received feedback, the trained model is retrained.

11. The computing device (701) according to claim 10, characterized in that, The training sample set (306) includes: A first subset (401) of training samples is used to train the model to predict pose information of images acquired in the target environment (201); A second subset (402) of the training samples, which does not intersect with the first subset (401) of the training samples, is used to train the model to provide confidence values ​​for pose information prediction.

12. The computing device (701) according to any one of claims 1, 2, 10 to 11, characterized in that: The computing device (701) is a mobile device and is used to train and / or retrain the model in collaboration with one or more other mobile devices.

13. The computing device (701) according to any one of claims 1, 2, 10 to 11, characterized in that, Used for: Request one or more 3D models (205) of the target environment (201) from one or more mapping platforms to generate 3D models.

14. A mobile device (501) for performing visual positioning in a target environment (201), characterized in that, The mobile device (501) is used for: Receive the trained model (503) from the computing device (701); Capture an image (502) in the target environment (201); The image (502) is input into the trained model (503), wherein the trained model (503) is used to predict pose information (504) corresponding to the image (502) and output a confidence value (505) associated with the predicted pose information (504). If the confidence value (505) is lower than the threshold, feedback information is provided to the computing device (701), the feedback information including the image (502) and the predicted pose information (504).

15. A method (1300) for supporting a mobile device (501) to perform visual positioning in a target environment (201), characterized in that, The method (1300) includes: Obtain (1301) one or more three-dimensional (3D) models (205) of the target environment (201). By modifying one or more 3D models (205) of the acquired target environment (201), a set of 3D models (305) of the target environment (201) is generated (1302). Determine (1303) a training sample set (306), wherein each training sample includes an image derived from the set of 3D models (305) and the corresponding pose information in the target environment (201); A model (1304) is trained based on the training sample set (306), wherein the trained model (503) is used to predict the pose information (504) of the image (502) acquired in the target environment (201) and outputs a confidence value (505) associated with the predicted pose information (504).

16. A method (1400) for a mobile device to perform visual positioning in a target environment (201), characterized in that, The method (1400) includes: Receive (1401) the trained model (503) from the computing device (701); Capture (1402) an image (502) in the target environment (201); The image (502) is input (1403) into the trained model (503), wherein the trained model (503) is used to predict pose information (504) corresponding to the image (502) and output a confidence value (505) associated with the predicted pose information (504). If the confidence value (505) is lower than the threshold, feedback information (1404) is provided to the computing device (701), the feedback information including the image (502) and the predicted pose information (504).

17. A computer program product, characterized in that, include: Program code, when executed on a computer, to perform the method according to claim 15 or 16.

Citation Information

Patent Citations

  • Facility surveillance systems and methods

    WO2020088739A1