Convergence rate optimization method of SLAM model

By building the SLAM model, using pre-trained data and volume rendering technology to optimize scene coding and pose updates, the problem of slow convergence speed of existing dense SLAM methods is solved, and more efficient SLAM model convergence and camera tracking accuracy are achieved.

CN120014458APending Publication Date: 2025-05-16ZHEJIANG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510090467.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing dense SLAM method using neural radiation fields relies on static scene characterization strategies, so that only scene information optimization is performed on individual images each time the input image is processed, which reduces the convergence speed of the SLAM method.

Method used

By constructing the SLAM model, using pre-trained data to generate a 3D point cloud collection, determining the scene encoding of the image, and performing volume rendering to generate prediction results, and building a loss function optimization model. In the map construction process, input images are collected in image order, volume rendering and synchronous pose updates, and model parameters are updated.

Benefits of technology

The convergence speed of the SLAM method is improved, the accuracy of camera tracking is improved, the computational burden is reduced, and continuous pose optimization is achieved for each sampled image without adding additional calculations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014458A_ABST
    Figure CN120014458A_ABST
Patent Text Reader

Abstract

The invention provides a convergence rate optimization method for an SLAM model, and the method employs a multi-image beam optimization adjustment strategy, reduces the calculation burden of all selected sampling images in a mode of replacing global pixels with local pixels in a map construction process, and achieves the convergence rate optimization of all selected sampling images through employing a trained loss function. According to the method, all the selected sampling images are subjected to batch synchronous pose updating, so that pose optimization can be continuously carried out on each sampling image in the mapping process under the condition of not increasing additional calculation burden, the camera tracking accuracy can be improved, and the convergence speed of the SLAM method is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the technical field of data transmission, and in particular relates to a method for optimizing the convergence speed of a SLAM model. Background Art

[0002] Simultaneous Localization and Mapping (SLAM) technology has significant application value in many fields such as autonomous driving, robot navigation, and virtual / augmented reality because it can accurately reconstruct the camera's motion trajectory and the three-dimensional map of the scene from an image sequence. Traditional dense SLAM methods fuse depth observations into explicit scene representations (such as voxels, surface units, etc.) and perform camera tracking while creating dense scene maps. However, due to the limitation of fixed resolution, traditional dense SLAM methods are very likely to lose reconstruction details at small scales, so they have been replaced by dense SLAM using neural radiance fields.

[0003] However, the neural radiation field technology used by most current dense SLAM methods generally relies on static scene representation strategies. Specifically, these methods tend to assume that the optimizable scene features are random variables sampled from a normal distribution. Each time such methods process an input image, they only optimize the scene information of each input image separately, thereby reducing the convergence speed of the corresponding SLAM method. Summary of the invention

[0004] In order to solve the problems existing in the existing dense SLAM methods using neural radiation fields, one or more embodiments of this specification describe a method for optimizing the convergence speed of a SLAM model.

[0005] According to a first aspect, a method for optimizing the convergence speed of a SLAM model is provided, comprising:

[0006] Constructing a SLAM model, obtaining pre-training data, wherein the pre-training data includes a pre-training image and a pre-training depth image corresponding to the pre-training image, generating a 3D point cloud set based on the pre-training data, and determining a first scene code of each first sampling point of the pre-training image in a global voxel feature grid through the 3D point cloud set, wherein the first scene code includes a scene geometry code and a scene appearance code;

[0007] Performing volume rendering on the scene encoding in the SLAM model to generate a prediction result, wherein the prediction result includes a predicted image and a predicted depth image corresponding to the pre-trained image;

[0008] Constructing a first loss function of the SLAM model, optimizing the first loss function based on a comparison difference between a prediction result and pre-training data, and obtaining an optimized SLAM model;

[0009] In the map building process of the optimized SLAM model, according to the image sequence of the first input images captured by the depth camera, at least a preset number of the first input images whose shooting time is after a preset time are continuously collected, and at least a preset number of the first input images whose shooting time is before a preset time are randomly collected, and volume rendering is performed on each of the collected first input images. According to the optimized first loss function, a synchronous pose update is performed on each of the first input images after volume rendering, and the map building process parameters of the optimized SLAM model are updated based on the synchronous pose update result.

[0010] Preferably, the generating of a 3D point cloud set based on the pre-training data, and determining a first scene code of each first sampling point of the pre-training image in a global voxel feature grid through the 3D point cloud set, wherein the first scene code includes a scene geometry code and a scene appearance code, including:

[0011] For any of the pre-trained depth images, determine image feature information of the pre-trained image corresponding to the pre-trained depth image, determine each first sampling point based on the image feature information, and draw each first sampling point of the pre-trained image into a world coordinate system corresponding to the feature three planes to obtain a 3D point cloud set;

[0012] Based on the prior encoder, the 3D point cloud set is geometrically encoded to obtain a global voxel feature grid;

[0013] Decoding the global voxel feature grid based on the decoder to obtain a first scene code of the global voxel feature grid;

[0014] Ray sampling is performed on the pre-trained image, and first scene codes corresponding to each first sampling point are obtained from the global voxel feature grid according to the ray sampling result.

[0015] Preferably, the geometric a priori encoding of the 3D point cloud set by the a priori encoder to obtain a global voxel feature grid includes:

[0016] Acquire prior data corresponding to the pre-trained image, the prior data comprising first prior data corresponding to scene appearance and second prior data corresponding to scene geometry;

[0017] Training the a priori encoder using the a priori data, and optimizing parameters of the a priori encoder according to the training result of the a priori encoder;

[0018] Performing geometric a priori encoding on local voxels by using the optimized a priori encoder to obtain a local voxel set, wherein the local voxels are voxels corresponding to all 3D point clouds in the 3D point cloud set;

[0019] The local voxel set is fused into the global voxel set, and a global voxel feature grid is constructed according to the fused global voxel set.

[0020] Preferably, the decoder includes an appearance decoder and a geometry decoder, the appearance decoder is used to decode the scene appearance code in the global voxel feature grid, and obtain the RGB values ​​corresponding to each of the first sampling points of the pre-trained image; the geometry decoder is used to decode the scene geometry code in the global voxel feature grid, and obtain the TSDF values ​​corresponding to each of the first sampling points of the pre-trained image.

[0021] Preferably, the performing ray sampling on the pre-trained image includes:

[0022] Obtaining the depth camera pose and sampling step size corresponding to the pre-trained image;

[0023] Determine at least two rays according to the depth camera pose and the sampling step, sample the pre-trained image along the ray direction, and map the second sampling point of each ray to the global voxel feature grid;

[0024] A first scene code corresponding to the second sampling point is obtained to generate a ray sampling result.

[0025] Preferably, the method further comprises: in the map building process of the optimized SLAM model, gradient updating the depth camera pose of the camera tracking process by using the optimized first loss function.

[0026] Preferably, the first loss function includes a second loss function, a third loss function and a fourth loss function; the second loss function is used to generate a rendered image that approaches the pre-trained image based on the prediction result and the pre-training data; the third loss function is used to predict the TSDF value of the first sampling point corresponding to the pre-training image and located in the truncation area near the ray surface, so that the predicted result of the TSDF value of the first sampling point approaches the TSDF value of the pre-training image; the fourth loss function is used to predict the TSDF value of the first sampling point corresponding to the pre-training image and located between the center of the depth camera and the truncation area, so that the TSDF value of the first sampling point approaches a predetermined fixed value of the truncation distance.

[0027] Preferably, the second loss function includes:

[0028] a fifth loss function, the fifth loss function being used to calculate color loss of the pre-trained image and the predicted image generated by volume rendering;

[0029] A sixth loss function is used to calculate the depth loss of the pre-training image and the predicted depth image generated by volume rendering.

[0030] Preferably, the method further comprises: in the camera tracking process of the optimized SLAM model, randomly sampling local pixels of the first input image, and replacing global pixels of the first input image based on the sampled local pixels.

[0031] Preferably, after performing synchronous posture updating on each of the first input images after volume rendering according to the optimized first loss function, the method further includes:

[0032] A second scene code of a two-dimensional depth image corresponding to a second input image is obtained, and decoder parameters are updated based on the second scene code, wherein the second input image is obtained by performing synchronous posture update on the first input image after volume rendering by the first loss function.

[0033] Beneficial effect: The method proposed in this application adopts a multi-image bundle optimization adjustment strategy. In the process of map construction, local pixels are used instead of global pixels to reduce the computational burden of all selected sampled images, and then the trained loss function is used to batch-update the synchronous pose of all selected sampled images. In this way, the pose of each sampled image in the mapping process can be continuously optimized without adding additional computational burden, thereby improving the accuracy of camera tracking and the convergence speed of the SLAM method. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0035] Figure 1 A flow chart of a method for optimizing the convergence speed of a SLAM model provided in an embodiment of the present application;

[0036] Figure 2 A schematic diagram of the principle of a method for optimizing the convergence speed of a SLAM model provided in an embodiment of the present application. DETAILED DESCRIPTION

[0037] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.

[0038] In the following introduction, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance. The following introduction provides multiple embodiments of the present application, and different embodiments can be replaced or combined, so the present application can also be considered to include all possible combinations of the same and / or different embodiments recorded. Therefore, if one embodiment includes features A, B, and C, and another embodiment includes features B and D, then the present application should also be considered to include embodiments containing one or more of all other possible combinations of A, B, C, and D, although the embodiment may not be clearly recorded in the following text.

[0039] The following description provides examples and does not limit the scope, applicability or examples set forth in the claims. Changes may be made to the functions and arrangements of the elements described without departing from the scope of the present application. Various processes or components may be appropriately omitted, substituted or added to each example. For example, the described method may be performed in an order different from the order described, and various steps may be added, omitted or combined. In addition, features described in some examples may be combined in other examples.

[0040] See also Figure 1 , Figure 1 1 is a flow chart of a method for optimizing the convergence speed of a SLAM model provided in an embodiment of the present application. In an embodiment of the present specification, the method includes:

[0041] S1. Build a SLAM model and obtain pre-training data, wherein the pre-training data includes a pre-training image and a pre-training depth image corresponding to the pre-training image, generate a 3D point cloud set based on the pre-training data, and determine the first scene code of each first sampling point of the pre-training image in the global voxel feature grid through the 3D point cloud set, wherein the first scene code includes a scene geometry code and a scene appearance code.

[0042] The execution subject of the present application may be a processor using the SLAM method, and the processor may be an ARM processor based on a reduced instruction set computing (RISC) architecture, an embedded device processor, and a mobile computer.

[0043] In the embodiments of this specification, the image taken by the depth camera is processed in real time by the processor using the method of this embodiment, and the first loss function corresponding to the method can also be trained. The training process can be directly trained using the processor CPU, or it can be trained by an external GPU device. The processor of the method of this embodiment uses the first loss function corresponding to the method of this embodiment to update the depth camera pose of the camera tracking process, and can simultaneously update the pose of all the first input images used in the map construction process, thereby updating the parameters of the map construction. The objects of this method include depth cameras. The depth camera is a camera that can obtain the distance information of an object. It can calculate the three-dimensional coordinates of the object by combining 2D image information and depth information.

[0044] The purpose of obtaining the pre-trained image and the pre-trained depth image corresponding to the pre-trained image is to train the first loss function used in this method. The training process requires a large number of pre-trained images to improve the accuracy of the first loss function. The purpose of generating a 3D point cloud set based on the pre-trained image and the pre-trained depth image is to associate the features of the three-dimensional pre-trained depth image with the corresponding two-dimensional pre-trained image in a three-dimensional space to obtain the scene geometry and scene appearance features corresponding to the two-dimensional pre-trained image. The global voxel feature grid is used to concentrate the voxel information corresponding to each sampling point of the pre-trained image, and can quickly obtain the required voxel information through the mapping relationship between the two-dimensional pre-trained image and the three-dimensional pre-trained depth image.

[0045] In one embodiment, the generating of a 3D point cloud set based on the pre-training data, determining a first scene code of each first sampling point of the pre-training image in a global voxel feature grid through the 3D point cloud set, wherein the first scene code includes a scene geometry code and a scene appearance code, comprises:

[0046] For any of the pre-trained depth images, determine image feature information of the pre-trained image corresponding to the pre-trained depth image, determine each first sampling point based on the image feature information, and draw each first sampling point of the pre-trained image into a world coordinate system corresponding to the feature three planes to obtain a 3D point cloud set;

[0047] Based on the prior encoder, the 3D point cloud set is geometrically encoded to obtain a global voxel feature grid;

[0048] Decoding the global voxel feature grid based on the decoder to obtain a first scene code of the global voxel feature grid;

[0049] Ray sampling is performed on the pre-trained image, and first scene codes corresponding to each first sampling point are obtained from the global voxel feature grid according to the ray sampling result.

[0050] In the embodiment of the present specification, to obtain the scene geometry coding and scene appearance coding information corresponding to each first sampling point of the pre-trained image, it is necessary to obtain the 3D point cloud set of each sampling point through the pre-trained image and the corresponding pre-trained depth image. In the process of generating the 3D point cloud set of the pre-trained image in the world coordinate system corresponding to the feature three planes, it is necessary to initialize a feature three plane for representing the scene appearance, and then reversely project the depth image into the world coordinate system according to the estimated camera pose, so as to obtain the 3D point cloud set.

[0051] In the actual process of obtaining the first scene coding information, some other information of the image itself, such as size and resolution, occupies a lot of computing resources. Directly using the loss function or model to train the entire image will consume a lot of computing power. Therefore, it is necessary to quickly extract the first scene coding information corresponding to the image through the operation of encoding first and then decoding.

[0052] The purpose of ray sampling the pre-trained image is to associate the pre-trained image with the corresponding pre-trained depth image, that is, to associate the voxel features of the first sampling point of the two-dimensional pre-trained image with the three-dimensional feature points of the global voxel feature grid, and map the corresponding first sampling point of the pre-trained image to the global voxel feature grid and the feature three-plane, so as to facilitate subsequent decoding and volume rendering, as well as the training of the first loss function. Here, the feature three-plane is three projection planes, which are used to store the global voxel feature grid and are the basis for establishing the world coordinate system.

[0053] In one implementation, the geometric a priori encoding of the 3D point cloud set by the a priori encoder to obtain a global voxel feature grid includes:

[0054] Acquire prior data corresponding to the pre-trained image, the prior data comprising first prior data corresponding to scene appearance and second prior data corresponding to scene geometry;

[0055] Training the a priori encoder using the a priori data, and optimizing parameters of the a priori encoder according to the training result of the a priori encoder;

[0056] Performing geometric a priori encoding on local voxels by using the optimized a priori encoder to obtain a local voxel set, wherein the local voxels are voxels corresponding to all 3D point clouds in the 3D point cloud set;

[0057] The local voxel set is fused into the global voxel set, and a global voxel feature grid is constructed according to the fused global voxel set.

[0058] In the embodiments of this specification, the a priori encoder is pre-trained with a priori data, which may include scene appearance information and scene geometry information of part of the pre-trained image captured by the depth camera, and may also include a priori data artificially produced based on the scene captured by the depth camera. The a priori encoder is pre-trained with a priori data so that the processor can obtain the basic structure of the scene before image optimization and have a preliminary understanding of the scene structure in the image. This process ensures that the method can be executed in real time when performing precise camera tracking and high-precision surface reconstruction.

[0059] The voxel set refers to the set of feature vectors with fixed length obtained by encoding the geometric prior at each local pixel corner vertex in the 3D point cloud set. The process of fusing the voxel set into the global voxel set is as described in the formula:

[0060]

[0061] Where V represents the geometric encoding feature vector, V g represents the global voxel set, V m represents the voxel set, W g represents the weight corresponding to the global voxel set, W m Represents the weight corresponding to the global voxel set. By combining the global voxel set with the feature three planes, the global voxel feature grid is obtained.

[0062] In one embodiment, the decoder includes an appearance decoder and a geometry decoder, wherein the appearance decoder is used to decode the scene appearance code in the global voxel feature grid and obtain the RGB values ​​corresponding to each of the first sampling points of the pre-trained image; the geometry decoder is used to decode the scene geometry code in the global voxel feature grid and obtain the TSDF values ​​corresponding to each of the first sampling points of the pre-trained image.

[0063] In an embodiment of the present specification, in an embodiment of the present specification, the scene geometry code in the global voxel feature grid is used to represent the position and geometry of the corresponding pre-trained image, and the scene appearance code in the global voxel feature grid is used to represent the color features of the corresponding pre-trained image. Both the scene geometry code and the scene appearance code in the global voxel feature grid need to be decoded before they can be used for the training of the first loss function. In the process of decoding the global voxel features by the decoder, the scene geometry code is decoded into a truncated signed distance field (TSDF) value. TSDF is a data structure used to represent the surface of an object in three-dimensional space. It divides the space into a regular voxel grid and stores a signed distance value for each voxel. This distance value represents the closest distance from the center of the voxel to the surface of the object. The voxels inside the object surface have negative values ​​and the outside has positive values. The scene appearance code is decoded into an RGB value, which represents the numerical value of the three color channels of the corresponding pre-trained image.

[0064] To obtain the TSDF value corresponding to each first sampling point of the pre-trained image, it is necessary to first obtain the 8 corner vertex features of the voxel corresponding to each sampling point, and then calculate the geometric feature vector of the point by trilinear interpolation. The calculation process is described in the formula:

[0065]

[0066] Among them, S P represents the TSDF value of the first sampling point P, TriLerp(,) represents the trilinear interpolation function, V g represents the global voxel set, Represents the decoding process of the geometry decoder.

[0067] Obtain the RGB values ​​corresponding to each first sampling point of the pre-trained image, first project each first sampling point onto three different color feature planes, and then retrieve the corresponding feature f by bilinear interpolation xy (p), f xz (p) and f yz (p), the final decoder Decode the appearance feature vector into RGB values. The calculation process is as described in the formula:

[0068]

[0069] Among them, f xy (p), f xz (p) and f yz (p) represents the color features of different feature planes, represents the decoding process of the appearance decoder, and f(p) represents the appearance feature vector of the first sampling point p.

[0070] In one implementation, the performing ray sampling on the pre-training image includes:

[0071] Obtaining the depth camera pose and sampling step size corresponding to the pre-trained image;

[0072] Determine at least two rays according to the depth camera pose and the sampling step, sample the pre-trained image along the ray direction, and map the second sampling point of each ray to the global voxel feature grid;

[0073] A first scene code corresponding to the second sampling point is obtained to generate a ray sampling result.

[0074] In the embodiment of this specification, the depth camera posture includes the shooting point position of the depth camera in the global voxel feature grid and the position of the lens shooting center axis. A number of rays can be obtained through the depth camera posture and sampling step size, and the pre-trained image is sampled along the rays, and the second sampling point is mapped to the global voxel feature grid and the feature three-plane. First, M pixels are sampled on the input image, and the camera ray direction vector d(x) corresponding to each pixel of the voxel is calculated through the back projection model and the corresponding camera posture. d ,y d ,z d ) and the camera center o(x o ,y o ,z o ), and get the ray set r m =o m +z p ×d m ,m∈{1,…,M}, where z p is the depth of the second sample point along the ray.

[0075] The camera ray direction vector d(x) corresponding to each pixel is calculated by back-projecting the model and the corresponding camera pose. d ,y d ,z d ) and the camera center o(x o ,y o ,z o ) process is shown in the formula:

[0076]

[0077] (x o ,y o ,z o ) = t k

[0078] Among them, u and v are the horizontal and vertical coordinates of the pixel point in the image coordinate system, c x 、c yare the coordinates of the image center on the x-axis and y-axis in the image coordinate system, respectively, x 、f y are the components of the camera focal length in the x-axis and y-axis directions, R k represents the rotation matrix, t k Represents a three-dimensional translation vector.

[0079] S2. Perform volume rendering on the scene encoding in the SLAM model to generate a prediction result, wherein the prediction result includes a predicted image and a predicted depth image corresponding to the pre-trained image.

[0080] In the embodiments of this specification, volume rendering is a technology that uses 3D data fields to generate two-dimensional images. Unlike traditional rasterization rendering that uses triangular facets for modeling, volume rendering uses voxels for modeling and samples and renders data along the light path through ray casting technology. Through volume rendering, the global voxel feature grid decoded by the decoder can be converted into a two-dimensional predicted image and predicted depth image, so that the first loss function can be trained in conjunction with the pre-trained image and the pre-trained depth image.

[0081] S3, constructing a first loss function of the SLAM model, optimizing the first loss function based on the comparison difference between the prediction result and the pre-training data, and obtaining an optimized SLAM model.

[0082] In the embodiments of this specification, Figure 1 As shown, the first loss function can be trained with a large number of data sets, such as using the real-world data set ScanNet for pre-training, and then re-training with the prediction results corresponding to each pre-training image and the pre-training data. After being trained with a certain number of training images, the first loss function can be updated with gradients according to the pose of the image frame taken in real time by the depth camera, thereby updating and optimizing the depth camera pose, and can also update the pose of each image of the pixel database used for map construction or the optimized frame of non-pixel data, thereby optimizing the parameters of the constructed map.

[0083] In one embodiment, the first loss function includes a second loss function, a third loss function and a fourth loss function; the second loss function is used to generate a rendered image that approaches the pre-trained image based on the predicted result and the actual result; the third loss function is used to predict the TSDF value of the first sampling point corresponding to the pre-trained image and located in the truncation area near the ray surface, so that the predicted result of the TSDF value of the first sampling point approaches the TSDF value of the pre-trained image; the fourth loss function is used to predict the TSDF value of the first sampling point corresponding to the pre-trained image and located between the center of the depth camera and the truncation area, so that the TSDF value of the first sampling point approaches a predetermined fixed value of the truncation distance.

[0084] In the embodiment of this specification, since the second sampling point on the ray is on the global voxel feature grid, the second sampling point belongs to a three-dimensional sampling point. Therefore, there is a first sampling point corresponding to the second sampling point on the ray on the pre-trained image. For ease of explanation, the sampling points used for calculating the first loss function are all the first sampling points. When the second sampling point on the ray is needed, the first sampling point corresponding to the second sampling point is also used. As an example, the first objective function can be a composite function composed of multiple functions such as an L1 loss function and an L2 loss function.

[0085] The first loss function is:

[0086] L=λ rgb L rgb +λ depth L depth +λ fs L fs +λ tsdf L tsdf

[0087] Where L rgb , L depth are all sub-functions of the second loss function, λ rgb With λ depth Respectively represent L rgb With L depth The preset parameters of fs Represents the preset parameters of the fourth loss function, L fs represents the fourth loss function; λ tsdf Represents the preset parameters of the third loss function, L tsdf Represents the third loss function. rgb With λ depth , fs With λ tsdf The specific value of is modified according to the actual training requirements. For example, when the most important goal is to improve the accuracy of the depth camera pose in the camera tracking process, L is set during the training process. rgb With Ldepth Parameters are frozen and no parameter update is performed.

[0088] The third loss function is:

[0089]

[0090] Among them, z p represents the depth of point p along the ray, is the true depth value of the corresponding pixel, represents the first sampling point in the truncation region near the ray surface, tr is the defined maximum TSDF value, S p is the TSDF value corresponding to the first sampling point predicted, and m represents the sequence number of the first sampling point.

[0091] In the embodiment of this specification, as an example, the third objective function may be a combination of one or more loss functions such as a mean square error function and a mean absolute error loss function. The fourth objective function may be a combination of one or more loss functions such as a mean square error function and a mean absolute error loss function.

[0092] The third loss function is used to predict the TSDF value of the first sampling point in the truncated area near the ray surface corresponding to the pre-training image, and the second sampling point in the truncated area near the ray surface also corresponds to the first sampling point of the pre-training image. The third loss function uses the geometric loss of the first sampling point corresponding to the pre-training image, and uses the real depth observation value to approximate the TSDF true value for supervision for the first sampling point in the truncated area near the ray surface. In this way, the third loss function can more accurately learn the scene surface shape corresponding to the global voxel feature grid.

[0093] The fourth loss function is:

[0094]

[0095] in, represents the first sampling point between the center of the depth camera and the truncated area, tr is the predetermined fixed value of the truncation distance, S p is the TSDF value corresponding to the first predicted sampling point.

[0096] The fourth loss function is used to predict the TSDF values ​​of the first sampling points located between the center of the depth camera and the truncated area corresponding to the pre-trained image, and force the TSDF value prediction of these first sampling points to be close to the predefined maximum TSDF value. In this way, the influence of unimportant feature points far from the scene surface corresponding to the pre-trained image is reduced, similar to reducing the influence of noise points, while reducing the amount of calculation of the first loss function, optimizing the efficiency of the method of this embodiment.

[0097] In one embodiment, the second loss function includes:

[0098] a fifth loss function, the fifth loss function being used to calculate color loss of the pre-trained image and the predicted image generated by volume rendering;

[0099] A sixth loss function, the fifth loss function is used to calculate the depth loss of the pre-training image and the predicted depth image generated by volume rendering.

[0100] In the embodiment of this specification, the fifth objective function can be a combination of one or more loss functions such as a mean square function and a mean absolute error loss function. The fifth loss function is:

[0101]

[0102] in, represents the actual observed color of the corresponding pixel, C m is the predicted color of the corresponding pixel;

[0103] The sixth loss function, the sixth objective function can be a combination of one or more loss functions such as a mean square function, a mean absolute error loss function, etc.

[0104] The sixth loss function is used to calculate the depth loss of the pre-trained image and the predicted depth image generated by volume rendering, and the fourth loss function is:

[0105]

[0106] in, is the true depth value of the corresponding pixel; D m is the predicted depth value of the corresponding pixel.

[0107] The fifth loss function is used to calculate the color loss of the pre-trained image and the predicted image generated by volume rendering. The difference between the color features of the predicted image and the pre-trained image can be obtained by the mean absolute error method, and then the fifth loss function can be updated with parameters to generate a rendered image with a color closer to the pre-trained image. The sixth loss function is used to calculate the depth loss of the predicted depth image generated by the pre-trained image and volume rendering. The depth loss is directly related to the scene geometry features of the predicted depth image. The difference between the appearance geometry features of the predicted depth image and the pre-trained depth image can be obtained by the mean absolute error method, and then the fourth loss function can be updated with parameters to cooperate with the second loss function to generate a rendered image with scene geometry features closer to the pre-trained image.

[0108] S4. In the map building process of the optimized SLAM model, according to the image sequence of the first input images captured by the depth camera, continuously collect at least a preset number of the first input images whose shooting time is after a preset time, randomly collect at least a preset number of the first input images whose shooting time is before a preset time, perform volume rendering on each of the collected first input images, and perform synchronous pose update on each of the first input images after volume rendering according to the optimized first loss function, and update the map building process parameters of the optimized SLAM model based on the synchronous pose update result.

[0109] In the embodiment of this specification, since the first input image is an image directly taken by the depth camera, the first input image includes a two-dimensional plane image and a three-dimensional depth image, rather than a single two-dimensional image. The preset time can be any time endpoint in the working process of the depth camera, and the preset number can be determined according to the actual situation, but the number cannot be too low. For example, if only 10 frames of the first input image are collected, the effect of map construction will be poor.

[0110] According to the sequence of the first input images captured by the depth camera, continuously collect at least a preset number of the first input images whose shooting time is after the preset time, and randomly collect at least a preset number of the first input images whose shooting time is before the preset time. This operation mode is to ensure the accuracy of map construction while improving the execution efficiency. For example, in each map construction process, a total of 200 frames of images captured by the camera are selected for joint optimization, including the 20 frames with the most recent camera shooting time, and 90 frames randomly selected from the frames with a common viewing area of ​​more than 10% with the current frame. And an additional 90 frames randomly selected from all historical frames to prevent a large amount of information about the frames with earlier shooting times from being forgotten, thereby ensuring the accuracy of map construction.

[0111] like Figure 2 As shown, the pose of each first input image after volume rendering is updated synchronously according to the first loss function, that is, the first input image used to be collected is bundle optimized. Bundle optimization is to optimize a group of images or data in a unified manner. Through this optimization strategy, the pose of each input frame can be continuously refined during all map construction processes without adding additional computational burden.

[0112] In the camera tracking process, the gradient of the depth camera pose of the camera tracking process can also be updated by optimizing the first loss function. Figure 2As shown, the first loss function is used to calculate the error between the estimated pose and the true pose during the camera tracking process, and the pose is gradient updated through this error to optimize the camera position. When the first loss function is used to optimize the depth camera pose, the geometric feature encoding, appearance feature encoding and decoder parameters of the pre-trained image and the pre-trained depth image are frozen, and the gradient of the loss function relative to the camera pose parameters is calculated by back propagation, and the loss function is minimized only by updating the camera pose. In this way, the amount of calculation when using the first loss function in the camera tracking process can be reduced, and the efficiency of the method can be improved.

[0113] In one embodiment, the method further includes: in a camera tracking process of the optimized SLAM model, randomly sampling local pixels of the first input image, and replacing global pixels of the first input image based on the sampled local pixels.

[0114] In the embodiments of this specification, partial pixels are used to replace global pixels of an image during map construction. There is no need to select key frames. Instead, a small number of pixels are sampled from each input frame to replace the frame. This allows the system to optimize the pose of each frame without increasing the computational load.

[0115] In one embodiment, after synchronously updating the pose of each of the first input images after volume rendering according to the optimized first loss function, the method further includes:

[0116] A second scene code of a two-dimensional depth image corresponding to a second input image is obtained, and decoder parameters are updated based on the second scene code, wherein the second input image is obtained by performing synchronous posture update on the first input image after volume rendering by the first loss function.

[0117] In the embodiments of this specification, Figure 2 As shown, in the map construction process of the optimized SLAM model, the first loss function will render the volume of the first input image to obtain the second input image, and synchronize the pose update of the second input image. The second input image can be understood as the corresponding predicted image generated by the first loss function based on the first input image. Since it is an image directly taken by the depth camera, the first input image includes a two-dimensional plane image and a depth image, rather than a single two-dimensional image. Therefore, the first loss function will encode the second scene of each second input image according to the prediction result. The second scene encoding includes geometric feature encoding and appearance feature encoding, and back-propagates to the decoder, so that the decoder can be updated with parameters to further improve the accuracy of the generated predicted image.

[0118] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.

[0119] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0120] If the unit of the integrated method is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a memory, including a number of instructions to enable a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned memory includes: U disk, read-only memory (ROM), random access memory (RAM), mobile hard disk, disk or optical disk and other media that can store program codes.

[0121] Those skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by entering a program to instruct related hardware, and the program can be stored in a computer-readable memory, which can include: a flash disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc. The above is only an exemplary embodiment of the present disclosure and cannot be used to limit the scope of the present disclosure.

[0122] Those skilled in the art will readily appreciate the embodiments of the present disclosure after considering the specification and practicing the disclosure herein. This application is intended to cover any variations, uses or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not described in the present disclosure. The specification and examples are intended to be exemplary only, and the scope and spirit of the present disclosure are defined by the claims.

Claims

1. A method for optimizing the convergence speed of a SLAM model, characterized in that: include: Constructing a SLAM model, obtaining pre-training data, wherein the pre-training data includes a pre-training image and a pre-training depth image corresponding to the pre-training image, generating a 3D point cloud set based on the pre-training data, and determining a first scene code of each first sampling point of the pre-training image in a global voxel feature grid through the 3D point cloud set, wherein the first scene code includes a scene geometry code and a scene appearance code; Performing volume rendering on the scene encoding in the SLAM model to generate a prediction result, wherein the prediction result includes a predicted image and a predicted depth image corresponding to the pre-trained image; Constructing a first loss function of the SLAM model, optimizing the first loss function based on a comparison difference between a prediction result and pre-training data, and obtaining an optimized SLAM model; In the map building process of the optimized SLAM model, according to the image sequence of the first input images captured by the depth camera, at least a preset number of the first input images whose shooting time is after a preset time are continuously collected, and at least a preset number of the first input images whose shooting time is before a preset time are randomly collected, and volume rendering is performed on each of the collected first input images. According to the optimized first loss function, a synchronous pose update is performed on each of the first input images after volume rendering, and the map building process parameters of the optimized SLAM model are updated based on the synchronous pose update result.

2. The method according to claim 1, characterized in that The step of generating a 3D point cloud set based on the pre-training data and determining a first scene code of each first sampling point of the pre-training image in a global voxel feature grid through the 3D point cloud set includes: For any of the pre-trained depth images, determine image feature information of the pre-trained image corresponding to the pre-trained depth image, determine each first sampling point based on the image feature information, and draw each first sampling point of the pre-trained image into a world coordinate system corresponding to the feature three planes to obtain a 3D point cloud set; Based on the prior encoder, the 3D point cloud set is geometrically encoded to obtain a global voxel feature grid; Decoding the global voxel feature grid based on the decoder to obtain a first scene code of the global voxel feature grid; Ray sampling is performed on the pre-trained image, and first scene codes corresponding to each first sampling point are obtained from the global voxel feature grid according to the ray sampling result.

3. The method according to claim 2, characterized in that The geometric prior encoding of the 3D point cloud set by the prior encoder to obtain a global voxel feature grid includes: Acquire prior data corresponding to the pre-trained image, the prior data comprising first prior data corresponding to scene appearance and second prior data corresponding to scene geometry; Training the a priori encoder using the a priori data, and optimizing parameters of the a priori encoder according to the training result of the a priori encoder; Performing geometric a priori encoding on local voxels by using the optimized a priori encoder to obtain a local voxel set, wherein the local voxels are voxels corresponding to all 3D point clouds in the 3D point cloud set; The local voxel set is fused into the global voxel set, and a global voxel feature grid is constructed according to the fused global voxel set.

4. The method according to claim 2, characterized in that: The decoder includes an appearance decoder and a geometry decoder. The appearance decoder is used to decode the scene appearance code in the global voxel feature grid and obtain the RGB values ​​corresponding to each of the first sampling points of the pre-trained image; the geometry decoder is used to decode the scene geometry code in the global voxel feature grid and obtain the TSDF values ​​corresponding to each of the first sampling points of the pre-trained image.

5. The method according to claim 2, characterized in that: The performing ray sampling on the pre-training image comprises: Obtaining the depth camera pose and sampling step size corresponding to the pre-trained image; Determine at least two rays according to the depth camera pose and the sampling step, sample the pre-trained image along the ray direction, and map the second sampling point of each ray to the global voxel feature grid; A first scene code corresponding to the second sampling point is obtained to generate a ray sampling result.

6. The method according to claim 1, characterized in that The method further comprises: In the map building process of the optimized SLAM model, the depth camera pose of the camera tracking process is gradient updated by using the optimized first loss function.

7. The method according to claim 1, characterized in that The first loss function includes a second loss function, a third loss function and a fourth loss function; the second loss function is used to generate a rendered image that approaches the pre-training image based on the prediction result and the pre-training data; the third loss function is used to predict the TSDF value of the first sampling point corresponding to the pre-training image and located in the truncation area near the ray surface, so that the prediction result of the TSDF value of the first sampling point approaches the TSDF value of the pre-training image; the fourth loss function is used to predict the TSDF value of the first sampling point corresponding to the pre-training image and located between the center of the depth camera and the truncation area, so that the TSDF value of the first sampling point approaches a predetermined fixed value of the truncation distance.

8. The method according to claim 7, characterized in that The second loss function includes: a fifth loss function, the fifth loss function being used to calculate color loss of the pre-trained image and the predicted image generated by volume rendering; A sixth loss function is used to calculate the depth loss of the pre-training image and the predicted depth image generated by volume rendering.

9. The method according to claim 1, characterized in that: The method further comprises: In the camera tracking process of the optimized SLAM model, local pixels of the first input image are randomly sampled, and global pixels of the first input image are replaced based on the sampled local pixels.

10. The method according to claim 1, characterized in that After synchronously updating the pose of each of the first input images after volume rendering according to the optimized first loss function, the method further includes: A second scene code of a two-dimensional depth image corresponding to a second input image is obtained, and decoder parameters are updated based on the second scene code, wherein the second input image is obtained by performing synchronous posture update on the first input image after volume rendering by the first loss function.