A Visual Localization Method Based on Inverse Neural Radiance Field

Through the visual positioning method based on the inverse neural radiation field, the deep learning neural network and the improved Transformer model are used to directly reversely solve the camera position pose, solving the problem of insufficient real-time and adaptability in the existing technology, and achieving efficient visual positioning in various scenarios.

CN116630420BActive Publication Date: 2025-06-10HANGZHOU EBOYLAMP ELECTRONICS CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310496432.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-05
Publication Date
2025-06-10
Estimated Expiration
2043-05-05

AI Technical Summary

Technical Problem

The existing visual positioning technology has shortcomings in real-time and adaptability, especially in scenarios where image continuity and texture characteristics are obvious, and the initialization time is long and may fail, so it cannot adapt to high-real-time scenarios.

Method used

A visual positioning method based on inverse neural radiation field is proposed. By constructing a deep learning neural network model, it uses three-dimensional rendering data sets and artificially annotated real scene data sets for training, and combines the improved Transformer model and specific artificial intelligence acceleration hardware adaptation to directly reversely solve the camera position pose.

Benefits of technology

It realizes efficient visual positioning without the need for image continuity and obvious texture features, improves real-time and adaptability, and can accurately obtain camera position information in various scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116630420B_ABST
    Figure CN116630420B_ABST
Patent Text Reader

Abstract

The present invention discloses a visual positioning method based on an inverse neural radiance field, including collecting a three-dimensional rendering data set and an artificially collected and labeled real-scene data set. This visual positioning method based on the inverse neural radiance field uses the images captured by the camera as input, directly reversely solves the pose of the camera itself to achieve visual positioning, and solves the problems of poor real-time performance, the need for continuous images, and obvious texture features in the existing visual positioning method based on image feature points; during the training process of the deep learning neural network model, it is divided into pre-training and correction training. During the pre-training process, a three-dimensional rendering data set is used, which is relatively simple to obtain and the obtained scenes are relatively rich and diverse. Then, during the correction training process, an artificially collected and labeled real-scene data set is used, so that the deep learning neural network model can not only have the ability to adapt to multiple scenes but also have a more accurate effect for real scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of visual positioning, and particularly relates to a visual positioning method based on an inverse neural radiance field. Background Art

[0002] In recent years, the industry has studied a technology that can synthesize images from new perspectives through a deep learning neural network given camera pose parameters, which is called a neural radiance field. Among them, the deep learning neural network is a multi-layer perceptron structure (MLP).

[0003] Visual positioning technology is one of the main technologies for intelligent processing of video images in the current industry, and has received extensive attention and applications in civil and military defense fields such as autonomous driving and autonomous navigation of unmanned aerial vehicles. Among the current mainstream visual positioning technologies, typical applications such as the visual SLAM method are basically based on image feature point extraction and feature point matching, and then combined with camera parameters to explicitly model and calculate the position of itself. This method has great limitations and is applicable only in scenarios where the captured images are continuous and the images have obvious texture features; in addition, this method has a long initialization time and even fails to initialize, and cannot adapt to scenarios with high real-time requirements. Summary of the Invention

[0004] The purpose of the present invention is to propose a visual positioning method based on an inverse neural radiance field for solving the problems raised in the background art.

[0005] To achieve the above purpose, the technical solution adopted by the present invention is as follows:

[0006] A visual positioning method based on an inverse neural radiance field proposed by the present invention includes:

[0007] Collect a three-dimensional rendering data set and a manually collected and labeled real-scene data set.

[0008] Construct a deep learning neural network model including a first convolutional neural network, a second convolutional neural network, a multi-layer perceptron structure, an encoder, a decoder, and a feed-forward neural network.

[0009] Use the three-dimensional rendering data set and the manually collected and labeled real-scene data set to train the deep learning neural network model.

[0010] Perform model lightweight processing and specific artificial intelligence acceleration hardware adaptation on the trained deep learning neural network model.

[0011] Deploy the adapted deep learning neural network model to a computing device, input an image captured by a camera into the computing device deployed with the deep learning neural network model, and then the computing device infers and calculates the pose information of the camera.

[0012] Preferably, collect a three-dimensional rendering data set and a manually collected and labeled real-scene data set, including:

[0013] The three-dimensional rendering data set generates images of different scenes and corresponding virtual camera pose information through a three-dimensional rendering engine, and the number of different scenes is not less than 100, and the number of images generated by the virtual camera for each scene is not less than 500.

[0014] The manually collected and labeled real-scene data set obtains images and corresponding real camera pose information by shooting a specific scene with a real camera, and the number of obtained images is not less than 10,000.

[0015] Preferably, construct a deep learning neural network model including a first convolutional neural network, a second convolutional neural network, a multi-layer perceptron network structure, an encoder, a decoder, and a feed-forward neural network, including:

[0016] The first convolutional neural network is ResNet50 and adds a fully connected layer with 128-dimensional output at the last layer, and the second convolutional neural network is ResNet18 and adds a fully connected layer with 128-dimensional output at the last layer.

[0017] The number of layers of the multi-layer perceptron network structure is at least four, and the input feature dimension is 128, the output feature dimension is 128, and at least one layer in the hidden layer has no less than 1024 neurons.

[0018] Perform feature fusion by corresponding feature weighted addition of the first feature vector output by the first convolutional neural network and the second feature vector output by the second convolutional neural network to obtain a first fused feature vector, and the weight of the first feature vector is w 1 , the weight of the second feature vector is w 2 , and w 1 +w 2 =1.0.

[0019] Input the first fused feature vector into the encoder for feature extraction to obtain a first feature encoded output vector.

[0020] Input the first feature encoded output vector into the multi-layer perceptron network structure to obtain a second feature encoded output vector.

[0021] Input the second feature encoded output vector into the decoder, and input a preset number m of query key-value vectors into the decoder, and obtain m decoded feature outputs after decoding.

[0022] Finally, input the m decoded feature outputs into the corresponding number of feed-forward neural networks in sequence to obtain m pose information of the camera, where the value of m is 6.

[0023] Preferably, the deep learning neural network model is trained using a three-dimensional rendering dataset and a manually collected and labeled real-scene dataset, including:

[0024] First, the deep learning neural network model is pre-trained using the three-dimensional rendering dataset to obtain a pre-trained deep learning neural network model, and then the pre-trained deep learning neural network model is corrected and trained using the manually collected and labeled real-scene dataset to obtain the finally trained deep learning neural network model.

[0025] Preferably, during the pre-training using the three-dimensional rendering dataset and the correction training using the manually collected and labeled real-scene dataset, the following steps are used for training:

[0026] Randomly initialize the parameter values and query key value vectors of the deep learning neural network model, select images from the three-dimensional rendering dataset or the manually collected and labeled real-scene dataset, and uniformly scale the selected images to a size of 1024×1024 pixels. The images are divided into 8×8 grid regions to obtain the training input of the first convolutional neural network;

[0027] Extract a proportion of r grids from the input of the first convolutional neural network for random occlusion to obtain the training input of the second convolutional neural network, where 0.3 < r < 0.6;

[0028] Input the training input of the first convolutional neural network and the training input of the second convolutional neural network into the first convolutional neural network and the second convolutional neural network in sequence, and train the deep learning neural network model or the pre-trained deep learning neural network model in combination with the loss function. The formula of the loss function is as follows:

[0029] L = λ 1 L loc + λ 2 L pos

[0030]

[0031]

[0032] Among them, L represents the loss function value, L loc represents the camera position loss value, L loc represents the camera pose loss value, x, y, z respectively represent the true values of the camera position coordinates, x, y, z respectively represent the predicted values of the camera position coordinates, d x , d y , d z represent the true values of the rotation angles of the camera around the three coordinate axes when generating the scene image, which are used to represent the three camera poses, d x ', d y ', dz ' represents the predicted values of the poses of three cameras, λ 1 , λ 2 successively represent the weights of the camera position loss and the pose loss, and 0.5 < λ 1 < 1.0, 3 < λ 2 < 7;

[0033] According to the loss function value, use backpropagation to calculate the magnitude of the parameter gradient of the deep learning neural network model, and use the gradient to update the parameter values of the deep learning neural network model or the pre-trained deep learning neural network model until the loss function converges or reaches the number of training times to stop training, obtaining the pre-trained deep learning neural network model or the finally trained deep learning neural network model. And during the training process of the pre-trained deep learning neural network model with the artificially collected and labeled real-scene dataset, freeze the encoder parameters and decoder parameters so that they do not update during the backpropagation training process.

[0034] Preferably, perform model lightweight processing and specific artificial intelligence acceleration hardware adaptation on the trained deep learning neural network model, including:

[0035] Quantize the parameters of the trained deep learning neural network model so that the data type of the deep learning neural network model parameters is stored as an integer type from the floating-point type to obtain the first deep learning quantization parameter model.

[0036] Use the tool chain supporting the artificial intelligence acceleration hardware to perform model conversion on the first deep learning quantization parameter model to obtain the second deep learning quantization parameter model adapted to the selected artificial intelligence acceleration hardware.

[0037] Preferably, deploy the adapted deep learning neural network model to a computing device, input the image captured by the camera into the computing device where the deep learning neural network model is deployed, and then the computing device performs inference calculation to obtain the pose information of the camera, including:

[0038] Load the second deep learning quantization parameter model into the memory of the computing device.

[0039] Obtain the image captured by the camera and scale the image to a size of 1024×1024 pixels.

[0040] Divide the scaled image into 8×8 grid regions to obtain the first input image.

[0041] Select a grid with a ratio of r from the first input image to uniformly occlude the first input image to obtain the second input image.

[0042] Input the first input image and the second input image into a computing device loaded with a second deep learning quantization parameter model, and finally the computing device calculates the pose information of the camera itself.

[0043] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0044] 1. This visual positioning method based on inverse neural radiance fields uses the images captured by the camera as input, directly inversely solves the pose of the camera itself to achieve visual positioning, and solves the problems of poor real-time performance, the need for continuous images, and obvious texture features in the visual positioning method based on image feature points in the prior art;

[0045] 2. This visual positioning method based on inverse neural radiance fields uses an improved Transformer deep learning neural network model. Compared with the existing Transformer deep learning neural network model, it has fewer parameters, can well fit high-frequency information, and both the original image and the randomly occluded image are used for feature extraction and fusion during the training process, and it has stronger adaptability to object occlusion in the actual application scenario;

[0046] 3. In this visual positioning method based on inverse neural radiance fields, the training process of the deep learning neural network model is divided into pre-training and correction training. During the pre-training process, a three-dimensional rendering data set is used, which is relatively simple to obtain and the obtained scenes are relatively rich and diverse. Then, in the correction training process, a real-scene data set collected and labeled manually is used, so that the deep learning neural network model can not only have the ability to adapt to multiple scenes but also have more accurate effects for real scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 It is a schematic structural diagram of the Transformer deep learning neural network model in the prior art of the present invention;

[0048] Figure 2 It is a block diagram of the deep learning neural network model in the visual positioning method based on inverse neural radiance fields of the present invention;

[0049] Figure 3 It is a schematic diagram of processing the input training image during the training process of the present invention;

[0050] Figure 4 It is a schematic diagram of processing the input image after the training of the present invention;

[0051] Figure 5 It is a schematic diagram of comparing the real position information with the position information obtained by the method of the present invention during the test experiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0053] It should be noted that when a component is referred to as being "connected" to another component, it can be directly connected to the other component or there may also be an intermediate component. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs. The terms used herein in the specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application.

[0054] As Figure 2-4 shown, a visual positioning method based on an inverse neural radiance field includes:

[0055] Step S1, collect a three-dimensional rendering data set and a manually collected and labeled real-scene data set.

[0056] Specifically, the three-dimensional rendering data set is generated by a three-dimensional rendering engine to generate different scene images and corresponding virtual camera pose information, and the number of different scenes is not less than 100, and the number of images generated by each scene through the virtual camera is not less than 500. In this embodiment, the three-dimensional rendering engine used is Three.js. The number of different scenes generated by the Three.js three-dimensional rendering engine is 150, and the number of images rendered by adjusting the position and shooting angle of the virtual camera for each scene is 600. The total number of data sets collected is 90,000. The pose information includes the camera position and the camera attitude. Among them, the camera position is a three-dimensional coordinate, expressed as (x, y, z), and the camera attitude is the rotation angle of the camera around the three coordinate axes when generating the scene photo, expressed as (d x , d y , d z );

[0057] The manually collected and labeled real-scene data set is obtained by using a real camera to shoot a specific scene to obtain images and corresponding real camera pose information, and the number of images obtained is not less than 10,000. In this embodiment, the number of data sets collected is 10,342, and the specific scene is the scene for actual deployment applications.

[0058] Step S2, construct a deep learning neural network model including a first convolutional neural network, a second convolutional neural network, a multi-layer perceptron structure, an encoder, a decoder, and a feed-forward neural network.

[0059] Specifically, the deep learning neural network model in this embodiment is an improved Transformer deep learning neural network model. The structural schematic diagram of the existing Transformer deep learning neural network model is as shown in Figure 1 . Compared with the existing Transformer deep learning neural network model, it has fewer parameters and can well fit high-frequency information. The first convolutional neural network is ResNet50 with a fully connected layer added at the last layer to output 128 dimensions, and the second convolutional neural network is ResNet18 with a fully connected layer added at the last layer to output 128 dimensions;

[0060] The number of layers of the multi-layer perceptron network structure is at least four, and the input feature dimension is 128, the output feature dimension is 128, and at least one layer in the hidden layer has no less than 1024 neurons.

[0061] In this embodiment, the multi-layer perceptron network structure has four layers, including a 128-dimensional input layer, a 128-dimensional output layer, and 1024-dimensional and 512-dimensional hidden layers.

[0062] The first feature vector output by the first convolutional neural network and the second feature vector output by the second convolutional neural network are subjected to feature fusion by corresponding feature weighted addition to obtain a first fusion feature vector, and the weight of the first feature vector is w 1 , the weight of the second feature vector is w 2 , and w 1 +w 2 =1.0.

[0063] In this embodiment, the weight of the first feature vector is w 1 =0.75, and the weight of the second feature vector is w 2 =0.25.

[0064] The first fusion feature vector is input into the encoder for feature extraction to obtain a first feature encoded output vector.

[0065] The first feature encoded output vector is input into the multi-layer perceptron network structure to obtain a second feature encoded output vector.

[0066] The second feature encoded output vector is input into the decoder, and a preset number m of query key-value vectors (Query) are input into the decoder. After decoding, m decoded feature outputs are obtained.

[0067] Finally, the m decoded feature outputs are sequentially input into the feedforward neural network with the corresponding number to obtain the m pose information of the camera, where the value of m is 6.

[0068] In this embodiment, the six pose information of the camera obtained are three representing the position coordinates of the camera, and the other three representing different poses of the camera. The query key-value vector is a trainable parameter, and the initial value of the elements of the query key-value vector is a number greater than 0 and less than 1.

[0069] Step S3: Use the three-dimensional rendering dataset and the manually collected and labeled real-scene dataset to train the deep learning neural network model.

[0070] Specifically, first, use the three-dimensional rendering dataset to pre-train the deep learning neural network model to obtain a pre-trained deep learning neural network model, and then use the manually collected and labeled real-scene dataset to correct and train the pre-trained deep learning neural network model to obtain the finally trained deep learning neural network model (in the pre-training process, the three-dimensional rendering dataset is used, which is relatively easy to obtain and the obtained scenes are relatively rich and diverse. Then, in the correction training process, the manually collected and labeled real-scene dataset is used, so that the deep learning neural network model can not only have the ability to adapt to multiple scenes but also have a more accurate effect for real scenes).

[0071] In the pre-training process using the three-dimensional rendering dataset, the following steps are used for training:

[0072] Step S3.1: Randomly initialize the parameter values of the deep learning neural network model and the query key-value vector, select images from the three-dimensional rendering dataset, and uniformly scale the selected images to a size of 1024×1024 pixels. The images are divided into 8×8 grid regions (i.e., 64 grid regions in total, 8 rows and 8 columns) to obtain the training input of the first convolutional neural network.

[0073] Step S3.2: Randomly occlude a proportion r of the grids extracted from the input of the first convolutional neural network to obtain the training input of the second convolutional neural network, where 0.3 < r < 0.6. In this embodiment, r = 0.5.

[0074] Step S3.3: Input the training input of the first convolutional neural network and the training input of the second convolutional neural network into the first convolutional neural network and the second convolutional neural network in sequence (extract the high-dimensional feature vectors of the images through each convolutional neural network and obtain the six pose information of the camera by advancing the deep learning neural network model), and train the deep learning neural network model in combination with the loss function (calculate the loss using the camera pose information corresponding to the images selected from the three-dimensional rendering dataset), and the formula of the loss function is as follows:

[0075] L = λ 1 L loc +λ 2 L pos

[0076]

[0077]

[0078] Among them, L represents the loss function value, L loc represents the camera position loss value, L loc represents the camera pose loss value, x, y, z respectively represent the true values of the camera position coordinates, x, y, z respectively represent the predicted values of the camera position coordinates, d x , d y , d z represents the true values of the rotation angles of the camera around the three coordinate axes when generating the scene image, which are used to represent the three camera poses, d x ', d y ', d z ' represents the predicted values of the three camera poses, λ 1 , λ 2 successively represent the weights of the camera position loss and the pose loss, and 0.5 < λ 1 < 1.0, 3 < λ 2 < 7;

[0079] In this embodiment, λ 1 = 0.6, λ 2 = 5.

[0080] Step S3.4: According to the loss function value, use backpropagation to calculate the magnitude of the parameter gradient of the deep learning neural network model, and use the gradient to update the parameter values of the deep learning neural network model until the loss function converges or reaches the number of training times, and then stop training to obtain the pre-trained deep learning neural network model;

[0081] During the training process of correcting with the artificially collected and labeled real scene dataset, the following steps are adopted for training:

[0082] Step S3.5: Randomly initialize the parameter values and query key value vectors of the pre-trained deep learning neural network model;

[0083] Step S3.6: Freeze the encoder parameters and decoder parameters so that they are not updated during the backpropagation training process. Select images and the corresponding real camera pose information from the artificially collected and labeled real scene dataset as training data, and repeat steps S3.1 - S3.4 to obtain the finally trained deep learning neural network model.

[0084] Step S4: Perform model lightweight processing and specific artificial intelligence acceleration hardware adaptation on the trained deep learning neural network model.

[0085] Specifically, the parameters of the trained deep learning neural network model are quantized so that the data type of the deep learning neural network model parameters is stored as an integer from the floating-point type to obtain the first deep learning quantization parameter model;

[0086] Use the tool chain supporting the artificial intelligence acceleration hardware to convert the first deep learning quantization parameter model to obtain the second deep learning quantization parameter model adapted to the selected artificial intelligence acceleration hardware. In this embodiment, the selected artificial intelligence acceleration hardware is the Huawei Atlas200 dedicated neural network inference hardware, and the model is converted into a file in the.om format using the model conversion tool chain provided by Huawei so that the model parameters can be correctly loaded into the memory for inference calculation.

[0087] Step S5: Deploy the adapted deep learning neural network model to the computing device, use the camera to capture an image and input it into the computing device where the deep learning neural network model is deployed, and then the computing device performs inference calculation to obtain the pose information of the camera.

[0088] Specifically, in step S5.1: Load the second deep learning quantization parameter model into the memory of the computing device;

[0089] Step S5.2: Obtain the image captured by the camera and scale the image to a size of 1024×1024 pixels;

[0090] Step S5.3: Divide the scaled image into 8×8 grid regions to obtain the first input image (i.e., the input of the first convolutional neural network);

[0091] Step S5.4: Uniformly occlude the first input image for the grids with a ratio of r in the first input image to obtain the second input image (i.e., the input of the second convolutional neural network), where r = 0.5;

[0092] Step S5.5: Input the first input image and the second input image into the computing device loaded with the second deep learning quantization parameter model, and finally the computing device calculates the pose information of the camera itself.

[0093] Finally, in order to verify the effect of the method of the present invention, a test experiment was designed, and the present invention uses a quadrotor UAV for low-altitude flight vision positioning test.

[0094] The specific steps are as follows:

[0095] Step 1: Use a theodolite to measure the accurate position of the marked point and establish a coordinate system with this position as the coordinate origin;

[0096] Step 2: The UAV is at different heights (50m, 100m, 150m), and at each height, sensor image data with different attitude angles (0°, 20°, 40°) is captured;

[0097] Step 3: Using the image data as input, obtain the pose information of the UAV itself through the method of the present invention for visual positioning;

[0098] Step 4: Taking the UAV's own GPS positioning information as the ground truth, convert it to the coordinates in the coordinate system established in Step 1 as the ground truth coordinates, and compare the ground truth coordinates with the pose information obtained through the method of the present invention;

[0099] The obtained results are as Figure 5 shown. The dashed line represents the real position information, and the solid line represents the position information obtained through the method of the present invention. It can be seen that the visual positioning information obtained through the method of the present invention basically coincides with the real position information, with high accuracy. This proves the effectiveness of the method of the present invention.

[0100] This visual positioning method based on inverse neural radiance fields uses the images captured by the camera as input, directly reversely solves the pose of the camera itself to achieve visual positioning, and solves the problems of poor real-time performance, the need for continuous images, and obvious texture features in the existing visual positioning methods based on image feature points; this visual positioning method based on inverse neural radiance fields adopts an improved Transformer deep learning neural network model. Compared with the existing Transformer deep learning neural network models, it has fewer parameters, can better fit high-frequency information, and both the original image and the randomly occluded images are used for feature extraction and fusion during the training process, and it has stronger adaptability to object occlusion in the actual application scenario; in this visual positioning method based on inverse neural radiance fields, the training process of the deep learning neural network model is divided into pre-training and correction training. During the pre-training process, a three-dimensional rendering dataset is used, which is relatively easy to obtain and the obtained scenes are rich and diverse. Then, during the correction training process, a manually collected and labeled real-scene dataset is used, so that the deep learning neural network model can not only have the ability to adapt to multiple scenes but also have more accurate effects for real scenes.

[0101] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0102] The above-described embodiments only express the embodiments of the present application that are described in more specific and detailed terms, but should not be construed as limiting the scope of the patent application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A visual localization method based on inverse neural radiance fields, characterized in that: The visual localization method based on inverse neural radiance fields includes: Collecting a three-dimensional rendering dataset and an artificially collected and labeled real-scene dataset; Constructing a deep learning neural network model including a first convolutional neural network, a second convolutional neural network, a multi-layer perceptron structure, an encoder, a decoder, and a feed-forward neural network; Training the deep learning neural network model using the three-dimensional rendering dataset and the artificially collected and labeled real-scene dataset; Performing model lightweight processing and specific artificial intelligence acceleration hardware adaptation on the trained deep learning neural network model; Deploying the adapted deep learning neural network model to a computing device, inputting an image captured by a camera into the computing device deployed with the deep learning neural network model, and then the computing device performs inference calculation to obtain the pose information of the camera; Among them, the construction of the deep learning neural network model including a first convolutional neural network, a second convolutional neural network, a multi-layer perceptron network structure, an encoder, a decoder, and a feed-forward neural network includes: The first convolutional neural network is ResNet50 and a fully connected layer with a 128-dimensional output is added to the last layer, and the second convolutional neural network is ResNet18 and a fully connected layer with a 128-dimensional output is added to the last layer; The number of layers of the multi-layer perceptron network structure is at least four, and the input feature dimension is 128, the output feature dimension is 128, and at least one hidden layer has no less than 1024 neurons; The first feature vector output by the first convolutional neural network and the second feature vector output by the second convolutional neural network are subjected to feature fusion by corresponding feature weighted addition to obtain a first fused feature vector, and the weight of the first feature vector is w 1 , the weight of the second feature vector is w 2 , and w 1 +w 2 = 1.0; Inputting the first fusion feature vector into the encoder for feature extraction to obtain a first feature encoding output vector; Inputting the first feature encoding output vector into the multi-layer perceptron network structure to obtain a second feature encoding output vector; Inputting the second feature encoding output vector into the decoder, and inputting a preset number m of query key-value vectors into the decoder, and obtaining m decoded feature outputs after decoding; Finally, inputting the m decoded feature outputs into the corresponding number of feed-forward neural networks in sequence to obtain the m pose information of the camera, where the value of m is 6; The training of the deep learning neural network model using the three-dimensional rendering dataset and the artificially collected and labeled real-scene dataset includes: First, pre-training the deep learning neural network model using the three-dimensional rendering dataset to obtain a pre-trained deep learning neural network model, and then correcting and training the pre-trained deep learning neural network model using the artificially collected and labeled real-scene dataset to obtain the finally trained deep learning neural network model; During the pre-training using the three-dimensional rendering dataset and the correction training using the artificially collected and labeled real-scene dataset, the following steps are used for training: Randomly initializing the parameter values of the deep learning neural network model and the query key-value vectors, selecting an image from the three-dimensional rendering dataset or the artificially collected and labeled real-scene dataset, and uniformly scaling the selected image to a size of 1024×1024 pixels, and dividing the image into 8×8 grid regions to obtain the training input of the first convolutional neural network; Extract a grid with a proportion of r from the input of the first convolutional neural network for random occlusion to obtain the training input of the second convolutional neural network, where 0.3 < r < 0.6; Input the training input of the first convolutional neural network and the training input of the second convolutional neural network into the first convolutional neural network and the second convolutional neural network in sequence, and train the deep learning neural network model or the pre-trained deep learning neural network model in combination with the loss function. The formula of the loss function is as follows: L = λ 1 L loc + λ 2 L pos Among them, L represents the loss function value, L loc represents the camera position loss value, L loc represents the camera attitude loss value, x, y, z respectively represent the true values of the camera position coordinates, x, y, z respectively represent the predicted values of the camera position coordinates, d x , d y , d z represents the true values of the rotation angles of the camera around the three coordinate axes when generating the scene image, used to represent the three camera attitudes, d x ', d y ', d z ' represents the predicted values of the three camera attitudes, λ 1 , λ 2 successively represent the weights of the camera position loss and the attitude loss, and 0.5 < λ 1 < 1.0, 3 < λ 2 < 7; According to the loss function value, use backpropagation to calculate the parameter gradient magnitude of the deep learning neural network model, and use the gradient to update the parameter values of the deep learning neural network model or the pre-trained deep learning neural network model until the loss function converges or reaches the number of training times to stop training, obtaining the pre-trained deep learning neural network model or the finally trained deep learning neural network model. During the training process of the pre-trained deep learning neural network model with the artificially collected and labeled real-scene dataset, freeze the encoder parameters and decoder parameters so that they do not update during the backpropagation training process.

2. The visual localization method based on the inverse neural radiance field according to claim 1, characterized in that: The collection of the three-dimensional rendered dataset and the artificially collected and labeled real-scene dataset includes: The three-dimensional rendered dataset is generated by a three-dimensional rendering engine to generate different scene images and corresponding virtual camera pose information, and the number of different scenes is not less than 100, and the number of images generated by the virtual camera for each scene is not less than 500; The artificially collected and labeled real-scene dataset is obtained by using a real camera to capture images of a specific scene and the corresponding real camera pose information, and the number of acquired images is not less than 10,000.

3. The visual localization method based on the inverse neural radiance field according to claim 1, characterized in that: The model lightweight processing and specific artificial intelligence acceleration hardware adaptation of the trained deep learning neural network model include: Quantize the parameters of the trained deep learning neural network model so that the data type of the deep learning neural network model parameters is stored as an integer from a floating-point type to obtain the first deep learning quantization parameter model; Use the tool chain supporting the artificial intelligence acceleration hardware to perform model conversion on the first deep learning quantization parameter model to obtain the second deep learning quantization parameter model adapted to the selected artificial intelligence acceleration hardware.

4. The visual localization method based on the inverse neural radiance field according to claim 3, characterized in that: Deploy the adapted deep learning neural network model to a computing device, input the image captured by the camera into the computing device deployed with the deep learning neural network model, and then the computing device performs inference calculation to obtain the pose information of the camera, including: Load the second deep learning quantization parameter model into the computing device memory; Obtain the image captured by the camera and scale the image to a size of 1024×1024 pixels; Divide the scaled image into 8×8 grid regions to obtain the first input image; Select a grid with a proportion of r from the first input image to uniformly occlude the first input image to obtain the second input image; Input the first input image and the second input image into a computing device loaded with a second deep learning quantization parameter model, and finally the computing device calculates the pose information of the camera itself.

Citation Information

Patent Citations

  • A weak texture three-dimensional object attitude estimation method and device

    CN109934847A

  • Visual servo method and device

    CN114942591A

  • Training method of incomplete information target recognition model and target recognition method

    CN115761444A