Monocular visual odometer method, device and equipment based on deep learning and medium
Through the monocular visual odometry method based on deep learning, the context information of two adjacent frames of images is extracted and timing analysis is performed, which solves the problems of high manual design cost and low pose estimation accuracy of traditional visual mileage calculation methods, and achieves higher pose estimation accuracy and calculation efficiency.
Patent Information
- Application Number
- CN202510198710.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-13
AI Technical Summary
The manual design algorithms and models of traditional visual mileage calculation methods lead to high labor costs and the accuracy of pose estimation is easily disturbed by external factors. The accuracy of pose estimation based on deep learning is generally low.
Using a monocular visual odometry method based on deep learning, the image set acquired by the monocular camera is sized, the context information of two adjacent frames of images is extracted, and the timing dimension information of the target features is analyzed using feature association weights and timing learning modules, and finally three-dimensional pose estimation is performed through the pose network module.
It improves the accuracy of position estimation, reduces the computational complexity caused by image size differences, enhances the ability to fusion of image information at different moments and different perspectives, and provides more comprehensive and accurate timing information.
Smart Images

Figure CN119991811A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a monocular visual odometer method, device, equipment and medium based on deep learning. Background Art
[0002] In the fields of robotics, unmanned driving, augmented reality and virtual reality, simultaneous localization and mapping (SLAM) is a key technology for achieving navigation and positioning. Among them, visual SLAM uses cameras as sensors to simultaneously complete positioning and mapping in unknown environments. Visual odometry (VO) is its core part, which determines the pose by processing image sequences. The accuracy of its pose estimation directly affects the subsequent mapping process.
[0003] Traditional visual odometer calculation methods rely on manually designed algorithms and models. Since the calculation process involves a large number of mathematical and physical operations, the labor cost is high, and the accuracy of pose estimation is easily affected by external factors such as occlusion, textureless scenes, and lighting changes, resulting in limited mapping accuracy. With the breakthrough of deep learning technology in the field of computer vision, the visual odometer calculation method based on deep learning has emerged. It obtains pose information through autonomous training of neural networks. However, the accuracy of its pose estimation is generally low. Summary of the invention
[0004] In view of this, the purpose of the present invention is to provide a monocular visual odometer method, device, equipment and medium based on deep learning, which can improve the accuracy of pose estimation. The specific scheme is as follows:
[0005] In a first aspect, the present application discloses a monocular visual odometer method based on deep learning, comprising:
[0006] Resize the image set collected by the monocular camera and obtain two adjacent frames of the resized image;
[0007] The context information of two adjacent frames is extracted through the feature extraction module of the preset monocular visual odometer model, and the feature association weight of the context information is obtained. Then, the target feature is obtained according to the feature association weight and the convolution operation result of the standard convolution layer on the context information.
[0008] The temporal dimension information of the target features is analyzed through the temporal learning module of the preset monocular visual odometer model to obtain the temporal features for camera pose estimation;
[0009] The three-dimensional translational degree of freedom and rotational degree of freedom of the temporal features are estimated through the posture network module of the monocular visual odometry model to obtain a first translation vector and a first rotation vector, and a six-degree-of-freedom posture is obtained according to the first translation vector and the first rotation vector.
[0010] Optionally, the feature extraction module includes a context information acquisition unit and a self-attention mechanism unit, and the context information of two adjacent frames of images is extracted by the feature extraction module of the preset monocular visual odometer model, and the feature association weight of the context information is obtained, including:
[0011] The context information of two adjacent frames of images is extracted by multiple cascaded dilated convolution subunits in the context information acquisition unit; wherein different dilated convolution subunits correspond to different dilation coefficients, and, in each dilated convolution subunit, the input data of the dilated convolution subunit is operated based on the convolution kernel size of the dilated convolution subunit and the dilation coefficient of the dilated convolution subunit to obtain the output result of each dilated convolution subunit, for the first dilated convolution subunit, the input data of the dilated convolution subunit is the two adjacent frames of images, and for non-first dilated convolution subunits, the input data of the dilated convolution subunit is the output result of the previous dilated convolution subunit;
[0012] The context information is convolved separately through multiple target convolution layers in the self-attention mechanism unit to obtain the output result of each target convolution layer, and the feature association weight is calculated according to the output result of each target convolution layer.
[0013] Optionally, the target feature is obtained according to the feature association weight and the convolution operation result of the standard convolution layer on the context information, including:
[0014] The context information is convolved through the standard convolutional layer in the self-attention mechanism unit to obtain the convolution result;
[0015] The feature association weights and convolution operation results are cascaded to obtain the target features.
[0016] Optionally, the temporal dimension information of the target features is analyzed by a temporal learning module of a preset monocular visual odometer model to obtain temporal features for camera pose estimation, including:
[0017] The temporal dimension information of the target features is analyzed in sequence through the deep convolution layer, deep expansion convolution layer, and target convolution layer in the temporal learning module, and the output result of the target convolution layer is cascaded with the target features to obtain the temporal features for camera pose estimation.
[0018] Optionally, a three-dimensional translational degree of freedom and a rotational degree of freedom are estimated for the time series features through a posture network module of a monocular visual odometer model to obtain a first translation vector and a first rotation vector, and a six-degree-of-freedom posture is obtained according to the first translation vector and the first rotation vector, including:
[0019] The three-dimensional translational degree of freedom and rotational degree of freedom of the temporal features are estimated through multiple fully connected layers in the pose network module to obtain the first translation vector and the first rotation vector, and the six-degree-of-freedom pose is obtained based on the cascade result of the first translation vector and the first rotation vector.
[0020] Optional, deep learning-based monocular visual odometry method, also includes:
[0021] Inputting the sample image features into the model to be trained, so as to obtain a second translation vector and a second rotation vector predicted by the model to be trained;
[0022] Calculating a first training loss according to the second translation vector and the actual translation vector, and calculating a second training loss according to the second rotation vector and the actual rotation vector;
[0023] The total training loss is calculated according to the first training loss and the second training loss until a preset monocular visual odometry model that meets the loss condition is obtained.
[0024] Optionally, calculating a first training loss according to the second translation vector and the actual translation vector, and calculating a second training loss according to the second rotation vector and the actual rotation vector, includes:
[0025] Calculate a first vector difference between the second translation vector and the actual translation vector, and calculate a first training loss according to an L2 norm of the first vector difference;
[0026] A second vector difference between the second rotation vector and the actual rotation vector is calculated, and a second training loss is calculated according to the product of the target balance coefficient and the L2 norm of the second vector difference.
[0027] In a second aspect, the present application discloses a monocular visual odometer device based on deep learning, comprising:
[0028] An image acquisition module is used to resize the image set collected by the monocular camera and obtain two adjacent frames of the resized image;
[0029] A feature extraction module is used to extract context information of two adjacent frames of images through a feature extraction module of a preset monocular visual odometer model, obtain feature association weights of the context information, and then obtain target features based on the feature association weights and the convolution operation results of the standard convolution layer on the context information;
[0030] A timing analysis module is used to analyze the timing dimension information of target features through a timing learning module of a preset monocular visual odometer model to obtain timing features for camera pose estimation;
[0031] The pose estimation module is used to perform three-dimensional translational freedom estimation and rotational freedom estimation on the temporal features through the pose network module of the monocular visual odometer model, obtain a first translation vector and a first rotation vector, and obtain a six-degree-of-freedom pose according to the first translation vector and the first rotation vector.
[0032] In a third aspect, the present application discloses an electronic device, comprising:
[0033] Memory, used to store computer programs;
[0034] A processor is used to execute a computer program to implement the aforementioned deep learning-based monocular visual odometer method disclosed above.
[0035] In a fourth aspect, the present application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the aforementioned deep learning-based monocular visual odometer method disclosed above is implemented.
[0036] It can be seen that the present application proposes a monocular visual odometer method based on deep learning, including: resizing the image set collected by the monocular camera, and obtaining two adjacent frames of images after adjustment; extracting the context information of the two adjacent frames of images through a feature extraction module of a preset monocular visual odometer model, and obtaining the feature association weight of the context information, and obtaining the target feature according to the feature association weight and the convolution operation result of the context information by the standard convolution layer; analyzing the temporal dimension information of the target feature through a temporal learning module of a preset monocular visual odometer model to obtain the temporal feature for camera pose estimation; performing three-dimensional translational degree of freedom estimation and rotational degree of freedom estimation on the temporal feature through the pose network module of the monocular visual odometer model to obtain a first translation vector and a first rotation vector, and obtaining a six-degree-of-freedom pose according to the first translation vector and the first rotation vector. It can be seen that in this application, by resizing the image set collected by the monocular camera and obtaining two adjacent frames of images, it is helpful to reduce the computational complexity caused by the difference in image size and improve the efficiency of algorithm processing. At the same time, the use of two adjacent frames of images can capture richer motion information. Compared with traditional methods, key information can be extracted more effectively, which improves the accuracy of pose estimation to a certain extent. Furthermore, by extracting the context information of two adjacent frames of images, the image information at different times and different perspectives can be better fused. By obtaining the feature association weights of the context information, different feature information can be given different degrees of importance, thereby accurately focusing on the extraction of key features. In addition, through the temporal learning module in the deep learning model, this application can better capture the time dependency in the image sequence, thereby effectively analyzing the temporal changes of the target features, and providing more comprehensive and accurate temporal information for pose estimation, so as to further improve the accuracy of pose estimation. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0038] Figure 1 A flow chart of a monocular visual odometer method based on deep learning disclosed in this application;
[0039] Figure 2 A schematic diagram of a context convolution structure disclosed in this application;
[0040] Figure 3 This is a schematic diagram of the structure of an attention mechanism module disclosed in this application;
[0041] Figure 4 A schematic diagram of the structure of a time series learning module disclosed in this application;
[0042] Figure 5 This is a framework diagram of an end-to-end monocular visual mileage calculation method disclosed in this application;
[0043] Figure 6 A schematic diagram of a monocular visual odometer device based on deep learning disclosed in this application;
[0044] Figure 7 This is a structural diagram of an electronic device disclosed in this application. DETAILED DESCRIPTION
[0045] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0046] Traditional visual odometer calculation methods rely on manually designed algorithms and models. Since the calculation process involves a large number of mathematical and physical operations, the labor cost is high, and the accuracy of pose estimation is easily affected by external factors such as occlusion, textureless scenes, and lighting changes, resulting in limited mapping accuracy. With the breakthrough of deep learning technology in the field of computer vision, the visual odometer calculation method based on deep learning has emerged. It obtains pose information through autonomous training of neural networks. However, the accuracy of its pose estimation is generally low.
[0047] To this end, an embodiment of the present application proposes a monocular visual odometer solution based on deep learning, which can improve the accuracy of pose estimation.
[0048] The present application embodiment discloses a monocular visual odometer method based on deep learning, see Figure 1 As shown, the method includes:
[0049] Step S11: resizing the image set captured by the monocular camera and obtaining two adjacent frames of images after the resizing.
[0050] In this embodiment, the sample image features are input into the model to be trained, so as to predict the second translation vector and the second rotation vector through the model to be trained, and the first training loss is calculated according to the second translation vector and the actual translation vector, and the second training loss is calculated according to the second rotation vector and the actual rotation vector, and then the total training loss is calculated according to the first training loss and the second training loss, until a preset monocular visual odometer model that meets the loss condition is obtained. When training the model, the AdamW optimizer is used to train the neural network, and the initial learning rate is set to , as the number of training rounds increases, the learning rate gradually decays. When calculating the training loss, first calculate the first vector difference between the second translation vector and the actual translation vector, and calculate the first training loss based on the L2 norm of the first vector difference, then calculate the second vector difference between the second rotation vector and the actual rotation vector, and calculate the second training loss based on the product of the target balance coefficient and the L2 norm of the second vector difference. By deriving the training process through this loss calculation method based on the L2 norm of the vector difference and the balance coefficient, the model can show good results in the prediction of translation and rotation vectors. Taking the calculation of translation vector loss as an example, the L2 norm can intuitively reflect the deviation between the predicted value and the actual value in space. The model continuously minimizes this deviation during training, thereby accurately learning the translation information. In the calculation of rotation vector loss, the introduction of the balance coefficient in the model can more reasonably weigh the learning intensity of different parameters and effectively improve the accuracy and stability of the prediction. The following is a specific loss function formula:
[0051] ;
[0052] in, and Represent the predicted second translation vector and the actual translation vector respectively, and Represent the predicted second rotation vector and the actual rotation vector respectively, represents the L2 norm, Indicates the target balance coefficient, which can be set according to actual conditions.
[0053] In this embodiment, after the preset monocular visual odometer model is trained, all images in the image set collected by the monocular camera are resized, and two adjacent frames of images after the adjustment are obtained, and then the two adjacent frames of images are input into the preset monocular visual odometer model. It should be noted that when training the model, the images involved in the sample image features used are also resized images. In some embodiments, the size of the images in the image set can be adjusted to 320×96.
[0054] Step S12: extract the context information of two adjacent frames of images through the feature extraction module of the preset monocular visual odometer model, obtain the feature association weight of the context information, and then obtain the target feature according to the feature association weight and the convolution operation result of the standard convolution layer on the context information.
[0055] In this embodiment, the context information of two adjacent frames of images is extracted by multiple cascaded dilated convolution subunits in the context information acquisition unit; wherein different dilated convolution subunits correspond to different dilation coefficients, and in each dilated convolution subunit, the input data of the dilated convolution subunit is operated based on the convolution kernel size of the dilated convolution subunit and the dilation coefficient of the dilated convolution subunit to obtain the output result of each dilated convolution subunit, for the first dilated convolution subunit, the input data of the dilated convolution subunit is the two adjacent frames of images, and for non-first dilated convolution subunits, the input data of the dilated convolution subunit is the output result of the previous dilated convolution subunit. See Figure 2 ,This application cascades N dilated convolution subunits with different dilation coefficients R to form a context convolution, where ,N=1,2,…,n,R=,r1,r2,…r n The calculation process of each layer of context convolution is as follows:
[0056] ;
[0057] ;
[0058] in, represents the output feature of the i-th dilated convolution subunit, O represents the final output feature after the n dilated convolution subunits are cascaded, F represents the input data, express The convolution kernel of the dilated convolution subunit of size, is the expansion coefficient of the i-th dilated convolution subunit, k is the convolution kernel size of the standard convolution kernel, Represents a convolution operation. This application uses 9 layers of contextual convolutions, each of which is composed of 3 layers of dilated convolution subunits with different dilation coefficients. Each layer of contextual convolution is followed by a ReLU (Rectified Linear Unit) activation function, which together constitute a contextual information acquisition unit.
[0059] In this embodiment, multiple target convolutional layers in the self-attention mechanism unit perform convolution operations on the context information respectively to obtain the output results of each target convolutional layer, and the feature association weight is calculated according to the output results of each target convolutional layer. Figure 3 As shown in the figure, the self-attention mechanism unit contains three 1×1 convolutional layers (also known as target convolutional layers). , , , the output features of the context convolutional neural network are convolved through the above three 1×1 convolutional layers to obtain the query Q, key K and value V, and then the self-attention mechanism unit output m is obtained through the softmax layer. The self-attention mechanism output m is cascaded with the standard convolutional layer output Conv(O) to obtain the final output y (target feature) of the feature extraction module. The self-attention mechanism unit output m is the feature association weight, and the standard convolutional layer output Conv(O) is the convolution operation result obtained by convolving the context information through the standard convolutional layer in the self-attention mechanism unit. The feature extraction module includes a context information acquisition unit and a self-attention mechanism unit. The self-attention mechanism unit includes multiple target convolutional layers and a standard convolutional layer.
[0060] ;
[0061] ;
[0062] in, is the transpose of K, is the number of columns of matrices Q and K, O is the output of the context information acquisition unit, Conv( ) is the convolution operation performed by the standard convolution layer, and concat( ) represents cascading. The self-attention mechanism unit is connected after the context convolutional neural network and together with the context convolution constitutes the feature extraction module.
[0063] It can be seen that by extracting the contextual information of two adjacent frames of images, the image information at different times and perspectives can be better fused. By obtaining the feature association weights of the contextual information, different feature information can be given different levels of importance, thereby accurately focusing on the extraction of key features. In this way, the accuracy of the monocular visual odometer method based on deep learning is improved.
[0064] Step S13: Analyze the temporal dimension information of the target features through the temporal learning module of the preset monocular visual odometer model to obtain the temporal features for camera pose estimation.
[0065] The temporal dimension information of the target features is analyzed in sequence through the deep convolution layer, deep expansion convolution layer, and target convolution layer in the temporal learning module, and the output result of the target convolution layer is cascaded with the target features to obtain the temporal features for camera pose estimation. After the feature extraction module performs feature extraction, the temporal learning module is used for sequence learning. Specifically, the temporal learning module contains a layer of deep convolution, a layer of deep expansion convolution, and a layer of 1×1 convolution. Its structural block diagram is shown below: Figure 4As shown in the figure, the target features are sequentially captured through deep convolution, deep dilated convolution, and 1×1 convolution to capture long-term dependencies for sequence learning. Then, the output of the 1×1 convolution layer is element-wise multiplied with the input of the temporal learning module to obtain the final output of the temporal learning module, i.e., the temporal features.
[0066] Step S14: perform three-dimensional translational degree of freedom estimation and rotational degree of freedom estimation on the temporal features through the posture network module of the monocular visual odometer model to obtain a first translation vector and a first rotation vector, and obtain a six-degree-of-freedom posture according to the first translation vector and the first rotation vector.
[0067] In this embodiment, three-dimensional translational degree of freedom and rotational degree of freedom are estimated for the temporal features through multiple fully connected layers in the posture network module to obtain a first translation vector and a first rotation vector, and a six-degree-of-freedom posture is obtained based on the cascade result of the first translation vector and the first rotation vector.
[0068] See also Figure 5 As shown, the algorithm framework of this embodiment includes the following parts: 1) using context convolution to extract features from the input image sequence and enhance the context information of feature extraction; 2) introducing an attention mechanism module in the neural network framework to capture the long-term dependency and correlation between features, and together with the context convolutional neural network, constitute the feature extraction module of the monocular visual mileage calculation method framework proposed in this application; 3) the effective features extracted by the feature extraction module composed of the above-mentioned context convolution and attention mechanism are sequenced through the temporal learning module to obtain temporal features; 4) finally, a pose network module is formed by a fully connected layer to perform three-dimensional translation degree of freedom estimation and three-dimensional rotation degree of freedom estimation, and the translation vector and the rotation vector are cascaded to finally obtain a six-degree-of-freedom pose.
[0069] This application applies deep learning technology to monocular visual odometry to solve the problems of high manual calculation cost and susceptibility of pose estimation accuracy to environmental conditions in traditional monocular visual odometry calculation methods. In order to solve the problem of low pose estimation accuracy in traditional end-to-end monocular visual odometry calculation methods based on deep learning, this application abandons the standard convolutional neural network, adopts context convolution for image feature extraction, introduces the attention mechanism, and finally adopts the temporal learning module to learn temporal correlation and sequence dependency. 1) Compared with the traditional monocular visual odometry calculation method, this application applies deep learning technology to monocular visual odometry, avoiding the problems of complex calculation and high manual cost in traditional monocular visual odometry calculation methods. In addition, the use of deep learning technology in this application also avoids the prediction failure that may occur in traditional monocular visual odometry calculation methods under some more extreme environmental conditions. 2) Compared with the traditional monocular visual mileage calculation method based on deep learning, this application abandons the standard convolution and adopts contextual convolution for feature extraction, improving the ability of neural networks to integrate contextual information. At the same time, an attention mechanism module is introduced to capture the long-term dependency and internal correlation learning between features, which together with the contextual convolutional neural network constitutes a feature extraction module, which enhances the effectiveness of feature extraction and reduces the error of pose estimation. In addition, this application uses a temporal learning module for sequence learning to further improve the accuracy of pose estimation. Compared with the traditional end-to-end monocular visual mileage calculation method based on deep learning, the accuracy of pose estimation of the algorithm proposed in this application has been significantly improved.
[0070] Taking into account the complexity and changeability of actual scenes, two adjacent frames of images may have significant differences in lighting conditions and other aspects. In order to cope with this situation, the present application can extract image features such as brightness, contrast, and color histogram for evaluation. If the brightness difference between the two frames of images is large, or the color histogram distribution is significantly different, it indicates that the lighting conditions have changed. Based on the analysis results of these image characteristics, the fusion ratio of the output results of different dilated convolution subunits is adjusted in real time. Due to the different expansion coefficients of different dilated convolution subunits, the contextual information extracted differs in scale. The subunit with a small expansion coefficient extracts local detail features, while the subunit with a large expansion coefficient extracts more macroscopic global features. When lighting conditions change, local detail features may be greatly affected, and the weight of the output results of the subunit with a small expansion coefficient needs to be reduced to reduce the interference of lighting changes on feature extraction.
[0071] It can be seen that the present application proposes a monocular visual odometer method based on deep learning, including: resizing the image set collected by the monocular camera, and obtaining two adjacent frames of images after adjustment; extracting the context information of the two adjacent frames of images through a feature extraction module of a preset monocular visual odometer model, and obtaining the feature association weight of the context information, and obtaining the target feature according to the feature association weight and the convolution operation result of the context information by the standard convolution layer; analyzing the temporal dimension information of the target feature through a temporal learning module of a preset monocular visual odometer model to obtain the temporal feature for camera pose estimation; performing three-dimensional translational degree of freedom estimation and rotational degree of freedom estimation on the temporal feature through the pose network module of the monocular visual odometer model to obtain a first translation vector and a first rotation vector, and obtaining a six-degree-of-freedom pose according to the first translation vector and the first rotation vector. It can be seen that in this application, by resizing the image set collected by the monocular camera and obtaining two adjacent frames of images, it is helpful to reduce the computational complexity caused by the difference in image size and improve the efficiency of algorithm processing. At the same time, the use of two adjacent frames of images can capture richer motion information. Compared with traditional methods, key information can be extracted more effectively, which improves the accuracy of pose estimation to a certain extent. Furthermore, by extracting the context information of two adjacent frames of images, the image information at different times and different perspectives can be better fused. By obtaining the feature association weights of the context information, different feature information can be given different degrees of importance, thereby accurately focusing on the extraction of key features. In addition, through the temporal learning module in the deep learning model, this application can better capture the time dependency in the image sequence, thereby effectively analyzing the temporal changes of the target features and providing more comprehensive and accurate temporal information for pose estimation.
[0072] Correspondingly, the present application also discloses a monocular visual odometer device based on deep learning, see Figure 6 As shown, the device comprises:
[0073] The image acquisition module 11 is used to resize the image set collected by the monocular camera and obtain two adjacent frames of images after the resizing;
[0074] A feature extraction module 12 is used to extract context information of two adjacent frames of images through a feature extraction module of a preset monocular visual odometer model, obtain feature association weights of the context information, and then obtain target features according to the feature association weights and the convolution operation results of the standard convolution layer on the context information;
[0075] A timing analysis module 13 is used to analyze the timing dimension information of the target features through a timing learning module of a preset monocular visual odometer model to obtain timing features for camera pose estimation;
[0076] The pose estimation module 14 is used to perform three-dimensional translational freedom estimation and rotational freedom estimation on the temporal features through the pose network module of the monocular visual odometer model, obtain a first translation vector and a first rotation vector, and obtain a six-degree-of-freedom pose according to the first translation vector and the first rotation vector.
[0077] Among them, for more specific working processes of the above-mentioned modules, please refer to the corresponding contents disclosed in the aforementioned embodiments, which will not be repeated here.
[0078] It can be seen that the present application proposes a monocular visual odometer method based on deep learning, including: resizing the image set collected by the monocular camera, and obtaining two adjacent frames of images after adjustment; extracting the context information of the two adjacent frames of images through a feature extraction module of a preset monocular visual odometer model, and obtaining the feature association weight of the context information, and obtaining the target feature according to the feature association weight and the convolution operation result of the context information by the standard convolution layer; analyzing the temporal dimension information of the target feature through a temporal learning module of a preset monocular visual odometer model to obtain the temporal feature for camera pose estimation; performing three-dimensional translational degree of freedom estimation and rotational degree of freedom estimation on the temporal feature through the pose network module of the monocular visual odometer model to obtain a first translation vector and a first rotation vector, and obtaining a six-degree-of-freedom pose according to the first translation vector and the first rotation vector. It can be seen that in this application, by resizing the image set collected by the monocular camera and obtaining two adjacent frames of images, it is helpful to reduce the computational complexity caused by the difference in image size and improve the efficiency of algorithm processing. At the same time, the use of two adjacent frames of images can capture richer motion information. Compared with traditional methods, key information can be extracted more effectively, which improves the accuracy of pose estimation to a certain extent. Furthermore, by extracting the context information of two adjacent frames of images, the image information at different times and different perspectives can be better fused. By obtaining the feature association weights of the context information, different feature information can be given different degrees of importance, thereby accurately focusing on the extraction of key features. In addition, through the temporal learning module in the deep learning model, this application can better capture the time dependency in the image sequence, thereby effectively analyzing the temporal changes of the target features and providing more comprehensive and accurate temporal information for pose estimation.
[0079] Furthermore, an embodiment of the present application also provides an electronic device. Figure 7 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content in the diagram cannot be regarded as any limitation on the scope of use of the present application.
[0080] Figure 7A schematic diagram of the structure of an electronic device 20 provided in an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a display screen 23, an input and output interface 24, a communication interface 25, a power supply 26, and a communication bus 27. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the monocular visual odometer method based on deep learning disclosed in any of the aforementioned embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0081] In this embodiment, the power supply 26 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 25 can create a data transmission channel between the electronic device 20 and the external device, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of the present application, and is not specifically limited here; the input and output interface 24 is used to obtain external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs and is not specifically limited here.
[0082] In addition, the memory 22, as a carrier for resource storage, may be a read-only memory, a random access memory, a disk or an optical disk, etc., and the resources stored thereon may include a computer program 221, and the storage method may be temporary storage or permanent storage. Among them, the computer program 221 includes not only a computer program that can be used to complete the monocular visual odometer method based on deep learning performed by the electronic device 20 disclosed in any of the aforementioned embodiments, but also a computer program that can be used to complete other specific tasks.
[0083] Furthermore, an embodiment of the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the aforementioned deep learning-based monocular visual odometer method disclosed above is implemented.
[0084] For the specific steps of the method, reference may be made to the corresponding contents disclosed in the aforementioned embodiments, which will not be described in detail here.
[0085] The various embodiments in this application are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description.
[0086] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0087] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0088] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0089] The above is a detailed introduction to the monocular visual odometer method, device, equipment, and storage medium based on deep learning provided by the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for general technical personnel in this field, according to the idea of the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A monocular visual odometer method based on deep learning, characterized in that: include: Resize the image set collected by the monocular camera and obtain two adjacent frames of the resized image; Extracting the context information of the two adjacent frames of images through a feature extraction module of a preset monocular visual odometer model, and obtaining a feature association weight of the context information, and then obtaining a target feature according to the feature association weight and a convolution operation result of a standard convolution layer on the context information; Analyzing the temporal dimension information of the target feature through the temporal learning module of the preset monocular visual odometer model to obtain the temporal features for camera pose estimation; The three-dimensional translational degree of freedom and rotational degree of freedom of the temporal feature are estimated through the posture network module of the monocular visual odometer model to obtain a first translation vector and a first rotation vector, and a six-degree-of-freedom posture is obtained according to the first translation vector and the first rotation vector.
2. The monocular visual odometer method based on deep learning according to claim 1, characterized in that: The feature extraction module includes a context information acquisition unit and a self-attention mechanism unit. The feature extraction module of the preset monocular visual odometer model extracts the context information of the two adjacent frames of images and obtains the feature association weight of the context information, including: The context information of the two adjacent frames of images is extracted by multiple cascaded dilated convolution subunits in the context information acquisition unit; wherein different dilated convolution subunits correspond to different dilation coefficients, and, in each of the dilated convolution subunits, the input data of the dilated convolution subunit is operated based on the convolution kernel size of the dilated convolution subunit and the dilation coefficient of the dilated convolution subunit to obtain the output result of each of the dilated convolution subunits, for the first dilated convolution subunit, the input data of the dilated convolution subunit is the two adjacent frames of images, and for the non-first dilated convolution subunit, the input data of the dilated convolution subunit is the output result of the previous dilated convolution subunit; The context information is convolved separately through multiple target convolutional layers in the self-attention mechanism unit to obtain the output result of each target convolutional layer, and the feature association weight is calculated according to the output result of each target convolutional layer.
3. The monocular visual odometer method based on deep learning according to claim 2, characterized in that: The step of obtaining the target feature according to the feature association weight and the convolution operation result of the standard convolution layer on the context information includes: Performing a convolution operation on the context information through the standard convolution layer in the self-attention mechanism unit to obtain the convolution operation result; The feature association weight and the convolution operation result are cascaded to obtain the target feature.
4. The monocular visual odometer method based on deep learning according to claim 2, characterized in that: The analyzing the temporal dimension information of the target feature through the temporal learning module of the preset monocular visual odometer model to obtain the temporal features for camera pose estimation includes: The temporal dimension information of the target feature is analyzed in sequence through the deep convolution layer, the deep expansion convolution layer, and the target convolution layer in the temporal learning module, and the output result of the target convolution layer is cascaded with the target feature to obtain the temporal feature for camera pose estimation.
5. The monocular visual odometer method based on deep learning according to claim 1, characterized in that: The method of performing three-dimensional translational degree of freedom estimation and rotational degree of freedom estimation on the time series feature through the posture network module of the monocular visual odometer model to obtain a first translation vector and a first rotation vector, and obtaining a six-degree-of-freedom posture according to the first translation vector and the first rotation vector, includes: The three-dimensional translational degree of freedom and rotational degree of freedom of the temporal features are estimated through multiple fully connected layers in the posture network module to obtain a first translation vector and a first rotation vector, and a six-degree-of-freedom posture is obtained based on the cascade result of the first translation vector and the first rotation vector.
6. The monocular visual odometer method based on deep learning according to any one of claims 1 to 5, characterized in that: Also includes: Inputting the sample image features into the model to be trained, so as to predict a second translation vector and a second rotation vector through the model to be trained; Calculating a first training loss according to the second translation vector and the actual translation vector, and calculating a second training loss according to the second rotation vector and the actual rotation vector; The total training loss is calculated according to the first training loss and the second training loss until the preset monocular visual odometer model that meets the loss condition is obtained.
7. The monocular visual odometer method based on deep learning according to claim 6, characterized in that: The calculating the first training loss according to the second translation vector and the actual translation vector, and calculating the second training loss according to the second rotation vector and the actual rotation vector, comprises: Calculating a first vector difference between the second translation vector and the actual translation vector, and calculating a first training loss according to an L2 norm of the first vector difference; A second vector difference between the second rotation vector and the actual rotation vector is calculated, and a second training loss is calculated according to a product of a target balance coefficient and an L2 norm of the second vector difference.
8. A monocular visual odometer device based on deep learning, characterized in that: include: An image acquisition module is used to resize the image set collected by the monocular camera and obtain two adjacent frames of the resized image; A feature extraction module is used to extract the context information of the two adjacent frames of images through a feature extraction module of a preset monocular visual odometer model, and obtain a feature association weight of the context information, and then obtain a target feature according to the feature association weight and a convolution operation result of a standard convolution layer on the context information; A timing analysis module, used for analyzing the timing dimension information of the target feature through the timing learning module of the preset monocular visual odometer model to obtain the timing features for camera pose estimation; A posture estimation module is used to perform three-dimensional translational degree of freedom estimation and rotational degree of freedom estimation on the time series features through the posture network module of the monocular visual odometer model, obtain a first translation vector and a first rotation vector, and obtain a six-degree-of-freedom posture according to the first translation vector and the first rotation vector.
9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the deep learning-based monocular visual odometer method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: Used to store a computer program; wherein, when the computer program is executed by a processor, the deep learning-based monocular visual odometer method according to any one of claims 1 to 7 is implemented.