Monocular vision odometer method, apparatus and device, and storage medium
By constructing the loss function penalty term for polar geometric constraints in monocular visual mileage calculation method, the problem of neglected correlation in translation and rotation prediction is solved, and a higher pose estimation accuracy is achieved.
Patent Information
- Application Number
- CN202510220130.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-06
AI Technical Summary
The existing monocular visual mileage calculation method based on deep learning ignores the relationship and constraints between the two in translation and rotation prediction, resulting in low accuracy of six-degree-of-freedom pose estimation.
By constructing a loss function penalty term based on an overpole geometric constraint, the original loss function is modified to associate the translation prediction with the rotation prediction and constrain each other, thereby improving the accuracy of the prediction.
The accuracy of translation and rotation prediction in monocular visual odometer is improved, thereby improving the accuracy of six-degree-of-freedom pose estimation.
Smart Images

Figure CN120107683A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning technology, and in particular to a monocular visual odometer method, device, equipment and storage medium. Background Art
[0002] The existing end-to-end monocular visual odometry calculation framework based on deep learning proposes to extract image features through convolutional neural networks (CNN), and use the extracted image features to learn temporal information through recurrent neural networks (RNN), and finally regress the camera's six-degree-of-freedom pose estimation. The loss function used in the neural network training process usually treats the translation prediction and rotation prediction as two independent parts, and calculates the loss function separately, and then obtains the loss function finally used by the neural network through weighted summation. However, in actual situations, there is a correlation between translation and rotation in visual odometry. The calculation of the loss function in the existing monocular visual odometry calculation method based on deep learning technology ignores the connection and constraints between the translation and rotation of the visual odometry.
[0003] It can be seen that how to improve the accuracy of six-degree-of-freedom pose estimation by taking into account both the translation prediction accuracy and the rotation prediction accuracy is a problem that technical personnel in this field need to solve. Summary of the invention
[0004] The purpose of the embodiments of the present invention is to provide a monocular visual odometer method, device, equipment and storage medium, which can simultaneously improve the translation prediction accuracy and rotation prediction accuracy in the monocular visual odometer. The specific scheme is as follows:
[0005] In a first aspect, the present application discloses a monocular visual odometer method, comprising:
[0006] Performing time series feature extraction processing on image data collected by a monocular image acquisition device to obtain a target time series feature vector;
[0007] The objective loss function penalty term is constructed by using the epipolar geometry constraint, and the original loss function in the initial monocular visual odometry method is modified based on the objective loss function penalty term to obtain the monocular visual odometry method to be trained.
[0008] The target time series feature vector is input into the monocular visual odometry method to be trained for model training to obtain the target monocular visual odometry method, so as to predict the position and posture of the target object based on the target monocular visual odometry method.
[0009] Optionally, performing time series feature extraction processing on image data collected by a monocular image acquisition device to obtain a target time series feature vector includes:
[0010] Resizing and extracting features of image data collected by a monocular image acquisition device to obtain an image feature vector;
[0011] The image feature vector is input into a preset feature encoder for sequence learning to obtain the target time series feature vector.
[0012] Optionally, resizing and feature extraction are performed on image data collected by a monocular image acquisition device to obtain an image feature vector, including:
[0013] Using a monocular image acquisition device to acquire images to obtain an original image data set, and cropping the monocular image in the original image data set based on a preset image size to obtain an image sequence consisting of single-frame images of the same size;
[0014] The single-frame images of two adjacent frames in the image sequence are input into a convolutional neural network composed of a preset number of standard convolutional layers and a preset activation function to obtain the corresponding image feature vector.
[0015] Optionally, an objective loss function penalty term is constructed using the epipolar geometry constraint, and the original loss function in the initial monocular visual odometry method is modified based on the objective loss function penalty term to obtain the monocular visual odometry method to be trained, including:
[0016] Obtaining a first coordinate and a second coordinate of a target pixel on a normalized plane corresponding to two adjacent frames of images, and constructing a relative position relationship of the target pixel based on the first coordinate and the second coordinate;
[0017] Determine the predicted translation vector and predicted rotation matrix corresponding to the target pixel point in two adjacent frames based on the relative position relationship of the motion;
[0018] Determine the actual translation vector and actual rotation matrix corresponding to the target pixel point in two adjacent frames of images;
[0019] According to the actual translation vector, the actual rotation matrix, the predicted translation vector and the predicted rotation matrix, an epipolar geometric constraint associating the translation information and the rotation information of the target pixel is constructed;
[0020] Construct the penalty term of the objective loss function based on the preset penalty factor and epipolar geometry constraints;
[0021] According to the preset hyperparameters, scale parameters and penalty terms of the target loss function, the original adaptive loss function in the initial monocular visual odometry method is modified to obtain the monocular visual odometry method to be trained.
[0022] Optionally, the original adaptive loss function in the initial monocular visual odometry method is modified according to preset hyperparameters, scale parameters, and target loss function penalty terms to obtain the monocular visual odometry method to be trained, including:
[0023] A target loss function is generated according to preset hyperparameters, scale parameters and target loss function penalty terms, and the target loss function is used to replace the original adaptive loss function in the initial monocular visual odometry calculation method to obtain the monocular visual odometry calculation method to be trained.
[0024] Optionally, the target time series feature vector is input into the monocular visual odometry method to be trained for model training to obtain the target monocular visual odometry method, including:
[0025] Input the target time series feature vector into the fully connected layer of the monocular visual odometry method to be trained for training;
[0026] During the training process, the loss value of the monocular visual odometry method to be trained is determined based on the target loss function in the monocular visual odometry method to be trained, and the monocular visual odometry method to be trained is adjusted based on the loss value to obtain the target monocular visual odometry method.
[0027] Optionally, determining a loss value of the monocular visual odometry method to be trained based on a target loss function in the monocular visual odometry method to be trained, and adjusting the monocular visual odometry method to be trained based on the loss value to obtain a target monocular visual odometry method, including:
[0028] Obtain the six-degree-of-freedom pose prediction value of the monocular visual odometry method to be trained;
[0029] The loss value of the monocular visual odometry method to be trained is determined using the target loss function and the six-degree-of-freedom pose prediction value;
[0030] The monocular visual odometry method to be trained is optimized using a preset optimizer based on the loss value, and the monocular visual odometry method to be trained is iteratively trained for a preset number of rounds according to a learning rate decay strategy to obtain a target monocular visual odometry method.
[0031] In a second aspect, the present application discloses a monocular visual odometer device, comprising:
[0032] A feature extraction module is used to perform time series feature extraction processing on the image data collected by the monocular image acquisition device to obtain a target time series feature vector;
[0033] A loss function determination module is used to construct a target loss function penalty term using epipolar geometry constraints, and to modify the original loss function in the initial monocular visual odometry method based on the target loss function penalty term to obtain the monocular visual odometry method to be trained;
[0034] The model training module is used to input the target time series feature vector into the monocular visual odometry method to be trained for model training to obtain the target monocular visual odometry method, so as to predict the position and posture of the target object based on the target monocular visual odometry method.
[0035] In a third aspect, the present application discloses an electronic device, comprising:
[0036] Memory, used to store computer programs;
[0037] A processor is used to execute a computer program to implement the aforementioned monocular visual odometer method.
[0038] In a fourth aspect, the present application discloses a computer-readable storage medium for storing a computer program, which implements the aforementioned monocular visual odometer method when executed by a processor.
[0039] As can be seen from the above, the present invention performs time series feature extraction processing on image data collected by a monocular image acquisition device to obtain a target time series feature vector; constructs a target loss function penalty term using an epipolar geometry constraint, and modifies the original loss function in the initial monocular visual mileage calculation method based on the target loss function penalty term to obtain a monocular visual mileage calculation method to be trained; inputs the target time series feature vector into the monocular visual mileage calculation method to be trained for model training to obtain a target monocular visual mileage calculation method, so as to predict the position and posture of the target object based on the target monocular visual mileage calculation method.
[0040] It can be seen from the above technical scheme that the present invention is based on the epipolar geometry constraint between translation and rotation in the visual odometry, and designs a loss function penalty term suitable for the end-to-end monocular visual odometry calculation method based on deep learning. The translation prediction and the rotation prediction are linked, and they are constrained to each other to improve the translation prediction accuracy and the rotation prediction accuracy, thereby improving the accuracy of the six-degree-of-freedom pose estimation of the visual odometry. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0042] Figure 1A flow chart of a monocular visual odometer method provided by an embodiment of the present invention;
[0043] Figure 2 A schematic diagram of a convolutional neural network provided by an embodiment of the present invention;
[0044] Figure 3 A schematic diagram of the structure of a monocular visual mileage calculation method provided by an embodiment of the present invention;
[0045] Figure 4 A specific flow chart of a monocular visual odometer method provided by an embodiment of the present invention;
[0046] Figure 5 A schematic diagram of the structure of a monocular visual odometer method device disclosed in the present invention;
[0047] Figure 6 The present invention is a structural diagram of an electronic device disclosed in the present invention. DETAILED DESCRIPTION
[0048] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0049] The terms "including" and "having" in the specification of the present invention and the above-mentioned drawings, as well as any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device including a series of steps or units is not limited to the listed steps or units, but may include steps or units that are not listed.
[0050] In order to enable those skilled in the art to better understand the scheme of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0051] In the existing monocular visual odometer calculation method based on deep learning technology, the loss function usually regards the translation prediction and rotation prediction as two independent parts, and calculates the loss function separately, and then obtains the loss function finally used by the neural network through weighted summation. In actual situations, there is a correlation between translation and rotation in the visual odometer. The calculation of the loss function in the existing monocular visual odometer calculation method based on deep learning technology ignores the connection and constraints between the translation and rotation of the visual odometer, resulting in that the deep learning-based algorithm generally does not take into account the translation prediction accuracy and rotation prediction accuracy well, thereby reducing the accuracy of six-degree-of-freedom pose estimation. Therefore, this application will specifically introduce a monocular visual odometer method that can solve the above problems.
[0052] See also Figure 1 As shown, the embodiment of the present application discloses a monocular visual odometer method, including:
[0053] Step S11: performing time series feature extraction processing on the image data collected by the monocular image acquisition device to obtain a target time series feature vector.
[0054] In this embodiment, it should be noted that visual odometry (VO) is a key link in vision-based simultaneous localization and mapping (SLAM) technology, and plays a vital role in technologies such as mobile robots, autonomous driving, autonomous positioning and navigation. In recent years, deep learning has made outstanding achievements in the field of computer vision, and the visual mileage calculation method based on deep learning has also made significant progress. The image data collected by the monocular image acquisition device is subjected to time series feature extraction processing to obtain the target time series feature vector, including: resizing and extracting features of the image data collected by the monocular image acquisition device to obtain the image feature vector; the image feature vector is input into a preset feature encoder for sequence learning to obtain the target time series feature vector. That is, firstly, a series of pictures are collected by the monocular image acquisition device, and then the image data is adjusted in image size, and features are extracted from the adjusted image.
[0055] Specifically, the image data collected by the monocular image acquisition device is resized and feature extracted to obtain an image feature vector, including: using the monocular image acquisition device to collect images to obtain an original image data set, and cropping the monocular image in the original image data set based on a preset image size to obtain an image sequence composed of single-frame images of the same size; inputting the single-frame images of two adjacent frames in the image sequence into a convolutional neural network composed of a preset number of standard convolutional layers and a preset activation function to obtain a corresponding image feature vector. Among them, firstly, the original monocular RGB (Red, Green, Blue, i.e., three primary colors) image collected by the monocular image acquisition device is cropped, and the specific cropping size can be adjusted accordingly according to the actual requirements of subsequent model training. By cropping the monocular image, the image size of each monocular image in the original image data set is adjusted to a uniform size; in this way, it is convenient to perform image feature extraction later. Then, the images of two adjacent frames in the image sequence after resizing, i.e., the i-th frame image and the i+1-th frame image, are input into a preset convolutional neural network for feature extraction to obtain a corresponding image feature vector. It should be noted here that convolutional neural networks such as Figure 2 As shown in the figure, it includes 9 standard convolutional layers. Specifically, the first convolutional layer has a convolution kernel size of 7×7 and 64 channels, the second convolutional layer has a convolution kernel size of 5×5 and 128 channels, the third convolutional layer has a convolution kernel size of 5×5 and 256 channels, the fourth convolutional layer has a convolution kernel size of 3×3 and 256 channels, the fifth, sixth, seventh, and eighth convolutional layers have a convolution kernel size of 3×3 and 512 channels, and the ninth convolutional layer has a convolution kernel size of 3×3 and 1024 channels. A ReLU activation function (Rectified Linear Unit) is connected after each convolutional layer.
[0056] Step S12: constructing a target loss function penalty term using the epipolar geometry constraint, and modifying the original loss function in the initial monocular visual odometry method based on the target loss function penalty term to obtain the monocular visual odometry method to be trained.
[0057] In this embodiment, it is first necessary to consider that in this actual operation, there is a correlation between the translation and rotation in the visual odometer. However, the calculation of the loss function in the current monocular visual odometer calculation method based on deep learning technology ignores the connection and constraints between the translation and rotation of the visual odometer. Therefore, this embodiment considers constructing a certain loss function penalty term based on the epipolar geometry constraint between the translation and rotation in the visual odometer. That is, the epipolar geometry constraint is used to construct a target loss function penalty term, and the original loss function in the initial monocular visual mileage calculation method is modified based on the target loss function penalty term to obtain the monocular visual mileage calculation method to be trained, including: obtaining the first coordinate and the second coordinate of the target pixel point on the normalized plane corresponding to two adjacent frames of images, and constructing the motion relative pose relationship of the target pixel point based on the first coordinate and the second coordinate; determining the predicted translation vector and the predicted rotation matrix corresponding to the target pixel point on the two adjacent frames of images based on the motion relative pose relationship; determining the actual translation vector and the actual rotation matrix corresponding to the target pixel point on the two adjacent frames of images; constructing the epipolar geometry constraint associated with the translation information and the rotation information of the target pixel point according to the actual translation vector, the actual rotation matrix, the predicted translation vector and the predicted rotation matrix; constructing the target loss function penalty term based on the preset penalty factor and the epipolar geometry constraint; modifying the original adaptive loss function in the initial monocular visual mileage calculation method according to the preset hyperparameter, scale parameter and the target loss function penalty term to obtain the monocular visual mileage calculation method to be trained. The specific calculation process is as follows: First, the epipolar geometry constraint estimates the relative pose relationship of the monocular camera frame-to-frame motion based on the two-dimensional plane information of the image. The formula is as follows:
[0058] ;
[0059] in, , are the coordinates of the two pixels obtained in the two adjacent frames on the normalized plane, t is the translation vector, R is the rotation matrix, Denotes the outer product with t. Based on the above epipolar geometry constraints, the constraint relationship that associates the translation prediction with the rotation prediction can be derived as follows:
[0060] ;
[0061] in, and are the actual translation vector and rotation matrix respectively, and are the predicted translation vector and rotation matrix respectively. Further, according to the optimization theory, the formula of the penalty term C of the loss function is constructed as follows:
[0062] ;
[0063] in, Is the penalty factor. After obtaining the current loss function penalty term, the original adaptive loss function in the initial monocular visual mileage calculation method is modified to obtain the monocular visual mileage calculation method to be trained. It should be noted here that in this embodiment, according to the preset hyperparameters, scale parameters and target loss function penalty terms, the original adaptive loss function in the initial monocular visual mileage calculation method is modified to obtain the monocular visual mileage calculation method to be trained, including: generating a target loss function according to the preset hyperparameters, scale parameters and target loss function penalty terms, and replacing the original adaptive loss function in the initial monocular visual mileage calculation method with the target loss function to obtain the monocular visual mileage calculation method to be trained. That is, the target loss function is generated by preset hyperparameters, scale parameters and target loss function penalty terms, wherein the specific target loss function Loss is as follows:
[0064] ;
[0065] in, is a hyperparameter used to control the robustness of the adaptive loss function. is the scale parameter, which scales the value of the loss function. is the prediction error, that is, the difference between the predicted six-degree-of-freedom pose and the actual six-degree-of-freedom pose. After obtaining the target loss function, the target loss function can be used to replace the original adaptive loss function in the initial monocular visual mileage calculation method to obtain the monocular visual mileage calculation method to be trained.
[0066] It should be noted here that in actual operation, the order of step S11 and step S12 can be adjusted. That is, the original loss function in the initial monocular visual mileage calculation method can be modified first, and after obtaining the monocular visual mileage calculation method to be trained, the convolutional neural network in the monocular visual mileage calculation method to be trained is used to extract features from the image data to obtain the target time series feature vector. In this way, the entire image processing and model training are completed based on the model of the monocular visual mileage calculation method to be trained, which can save the number of models used.
[0067] Step S13: inputting the target time series feature vector into the monocular visual odometry method to be trained for model training to obtain the target monocular visual odometry method, so as to predict the position and posture of the target object based on the target monocular visual odometry method.
[0068] In this embodiment, the target time series feature vector is input into the monocular visual mileage calculation method to be trained for model training to obtain the target monocular visual mileage calculation method, including: inputting the target time series feature vector into the fully connected layer of the monocular visual mileage calculation method to be trained for training; during the training process, the loss value of the monocular visual mileage calculation method to be trained is determined based on the target loss function in the monocular visual mileage calculation method to be trained, and the monocular visual mileage calculation method to be trained is adjusted based on the loss value to obtain the target monocular visual mileage calculation method. It should be noted here that the fully connected network is composed of 4 layers of fully connected layers. Specifically, the number of nodes in the first two layers of fully connected layers are 256 and 128 respectively, and the number of nodes in the last two layers of fully connected layers are both 3. The outputs of the last two fully connected layers are cascaded to finally obtain the six-degree-of-freedom posture prediction value. Wherein, the loss value of the monocular visual odometry method to be trained is determined based on the target loss function in the monocular visual odometry method to be trained, and the monocular visual odometry method to be trained is adjusted based on the loss value to obtain the target monocular visual odometry method, including: obtaining the six-degree-of-freedom pose prediction value of the monocular visual odometry method to be trained; using the target loss function and the six-degree-of-freedom pose prediction value to determine the loss value of the monocular visual odometry method to be trained; optimizing the monocular visual odometry method to be trained based on the loss value using a preset optimizer, and performing a preset round of iterative training on the monocular visual odometry method to be trained according to a learning rate decay strategy to obtain the target monocular visual odometry method. That is, firstly, the predicted six-degree-of-freedom pose and the actual six-degree-of-freedom pose are compared, and the neural network is trained using an AdamW optimizer (Adam with Weight Decay, an optimizer with weight decay function) through the loss function according to deep learning theory, and the learning rate decay strategy is used, and the training of the neural network is completed after a total of preset rounds of iterations. In actual operation, the preset rounds can be set to 200 times. In actual operation, the conditions for stopping model training can be set based on the actual situation of model training. For example, by setting a threshold, when the loss function is less than the threshold in multiple consecutive training rounds, the default model has achieved the best performance, and training can be stopped directly at this time; or by introducing a regularization term in the modified loss function, during the training process, when the influence of the regularization term on the loss function gradually increases, the default model has achieved the best performance, and training can be stopped at this time; in this way, by setting a variety of conditions for stopping model training, model training can be carried out based on actual conditions to avoid the situation where the accuracy of six-degree-of-freedom pose estimation is insufficient due to overfitting or insufficient training of the model training.
[0069] In summary Figure 3As shown, the present invention adopts a deep learning end-to-end monocular visual mileage calculation method based on an improved loss function. Specifically, first, two adjacent frames of images are used as input, and a convolutional neural network is used to extract features from the input image sequence; then the features extracted by the convolutional neural network are input into the Transformer encoder to extract timing information; finally, through four fully connected layers, three-dimensional translational degrees of freedom prediction and three-dimensional rotational degrees of freedom prediction are performed, and the translation prediction and rotation prediction are cascaded to finally obtain a six-degree-of-freedom posture; it should be noted that the loss function uses an adaptive loss function, and a loss function penalty term based on epipolar geometry constraints is added to it to improve the accuracy of six-degree-of-freedom posture estimation.
[0070] It can be seen that in this embodiment, Figure 4 As shown, the image data collected by the monocular image acquisition device is processed by time series feature extraction to obtain a target time series feature vector; the target loss function penalty term is constructed by using the epipolar geometry constraint, and the original loss function in the initial monocular visual mileage calculation method is modified based on the target loss function penalty term to obtain the monocular visual mileage calculation method to be trained; the target time series feature vector is input into the monocular visual mileage calculation method to be trained for model training to obtain the target monocular visual mileage calculation method, so as to predict the position and posture of the target object based on the target monocular visual mileage calculation method.
[0071] It can be seen from the above technical scheme that the present invention is based on the epipolar geometry constraint between translation and rotation in the visual odometry, and designs a loss function penalty term suitable for the end-to-end monocular visual odometry calculation method based on deep learning. The translation prediction and the rotation prediction are linked, and they are constrained to each other to improve the translation prediction accuracy and the rotation prediction accuracy, thereby improving the accuracy of the six-degree-of-freedom pose estimation of the visual odometry.
[0072] refer to Figure 5 The present application also discloses a monocular visual odometer device, including:
[0073] The feature extraction module 11 is used to perform time series feature extraction processing on the image data collected by the monocular image acquisition device to obtain a target time series feature vector;
[0074] A loss function determination module 12 is used to construct a target loss function penalty term using the epipolar geometry constraint, and to modify the original loss function in the initial monocular visual odometry method based on the target loss function penalty term to obtain the monocular visual odometry method to be trained;
[0075] The model training module 13 is used to input the target time series feature vector into the monocular visual odometry method to be trained for model training to obtain a target monocular visual odometry method, so as to predict the position and posture of the target object based on the target monocular visual odometry method.
[0076] It can be seen that in this embodiment, based on the epipolar geometry constraint between translation and rotation in the visual odometry, a loss function penalty term suitable for the end-to-end monocular visual odometry calculation method based on deep learning is designed, which links the translation prediction with the rotation prediction and constrains each other to improve the translation prediction accuracy and the rotation prediction accuracy, thereby improving the accuracy of the six-degree-of-freedom pose estimation of the visual odometry.
[0077] Furthermore, the present application also discloses an electronic device. Figure 6 It is a structural diagram of an electronic device according to an exemplary embodiment, and the content in the figure cannot be regarded as any limitation on the scope of use of this application. The electronic device may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input and output interface 25 and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the monocular visual odometer method disclosed in any of the aforementioned embodiments. In addition, the electronic device in this embodiment may specifically be an electronic computer.
[0078] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device; the communication interface 24 can create a data transmission channel between the electronic device and the external device, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of the present application, and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs and is not specifically limited here.
[0079] In addition, the memory 22, as a carrier for storing resources, can be a read-only memory, a random access memory, a disk or an optical disk, etc. The resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.
[0080] The operating system 221 is used to manage and control various hardware devices on the electronic device and the computer program 222, which can be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program that can be used to complete the monocular visual odometer method performed by the electronic device disclosed in any of the aforementioned embodiments, the computer program 222 can further include a computer program that can be used to complete other specific tasks.
[0081] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the monocular visual odometer method disclosed above is implemented. For the specific steps of the method, reference can be made to the corresponding contents disclosed in the above embodiments, and no further description will be given here.
[0082] Furthermore, the present application also discloses a computer program product, including a computer program / instruction; wherein the computer program / instruction, when executed by a processor, implements the aforementioned disclosed alarm aggregation method. For the specific steps of the method, reference may be made to the corresponding contents disclosed in the aforementioned embodiments, and no further description will be given here.
[0083] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0084] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0085] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0086] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0087] The technical solution provided by the present application is introduced in detail above. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for general technicians in this field, according to the idea of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A monocular visual odometer method, characterized in that: include: Performing time series feature extraction processing on image data collected by a monocular image acquisition device to obtain a target time series feature vector; The objective loss function penalty term is constructed by using the epipolar geometry constraint, and the original loss function in the initial monocular visual odometry method is modified based on the objective loss function penalty term to obtain the monocular visual odometry method to be trained; The target time series feature vector is input into the monocular visual odometry method to be trained for model training to obtain a target monocular visual odometry method, so as to predict the position and posture of the target object based on the target monocular visual odometry method.
2. The monocular visual odometer method according to claim 1, characterized in that: The step of performing time series feature extraction processing on the image data collected by the monocular image acquisition device to obtain a target time series feature vector includes: Resizing and extracting features of image data collected by a monocular image acquisition device to obtain an image feature vector; The image feature vector is input into a preset feature encoder for sequence learning to obtain a target time series feature vector.
3. The monocular visual odometer method according to claim 2, characterized in that: The step of resizing and extracting features from the image data collected by the monocular image acquisition device to obtain an image feature vector includes: Using a monocular image acquisition device to acquire images to obtain an original image data set, and cropping the monocular image in the original image data set based on a preset image size to obtain an image sequence consisting of single-frame images of the same size; The single-frame images of two adjacent frames in the image sequence are input into a convolutional neural network composed of a preset number of standard convolutional layers and a preset activation function to obtain a corresponding image feature vector.
4. The monocular visual odometer method according to claim 1, characterized in that: The method of constructing a target loss function penalty term by using the epipolar geometry constraint, and modifying the original loss function in the initial monocular visual mileage calculation method based on the target loss function penalty term to obtain the monocular visual mileage calculation method to be trained, includes: Acquire a first coordinate and a second coordinate of a target pixel on a normalized plane corresponding to two adjacent frames of images, and construct a relative position relationship of the target pixel based on the first coordinate and the second coordinate; Determine the predicted translation vector and the predicted rotation matrix corresponding to the target pixel point in the two adjacent frames of images based on the relative position relationship of the motion; Determine the actual translation vector and the actual rotation matrix corresponding to the target pixel point in the two adjacent frames of images; Constructing an epipolar geometric constraint associating translation information and rotation information of the target pixel point according to the actual translation vector, the actual rotation matrix, the predicted translation vector and the predicted rotation matrix; Constructing a penalty term of the target loss function based on a preset penalty factor and the epipolar geometry constraint; According to the preset hyperparameters, scale parameters and the penalty term of the target loss function, the original adaptive loss function in the initial monocular visual mileage calculation method is modified to obtain the monocular visual mileage calculation method to be trained.
5. The monocular visual odometer method according to claim 4, characterized in that: The method of modifying the original adaptive loss function in the initial monocular visual mileage calculation method according to the preset hyperparameters, scale parameters and the penalty term of the target loss function to obtain the monocular visual mileage calculation method to be trained includes: A target loss function is generated according to preset hyperparameters, scale parameters and the target loss function penalty term, and the target loss function is used to replace the original adaptive loss function in the initial monocular visual mileage calculation method to obtain the monocular visual mileage calculation method to be trained.
6. The monocular visual odometer method according to claim 4 or 5, characterized in that: The step of inputting the target time series feature vector into the monocular visual odometry method to be trained for model training to obtain the target monocular visual odometry method comprises: Inputting the target time series feature vector into the fully connected layer of the monocular visual odometry method to be trained for training; During the training process, a loss value of the monocular visual odometry method to be trained is determined based on the target loss function in the monocular visual odometry method to be trained, and the monocular visual odometry method to be trained is adjusted based on the loss value to obtain a target monocular visual odometry method.
7. The monocular visual odometer method according to claim 6, characterized in that: The method of determining a loss value of the monocular visual odometry method to be trained based on the target loss function in the monocular visual odometry method to be trained, and adjusting the monocular visual odometry method to be trained based on the loss value to obtain a target monocular visual odometry method, comprises: Obtaining a six-degree-of-freedom pose prediction value of the monocular visual odometry method to be trained; Determine the loss value of the monocular visual odometry method to be trained by using the target loss function and the six-degree-of-freedom pose prediction value; The monocular visual odometry method to be trained is optimized using a preset optimizer based on the loss value, and the monocular visual odometry method to be trained is iteratively trained for a preset number of rounds according to a learning rate decay strategy to obtain a target monocular visual odometry method.
8. A monocular visual odometer device, characterized in that: include: A feature extraction module is used to perform time series feature extraction processing on the image data collected by the monocular image acquisition device to obtain a target time series feature vector; A loss function determination module, used to construct a target loss function penalty term using epipolar geometry constraints, and modify the original loss function in the initial monocular visual odometry method based on the target loss function penalty term to obtain the monocular visual odometry method to be trained; A model training module is used to input the target time series feature vector into the monocular visual odometry method to be trained for model training to obtain a target monocular visual odometry method, so as to predict the position and posture of the target object based on the target monocular visual odometry method.
9. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to execute the computer program to implement the steps of the monocular visual odometer method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the monocular visual odometer method according to any one of claims 1 to 7 are implemented.