Short sequence image-based visual odometer method with local constraint loss function
By processing short sequence images through a local constrained loss function and an end-to-end convolutional-recurrent neural network, the problems of trajectory drift and insufficient robustness in visual odometry are solved, and high-precision pose estimation and real-time output are achieved.
Patent Information
- Application Number
- CN202510929864.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-10-10
AI Technical Summary
Existing visual odometry methods are prone to trajectory drift in long-distance motion estimation and are insufficiently robust to illumination changes and dynamic scenes. The loss function design fails to effectively constrain the global trajectory accuracy and the correlation between rotation prediction and translation prediction.
A visual odometry method based on short-sequence images is proposed with a local constraint loss function. The short-sequence images are processed through an end-to-end convolutional-recurrent neural network, a local constraint loss function is designed to reduce the trajectory accumulation error, and the geometric consistency of pose estimation is enhanced through a decoupled rotation-translation joint optimization mechanism.
It significantly reduces the cumulative error, improves the accuracy and robustness of pose estimation, can output high-precision relative pose in real time, resists interference from lighting changes and improves adaptability to dynamic scenes.
Smart Images

Figure CN120765747A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of visual odometry and relates to a visual odometry method based on short sequence images with a local constraint loss function. Background Art
[0002] Visual odometry, a core technology in fields such as autonomous driving, virtual reality, and augmented reality, estimates camera pose by analyzing the motion relationship between consecutive image frames, thereby constructing motion trajectories. Monocular visual odometry, requiring only a single camera, offers the advantages of low cost and wide applicability. With the development of deep learning, learning-based monocular visual odometry has become a research hotspot.
[0003] The current mainstream methods are mainly divided into two categories:
[0004] Adjacent frame image pair processing framework: This method only uses two adjacent frames to estimate the relative pose. Although it can achieve real-time calculation, it ignores the motion continuity between multiple frames in the sequence, resulting in a significant increase in the cumulative error of long time series trajectories.
[0005] Complete sequence image processing framework: Although this method can utilize long sequence information, as the sequence length increases, the motion information is diluted, the noise interference is aggravated, and the computational complexity increases exponentially, making it difficult to meet real-time requirements.
[0006] In addition, the loss function design of existing methods has obvious defects: only the relative pose error between adjacent frames is optimized, and the global trajectory accuracy is not constrained; the correlation between translation prediction and rotation prediction is not established, resulting in decoupled optimization of the two and reducing the geometric consistency of pose estimation.
[0007] These problems cause existing methods to be prone to trajectory drift in long-distance motion estimation and lack robustness to interference such as lighting changes and dynamic scenes. Summary of the Invention
[0008] In light of this, the present invention aims to provide a short-sequence image-based visual odometry method with a locally constrained loss function for end-to-end estimation of the six-degree-of-freedom pose of a monocular camera. This method employs a novel framework for processing short-sequence images to address the issues existing in the two aforementioned frameworks. Based on this framework, a locally constrained loss function is designed to reduce trajectory cumulative error and aid in recovering VO relative translation, thereby simultaneously reducing both absolute trajectory error and relative pose error. This method achieves high pose accuracy and significantly reduces cumulative error.
[0009] In order to achieve the above object, the present invention provides the following technical solutions:
[0010] A visual odometry method based on short sequence images with a local constraint loss function includes the following steps:
[0011] S1: input monocular RGB continuous image;
[0012] S2: Split the continuous image into short sequence images of length n, where n ≥ 2;
[0013] S3: Input the short sequence image into the end-to-end convolutional-recurrent neural network and output the pose change based on the first frame image of the short sequence as the base coordinate system, where:
[0014] The posture change includes a rotation component and translational components e is the Euler angle, t is the translation vector;
[0015] The CNN consists of Conv1, Conv2, Conv3, Conv3_1, Conv4, Conv4_1, Conv5, Conv5_1, and Conv6 layers. The Bi-GRU has a hidden layer dimension of 1024 and two layers. The regression network consists of three fully connected layers, with input and output scaling from 2048 to 256, 256 to 64, and 64 to 3. The number of parameters is in the tens of millions. Table 1 shows a comparison of the average inference time with other algorithms.
[0016] Table 1
[0017] Method Time(s / frame) DeepVO 0.052 SC-SfMLearner 0.031 FeatDepth 0.018 Proposed 0.069
[0018] S4: Perform relative translation recovery on the translation component of the network output:
[0019]
[0020] in, Represents the translation vector of the i-th frame image relative to the i-1-th frame image; Represents the rotation matrix from the 1st frame coordinate system to the i-1th frame coordinate system; Represents the translation vector of the i-th frame image in the first frame coordinate system; Represents the translation vector of the i-1th frame image in the 1st frame coordinate system;
[0021] S5: The restored posture Accumulate as absolute trajectory:
[0022]
[0023] Where T is the homogeneous transformation matrix including rotation and translation;
[0024] The method also includes a training phase and an inference phase.
[0025] Furthermore, the end-to-end convolutional-recurrent neural network includes:
[0026] a Convolutional Neural Network (CNN) is used to extract geometric features of short sequence images, and its output is connected to a Bidirectional Gated Recurrent Unit (Bi-GRU);
[0027] the Bi-GRU is used to model the sequence dependency of the feature sequence output by the CNN, and its output is connected to a decoupled regression network;
[0028] the decoupled regression network includes a rotation regression network and a translation regression network that are independent of each other but are connected in parallel, and respectively output rotation components and translation components;
[0029] wherein the output of the CNN is directly input to the Bi-GRU, and the output of the Bi-GRU is simultaneously input to the rotation regression network and the translation regression network.
[0030] Further, the CNN adopts a design without pooling layers, reduces the size of the feature map through a convolutional layer with a stride of 2, and contains a residual block whose operation satisfies:
[0031]
[0032] wherein O is the output feature map; C o , H o and W o are the channel number, height and width of the output feature map, respectively; F is the input feature map; C i , H i and W i are the channel number, height and width of the input feature map, respectively; is a convolution operator; W is a convolution kernel, and superscripts k, p and s represent kernel size, padding size and stride, respectively; is a concatenation operation.
[0033] Further, the operation of the Bi-GRU satisfies:
[0034] r t =σ(W r x t +U r h t-1 +b r )
[0035] z t =σ(W z x t +U z h t-1 +b z )
[0036]
[0037] y t =W o h t +b o
[0038] Among them, r t is the reset gate output; z t is the update gate output; is the candidate hidden state; h t is the current hidden state; y t is the network output; x t is the current input; h t-1 is the hidden state at the previous moment; W, U are weight matrices; b r 、b z 、b h and b o They correspond to the reset gate, update gate, candidate state, and output bias vector respectively; σ is the Sigmoid function; ⊙ is element-by-element multiplication.
[0039] Furthermore, in S2, the short sequence splitting adopts a partial overlapping update mechanism in the inference stage: the latest acquired frame of image is overlapped with the last n-1 frames of the previous short sequence to generate a new short sequence input network.
[0040] Furthermore, the label design used in the training phase is:
[0041]
[0042] in, is the Euler angle of the i-th frame image in the world coordinate system; is the translation vector of the i-th frame image in the world coordinate system; The rotation matrix from the first frame coordinate system to the world coordinate system.
[0043] Furthermore, the loss function used in the training phase is a local constraint loss function:
[0044] loss=r_loss+α·t_loss,α=100
[0045]
[0046] in: The rotation component predicted by the network; Rotate the component for the label; is the translation component predicted by the network; is the label translation component; MSE is the mean square error function, defined as α is the translation loss weighting factor.
[0047] Furthermore, the data enhancement strategy adopted in the training phase includes: sliding and intercepting short sequences of length n with a step size of 2, and generating a training set after disrupting the order.
[0048] A visual odometry system, used to implement the method, comprising:
[0049] Image acquisition module, used to obtain monocular RGB continuous images;
[0050] A short sequence generation module connected to the image acquisition module;
[0051] An end-to-end convolutional-recurrent neural network module, connected to the short sequence generation module;
[0052] The pose restoration module is connected to the end-to-end convolutional-recurrent neural network module.
[0053] Furthermore, the end-to-end convolutional-recurrent neural network module includes:
[0054] CNN submodule: includes a design without pooling layers, reduces the feature map size through convolutional layers with a stride of 2, and includes residual blocks;
[0055] Bi-GRU submodule: uses a bidirectional gated recurrent unit, whose operations include reset gate, update gate and candidate hidden state calculation;
[0056] Decoupled regression submodule: Contains a parallel rotation regression network and a translation regression network.
[0057] The beneficial effects of the present invention are:
[0058] (1) Through an innovative short sequence image processing framework, the spatiotemporal correlation of multiple frames within a sequence is effectively integrated, overcoming the problem of missing sequence information in the adjacent frame image pair method. The framework establishes local motion constraints based on the short sequence head as the reference coordinate system, enabling the network to capture motion continuity and suppress drift in long-distance motion estimation.
[0059] (2) Design a local constraint loss function to simultaneously optimize the overall trajectory accuracy and the relative pose accuracy of the latest frame in a short sequence window. This function establishes the geometric correlation between translation prediction and rotation prediction through a decoupled rotation-translation joint optimization mechanism, thereby enhancing the physical rationality of pose estimation.
[0060] (3) Based on the improved convolutional-recurrent neural network architecture:
[0061] The residual convolution module enhances the geometric feature extraction capability and resists the interference of illumination changes.
[0062] Bidirectional gated recurrent units deeply mine sequence dependencies and improve robustness in dynamic scenes
[0063] Decoupled regression network separates and learns rotation / translation feature space to reduce motion blur
[0064] (4) Adopting partial overlapping short sequence update mechanism:
[0065] Only the new frame needs to be overlapped with the tail of the historical short sequence
[0066] Avoiding the computational overhead of full sequence reconstruction
[0067] Ensure real-time output of each frame pose
[0068] (5) In the complete process from original image input to absolute trajectory output:
[0069] Unified data enhancement strategy improves model generalization
[0070] Adaptive loss function balances rotation / translation optimization weights
[0071] The pose restoration module realizes the coordinate system transformation
[0072] (6) Local label design makes short sequences independent of motion segments
[0073] Two-stage constraint loss to simultaneously optimize local and global accuracy
[0074] Data enhancement strategy triples the effective training sample size
[0075] Other advantages, objects, and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art upon examination of the following description or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below with reference to the accompanying drawings, in which:
[0077] Figure 1 It is the overall framework diagram of the present invention;
[0078] Figure 2 for the relationship between short sequences, entire sequences, and image pairs;
[0079] Figure 3 It is a network framework diagram of the present invention;
[0080] Figure 4 The process of generating real-time short sequence input for the inference stage;
[0081] Figure 5 It is a data enhancement method;
[0082] Figure 6 The absolute trajectory results of sequence 09 and sequence 10; Figure 6 (a) is the absolute trajectory result of sequence 09; Figure 6 (b) is the absolute trajectory result of sequence 10;
[0083] Figure 7 Compare the results of the rotation of sequence 09 and sequence 10 with the true value; Figure 7 (a) Comparison between the rotation result of sequence 09 and the true value; Figure 7 (b) Comparison between the rotation result of sequence 10 and the true value. DETAILED DESCRIPTION
[0084] The following describes the embodiments of the present invention by means of specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and the following embodiments and features in the embodiments can be combined with each other without conflict.
[0085] Among them, the accompanying drawings are only for illustrative purposes and represent only schematic diagrams rather than actual pictures, and should not be understood as limiting the present invention. In order to better illustrate the embodiments of the present invention, some parts of the accompanying drawings may be omitted, enlarged or reduced, and do not represent the dimensions of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions may be omitted in the accompanying drawings.
[0086] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "back", etc. indicating directions or positional relationships, they are based on the directions or positional relationships shown in the drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operate in a specific direction. Therefore, the terms describing the positional relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.
[0087] The overall framework of the present invention is as follows Figure 1The input of the framework is monocular RGB consecutive images, and the output is real-time relative pose. The framework is divided into a training part (red) and an inference part (blue). In the training part, the consecutive images are first split into short sequence images and data augmentation is performed to obtain a training set. At the same time, the original pose truth value of the consecutive images is transformed to obtain a data label. Then, the proposed loss function is used to guide the training process to obtain the target network. In the inference part, the latest collected image is first overlapped with the previous short sequence part to form a new short sequence. Then, the new short sequence is input into the trained network as real-time input. The network output obtains the short sequence pose. Finally, the short sequence pose is restored to obtain the real-time relative pose. The real-time relative pose is accumulated with the previous pose to obtain the absolute trajectory. The specific details of each part will be explained in the following chapters.
[0088] I. VO method based on short sequence images
[0089] The object processed by the method of the present application is a short sequence monocular RGB image between adjacent image pairs and the entire sequence image, and the relationship among the three is as shown in Figure 2 In the training process, the method of the present application does not learn the relative pose in the coordinate system of the previous frame of the current frame as in general VO, but learns the adjacent frame pose change in the coordinate system of the first frame of the short sequence. By placing the poses of all images in the short sequence in the same reference system, the short sequence can be regarded as a separate motion independent of the entire sequence, which strengthens the connection between the poses in the short sequence and enables the network to fully utilize the sequence information between the short sequence images. In order to achieve real-time VO, in the inference process, the latest collected image is overlapped with the previous short sequence part to form a new short sequence as the real-time input of the network. After the network outputs the short sequence pose, the relative translation is restored to restore the pose of the latest image to the real-time relative pose in the coordinate system of the previous frame.
[0090] Figure 2 In the figure, the short sequence is represented by a purple square, the entire sequence is represented by a red square, and the image pair is represented by a blue square.
[0091] II. End-to-end convolutional-recurrent neural network
[0092] Figure 3 The network framework proposed by the present application is divided into input (blue square), CNN (orange square), RNN (green square), and regression network (purple square).
[0093] In the context of the VO problem, convolutional neural networks (CNNs) need to learn geometric features rather than appearance features in image pairs. A modified FlowNet network that can learn geometric optical flow features is used as a CNN. Because pooling layers may lose spatial information in the feature domain, no pooling layers are used in CNN, and convolutional layers with a stride of 2 are used instead to reduce the size of the feature map. Residual blocks are used to replace the ordinary convolutional layers in the FlowNet network, deepening the network while expanding the receptive field. The skip connections in the residual blocks allow the input to be directly added to the output across one or more convolutional layers, enabling the network to utilize feature information at different levels and enhancing the CNN's ability to integrate global contextual information. The architectural design of the proposed CNN is summarized in Table 2.
[0094] Table 2
[0095]
[0096] Input feature map F into each layer to perform the following operations:
[0097]
[0098] Where O is the output feature map and F is the input feature map. is the convolution operator. represents the i-th convolution kernel with size k, padding p, and stride s. It is a cascade operation, which means a jump connection. After each layer, the number of channels, height, and width of the original feature map F are reduced by C i ×H i ×W i Change to C o ×H o ×W o .
[0099] The short sequence features obtained after CNN are further fed into RNN to model the dependencies in the sequence. RNN has various applications, and the input-output relationships include one-to-many, many-to-one, many-to-many, etc. Some applications can even predict future values or sequences based on historical input sequences. RNN adopts a many-to-many form, and both input and output contain sequence dimensions. Most existing methods use LSTM as a recurrent neural network, but GRU usually performs better when processing short sequences. In the method of the present invention, all image information in the short sequence can be utilized, which means that RNN can use this information not only in the forward direction but also in the reverse direction, so the bidirectional GRU (Bi-GRU) is finally selected as the RNN. Input sequence X = [x1, x2,…, x t ]Enter Bi-GRU and perform the following operations:
[0100]
[0101] Among them, W, U are weight matrices, b is the bias, σ is the sigmoid function, and ⊙ represents element-by-element multiplication. t , z t They are the reset gate and update gate of GRU respectively, h t They are candidate hidden states and current hidden states respectively. The final The forward and reverse hidden states are combined to obtain the final hidden state. The final output Y=[y1,y2,…,y t ] is also a sequence of length t.
[0102] The sequence features output by RNN then enter the regression network composed of multi-layer perceptrons for the final pose regression. The final pose regression part of most existing architectures puts rotation and translation in one network to directly regress the 6-DOF pose. However, it is believed that rotation and translation are two different movements and should have different feature spaces. Therefore, rotation and translation are decoupled into two networks for regression. The rotation network and translation network finally output 3-DOF rotation and 3-DOF translation respectively. The overall network framework is composed of Figure 3 shown.
[0103] 3. Label Design
[0104] The proposed method is a supervised learning VO method. However, due to its unique framework, the network cannot directly use the ground-truth pose values of the training set images as labels. The label design must conform to the framework based on short sequence images. During training, batch-sized short sequences of images are fed into the network as input. Each batch of short sequences can be regarded as an independent camera motion. The label of each time step should depend only on the current short sequence. Measuring this motion cannot be done in the original world coordinate system. A new local base coordinate system is required. Therefore, the first frame of the short sequence is used as the base coordinate system, and the pose changes of adjacent frames in this coordinate system are used as labels.
[0105] Assume that the short sequence contains n images (I1, I2, ..., I n) , the original absolute pose truth value corresponding to each image is:
[0106]
[0107] Where R represents the rotation matrix, t represents the translation vector, w represents the world coordinate system, and n is the camera coordinate system corresponding to the nth image in the short sequence. It refers to the rotation matrix that transforms the vector in coordinate system n to the coordinate system w. It refers to the translation vector from the origin of coordinate system w to the origin of coordinate system n.
[0108] In three-dimensional rigid body motion, the rotation matrix requires nine parameters to represent the rotation, which is redundant and not conducive to network learning. Euler angles are another way to represent three degrees of freedom rotation using three separate angles. Euler angles consist of yaw, pitch, and roll angles. The rotation matrix of these three angles can be expressed as follows:
[0109]
[0110] Where ψ, θ and φ are the yaw, pitch and roll angles respectively. z , R y and R x The basic rotation matrices for rotating the body around the Z, Y, and X axes are ψ, θ, and φ, respectively. The conversion relationship between the total rotation matrix R and the Euler angle can be obtained by the following formula:
[0111] R=R z (ψ)R y (θ)R x (φ) (7)
[0112] Based on the above formula, the rotation matrix Converting to Euler angles (ψ, θ, φ) as labels increases intuitiveness and interpretability, reduces parameters and redundant data, and is more conducive to network learning. The original absolute pose truth after conversion is:
[0113]
[0114] Where e represents the Euler angle, I x ,I y ,I z Represents the displacement along the X-axis, Y-axis and Z-axis of the rigid body respectively. According to the label design idea of this section, the rigid body coordinate transformation is used to convert the true value of the pose into the first frame image of the short sequence as the base coordinate system, and the pose changes of adjacent frames in this coordinate system are used as labels:
[0115]
[0116] The subscript 1 represents the camera coordinate system of the first frame image I1 of the short sequence, and n is the camera coordinate system corresponding to the nth image in the short sequence. It is the rotation matrix that transforms the vector in coordinate system 1 to the world coordinate system w. The calculation of rigid body coordinate transformation. The final short sequence image (I1, I2, ..., I n ) The corresponding label is set to The rotation part is the relative rotation of two adjacent frames (the difference in Euler angles), and the translation part is the relative displacement of adjacent frames with the first frame of the short sequence as the base coordinate system.
[0117] 4. Relative Translation Recovery and Loss Function
[0118] When reasoning with a trained network, the network will attempt to output quantities close to the label. In the method of the present invention, the output is the relative pose of the short sequence mentioned in Section C, with the first frame image of the short sequence as the base coordinate system. However, the standard relative pose of VO uses the coordinate system of the previous frame image of the current frame as the base coordinate system, and calculates the pose change of the current frame relative to the previous frame. For the rotation part of the network output, the relative pose is not affected by the base coordinate system, and for the translation part, the following coordinate transformation is used to restore the output translation to a general relative translation:
[0119]
[0120] Where, the superscript T represents the transpose of the rotation matrix, Euler angles Converted to, It is obtained by summing the relative postures of all outputs before the short sequence n-1 frames:
[0121]
[0122] After restoring the output pose to the standard relative pose, the absolute pose can be obtained by accumulating the relative poses:
[0123]
[0124] (1): Euler angles Converted to, The sum of all relative postures output before the short sequence n-1 frames (9) is obtained:
[0125] (2):
[0127] R=R z (ψ)R y (θ)R x (φ)
[0128]
[0129] The [ψ,θ,φ] is the Euler angle φ which can be obtained by sinφ. The other angles are similar. (3):
[0131]
[0132] In order to achieve real-time VO, that is, the camera can obtain the corresponding relative pose in real time every time it collects a new frame of image, in the inference stage, the latest collected frame of image is partially overlapped with the previous short sequence to form a new short sequence as the real-time input of the network, such as Figure 4 shown.
[0133] First frame initialization rules, such as the processing mechanism when the n-frame buffer is not full:
[0134] Taking a short sequence length of 6 frames as an example, the first frame at the beginning of the run is usually used as the world coordinate system W. After the first six frames are accumulated and enter the network, the network output is as shown in formula (9):
[0135]
[0136] The first frame of the first 6 frames, with the subscript "1", is also the world coordinate system W of the entire motion trajectory;
[0137] The output is rewritten as:
[0138]
[0139] The first 6-frame short sequence is not like the following short sequences, which only perform relative translation recovery on the last time step. Instead, all 5 time steps are used as output. The posture part of these 5 time steps is already the standard relative posture. For the position part, formula (10) also needs to be used to restore it to the standard position.
[0140] The problem of sequence truncation or jump is not considered, and the default input image sequence is the entire complete motion process.
[0141] α is a scaling factor used to balance the weights of translation and rotation errors, and is set to 100 because the translation change (in m) in inter-frame motion is usually two orders of magnitude larger than the rotation (in radians).
[0142] In this way, as new images are continuously collected, the generated new short sequences enter the network and output the sequence pose in real time. Then, the real-time relative pose can be obtained by performing relative translation recovery on the output of the latest time step.
[0143] Figure 4 This is the process of generating real-time short sequence input during the inference phase. The image with the red border is the most recently acquired frame, and the image within the orange frame is the previous short sequence. The newly acquired frame is partially overlapped with the previous short sequence to form a new short sequence, which serves as the network's real-time input.
[0144] The recovery process of the relative translation of the latest time step of the short sequence is as follows. Assume that the length of the short sequence is n, and the sum of the outputs of the rotation part of all time steps of the short sequence is:
[0145]
[0146] The output of the latest time step is Subtracting the rotation part of the latest time step from the result of (13) yields:
[0147]
[0148] According to formula (10), we can get Just the translation part of the latest time step They are used together to recover relative translation, which means that the relative translation of the latest time step is not only related to its own output, but also to the relative posture output of all time steps in the short sequence. Therefore, while paying attention to the relative posture of the last time step, it is also necessary to pay attention to the overall posture of the short sequence. To improve the relative posture accuracy of the framework output, it is necessary to reduce the overall error of the short sequence while reducing the posture error of the last time step.
[0149] Based on the analysis of the output pose in this section and the decoupling of rotation and translation into two networks for regression proposed in Section B, a new loss function is proposed to guide model training to improve pose accuracy. The proposed loss function focuses on the last time step while constraining the short sequence as a local VO, which is called the local constraint loss function. The local constraint loss function is divided into two parts: rotation and translation. The loss function for the rotation part is:
[0150] r_loss=r_sumloss+r_lastloss (15)
[0151]
[0152] Where MSE stands for mean squared error, n represents the length of the short sequence, r_sumloss is the mean squared error between the sum of all Euler angle predictions in the short sequence window and the true value, and r_lastloss is the mean squared error between the Euler angle prediction value at the last time step and the true value. For the translation part, the loss function is set similarly.
[0153] t_loss=t_sumloss+t_lastloss (18)
[0154]
[0155] The total loss is:
[0156] loss=r_loss+α·t_loss (21)
[0157] Where α is a weighting factor used to balance the rotation and translation components, and is empirically chosen to be 100. Where α is a scaling factor used to balance the weights of translation and rotation errors, and is set to 100 because the translation change (in meters) in inter-frame motion is usually two orders of magnitude larger than the rotation (in radians).
[0158] By setting the loss function in this way, the sumloss component enables the network to minimize the overall pose error of short sequence windows during model training, achieving an effect similar to local optimization and reducing the cumulative error of VO. The lastloss component causes the network to pay special attention to reducing the relative pose error of the last time step as the real-time output. The decoupled translation prediction and rotation prediction are linked through the relative translation recovery process in Equation (10). The total loss simultaneously reduces the errors of r_loss and t_loss, and the relative translation error is also reduced.
[0159] 5. Specific implementation details
[0160] The network is implemented based on the PyTorch framework. During the training phase, the batch size is set to 8, the training cycle is set to 150, the Adam optimizer is used for optimization, and the learning rate is set to 10 -4 , using dropout and early stopping techniques to prevent overfitting, the network is trained using the KITTI dataset left eye original RGB image and the labels made from the provided absolute pose truth. The Euler angles used are the internal rotation Euler angles in the YXZ order, the network input image size is 640×192, and the short sequence length is set to 6. Based on the short sequence VO method, an adaptive data augmentation strategy is proposed, such as Figure 5 As shown in Figure 1, all images in the training set are sorted in order. A short sequence window with a sequence length of 6 is used to intercept the entire sequence with a step size of two images. The intercepted short sequences are then shuffled to obtain the desired training set. Although some images are reused after data augmentation, the labeling process described in Section C uses the first frame of the short sequence as the base coordinate system. As long as the first frame image is different, even repeated images will have different pose labels. Furthermore, as long as the sequence images are not completely repeated, the sequence information learned by the network should be completely different. After data augmentation, the resulting short sequence training set is expanded to three times the size of the original dataset, enriching the training data and reducing the risk of network overfitting.
[0161] Figure 5 The proposed data augmentation method uses a short sequence window with a sequence length of 6 to capture every three images in the entire sequence image, and the size of the resulting short sequence training set is expanded by 3 times.
[0162] Experimental results:
[0163] Experiments were conducted on the public dataset KITTI, with sequences 00, 01, 02, 05, 07, and 08 selected as training sets, and sequences 09 and 10 as test sets. Figure 6 and Figure 7 As shown, red is the true value and blue is the result of this method.
[0164] Figure 6 is the absolute trajectory result of sequence 09 and sequence 10, Figure 6 (a) is the absolute trajectory result of sequence 09, the left part is the XZ direction projection, the middle part is the XY direction projection, and the right part is the ZY direction projection; Figure 6 (b) is the absolute trajectory result of sequence 10. The left part is the XZ direction projection, the middle part is the XY direction projection, and the right part is the ZY direction projection.
[0165] Figure 7 Compare the results of the rotation of sequence 09 and sequence 10 with the true value. Figure 7 (a) Comparison of the rotation results of sequence 09 with the true value. The upper part is the rotation along the X axis, the middle part is the rotation along the Y axis, and the lower part is the rotation along the Z axis. Figure 7 (b) shows the comparison between the rotation results of sequence 10 and the true value. The upper part is the rotation along the X axis, the middle part is the rotation along the Y axis, and the lower part is the rotation along the Z axis.
[0166] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions, which should all be included in the scope of the claims of the present invention.
Claims
1. A visual odometry method based on short image sequences with a local constraint loss function, characterized by: The following steps are involved: S1: input monocular RGB continuous image; S2: Split the continuous image into short sequence images of length n, where n ≥ 2; S3: Input the short sequence image into the end-to-end convolutional-recurrent neural network and output the pose change based on the first frame image of the short sequence as the base coordinate system, where: The posture change includes a rotation component and translational components i∈[2,n], e is the Euler angle, t is the translation vector; S4: Perform relative translation recovery on the translation component of the network output: in, Represents the translation vector of the i-th frame image relative to the i-1-th frame image; Represents the rotation matrix from the 1st frame coordinate system to the i-1th frame coordinate system; Represents the translation vector of the i-th frame image in the first frame coordinate system; Represents the translation vector of the i-1th frame image in the 1st frame coordinate system; S5: The restored posture Accumulate as absolute trajectory: Where T is the homogeneous transformation matrix including rotation and translation; The method also includes a training phase and an inference phase.
2. The visual odometry method based on short image sequences with a local constraint loss function according to claim 1, characterized in that: The end-to-end convolutional-recurrent neural network includes: Convolutional neural network (CNN) is used to extract geometric features of short sequence images, and its output is connected to a bidirectional gated recurrent unit; Bidirectional gated recurrent unit (Bi-GRU), used to model the sequence dependency of the feature sequence output by CNN, with its output connected to the decoupled regression network; A decoupled regression network, including a rotation regression network and a translation regression network that are independent but connected in parallel, outputting rotational components and translational components respectively; The output of the CNN is directly used as the input of the Bi-GRU, and the output of the Bi-GRU is simultaneously input into the rotation regression network and the translation regression network.
3. The visual odometry method based on short image sequences with a local constraint loss function according to claim 2, characterized in that: The CNN adopts a design without pooling layers, reduces the feature map size through convolutional layers with a stride of 2, and includes residual blocks. Its operation satisfies: Among them, O is the output feature map; C o 、H o and W o are the number of channels, height and width of the output feature map respectively; F is the input feature map; C i 、H i and W i are the number of channels, height, and width of the input feature map respectively; is the convolution operator; W is the convolution kernel, and the superscripts k, p, and s represent the kernel size, padding size, and stride, respectively; It is a cascade operation.
4. The visual odometry method based on short sequence images with a local constraint loss function according to claim 2, characterized in that: The operation of the Bi-GRU satisfies: r t =σ(W r x t +U r h t-1 +b r ) z t =σ(W z x t +U z h t-1 +b z ) Among them, r t is the reset gate output; z t is the update gate output; is the candidate hidden state; h t is the current hidden state; y t is the network output; x t is the current input; h t-1 is the hidden state at the previous moment; W, U are weight matrices; b r 、b z 、b h and b o They correspond to the reset gate, update gate, candidate state, and output bias vector respectively; σ is the Sigmoid function; ⊙ is element-by-element multiplication.
5. The visual odometry method based on short sequence images with a local constraint loss function according to claim 1, characterized in that: In S2, the short sequence splitting adopts a partial overlapping update mechanism in the inference stage: the latest acquired frame of image is overlapped with the last n-1 frames of the previous short sequence to generate a new short sequence input network.
6. The visual odometry method based on short sequence images with a local constraint loss function according to claim 1, characterized in that: The label design used in the training phase is: in, is the Euler angle of the i-th frame image in the world coordinate system; is the translation vector of the i-th frame image in the world coordinate system; The rotation matrix from the first frame coordinate system to the world coordinate system.
7. The visual odometry method based on short image sequences with a local constraint loss function according to claim 1, characterized in that: The loss function used in the training phase is the local constraint loss function: loss=r_loss+α·t_loss,α=100 in: The rotation component predicted by the network; Rotate the component for the label; is the translation component predicted by the network; is the label translation component; MSE is the mean square error function, defined as α is the translation loss weighting factor.
8. The visual odometry method based on short sequence images with a local constraint loss function according to claim 1, characterized in that: The data enhancement strategy adopted in the training phase includes: sliding and intercepting short sequences of length n with a step size of 2, and generating a training set after disrupting the order.
9. A visual odometry system, configured to implement the method of any one of claims 1 to 8, characterized in that: include: Image acquisition module, used to obtain monocular RGB continuous images; A short sequence generation module connected to the image acquisition module; An end-to-end convolutional-recurrent neural network module, connected to the short sequence generation module; The pose restoration module is connected to the end-to-end convolutional-recurrent neural network module.
10. The visual odometry system according to claim 9, wherein: The end-to-end convolutional-recurrent neural network module includes: CNN submodule: includes a design without pooling layers, reduces the feature map size through convolutional layers with a stride of 2, and includes residual blocks; Bi-GRU submodule: uses a bidirectional gated recurrent unit, whose operations include reset gate, update gate and candidate hidden state calculation; Decoupled regression submodule: Contains a parallel rotation regression network and a translation regression network.