Unsupervised monocular vision pose estimation method based on multi-view spatial-temporal feature fusion
Through the multi-view spatial and temporal feature fusion method, ConvLSTM and feature fusion network are used to solve the problem of insufficient accuracy and adaptability of unsupervised monocular visual pose estimation in complex dynamic scenarios, and achieve higher pose estimation accuracy and robustness.
Patent Information
- Application Number
- CN202510451244.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-29
AI Technical Summary
Unsupervised monocular visual pose estimation networks have low accuracy and poor adaptability in pose prediction when dealing with complex dynamic scenarios, especially when occlusion, blur or drastic viewing angle changes, the model is difficult to maintain high accuracy and robustness.
The multi-view spatiotemporal feature fusion method is adopted to extract the temporal and spatial correlation of image sequences through the ConvLSTM network, and the feature fusion network is used to enhance high-dimensional motion features, including self-context feature enhancement modules and cross-feature enhancement modules, improving the adaptability and robustness of the model to complex scenes.
It significantly improves the accuracy and robustness of pose estimation, can better cope with occlusion and motion blur in complex dynamic environments, enhances the model's adaptability to dynamic scenes, and reduces the risk of overfitting in specific perspectives.
Smart Images

Figure CN120388218A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of unsupervised monocular visual pose estimation, and particularly relates to an unsupervised monocular visual pose estimation method for multi-view spatio-temporal feature fusion. Background Art
[0002] In the fields of unmanned driving, 3D reconstruction, and augmented reality (AR), accurately estimating the depth of objects in the environment and one's own pose (state estimation) has always been a technical difficulty. Traditional SLAM (Simultaneous Localization and Mapping) technology, although able to estimate pose and depth through feature point detection and geometric methods, often can only obtain sparse feature point depth information and cannot comprehensively cover all pixel points in the scene. In addition, traditional methods rely on expensive devices such as binocular vision or lidar, which pose problems of cost and complexity in the case of limited resources.
[0003] In recent years, monocular visual depth and pose estimation methods based on unsupervised learning have become a popular research direction. This method uses deep learning to estimate object depth and camera pose through a single camera video image, and has great potential to replace traditional methods. CADepth-Net (Channel Attention Depth Estimation Network) is one of the representatives of this method. It is a neural network for depth estimation tasks, mainly improving the accuracy and robustness of depth prediction and pose prediction through the channel attention mechanism. The application scenarios of this technology are very extensive and are applicable to autonomous driving, high-precision map reconstruction, 3D visual reconstruction, and AR / VR positioning, etc.
[0004] An existing technical solution is CADepth-Net (Channel Attention Depth Estimation Network). By introducing the channel attention mechanism, a deep convolutional neural network is used to extract high-level features of the image from different levels, including pose and depth information. By combining the depth features, the network can estimate the pose information of the camera, including translation and rotation, usually using an additional fully connected layer or regression module for pose output. The network is trained in an end-to-end manner and uses a dataset with depth labels, while optimizing the loss functions of depth estimation and pose estimation to ensure the coordinated improvement of both during the training process. Commonly used loss functions include mean square error (MSE) and other reprojection-based loss functions to ensure the accuracy of depth and pose.
[0005] Although CADepth-Net can achieve unsupervised monocular depth and pose estimation, there are still some deficiencies in dealing with complex dynamic scenes or long-term sequence data.
[0006] ①Insufficient temporal dependence. When the original model processes long-term continuous motion, it is difficult to maintain a good estimate of the global trajectory, resulting in an increase in cumulative error.
[0007] ② It is impossible to make full use of global feature information. In dynamic scenarios, especially when there are occlusions, blurs, or drastic changes in viewpoints, the model is easily restricted by local information, resulting in inaccurate pose estimation.
[0008] ③ The adaptability to dynamic scenarios is limited. The original model assumes that most of the scenarios are static and fails to model the motion patterns of dynamic objects, thus leading to a decrease in pose estimation accuracy in these scenarios.
[0009] ④ The feature expression ability is limited. In tasks that require simultaneous capture of multiple viewpoints or feature interactions, the feature expression ability of the original model is insufficient, affecting the accuracy of pose estimation. Summary of the Invention
[0010] In view of the above deficiencies in the prior art, an unsupervised monocular visual pose estimation method based on multi-view spatio-temporal feature fusion provided by the present invention solves the problems of low accuracy and poor adaptability in pose prediction of the unsupervised monocular visual pose estimation network when facing complex dynamic scenarios.
[0011] To achieve the above invention purpose, the technical solution adopted by the present invention is: an unsupervised monocular visual pose estimation method based on multi-view spatio-temporal feature fusion, including:
[0012] Obtain an image sequence, and according to the image sequence, splice adjacent frames pairwise to obtain a number of pairs of adjacent frame sequence data;
[0013] Process each pair of adjacent frame sequence data through a CADepth-Net channel attention depth estimation network to obtain a number of depth maps; and use a pose estimation network to perform convolutional layer processing on each pair of adjacent frame sequence data, and perform two rotations on each pair of adjacent frame sequence data after convolutional layer processing to obtain a high-dimensional motion feature map, a view feature map with width dimension priority, and a view feature map with height dimension priority, and perform feature fusion on the high-dimensional motion feature map, the view feature map with width dimension priority, and the view feature map with height dimension priority, and pass the result of feature fusion through a pose network decoder to obtain a number of pose transformation matrices;
[0014] Perform image registration on each depth map, each pose transformation matrix, and each pair of adjacent frame sequence data respectively to obtain a pose prediction result;
[0015] Calculate the loss according to the pose prediction result, and optimize the pose estimation network and the CADepth-Net channel attention depth estimation network to obtain a trained pose estimation network and a CADepth-Net channel attention depth estimation network;
[0016] Obtain the image to be measured, and use the trained pose estimation network and the CADepth-Net channel attention depth estimation network to complete pose estimation.
[0017] The beneficial effects of the present invention are as follows: After rotating the feature map to obtain a new perspective, the pose estimation network can mine deeper motion information and improve the processing ability of motion information; Through multi-view feature extraction, the model can show stronger adaptability in new scenarios and significantly reduce the risk of overfitting to a specific perspective; The addition of the ConvLSTM network helps to make full use of the temporal information between consecutive frames, fully model the dynamic correlation between the front and rear frames, and at the same time can reduce the prediction fluctuations and instabilities caused by noise; The context self-enhancement module of the feature fusion network significantly improves the accuracy of pose prediction by enhancing the representation ability of local features and better adapts to complex scenarios. The cross-feature enhancement module performs excellently in complex dynamic environments, enhances the model's ability to handle challenges such as occlusion and motion blur, and improves the overall robustness.
[0018] Further, the obtaining of a plurality of pose transformation matrices is specifically as follows:
[0019] Process the data of each adjacent frame sequence through a convolutional layer to obtain a plurality of high-dimensional motion feature maps;
[0020] Rotate each high-dimensional motion feature map in two different directions respectively to obtain a perspective feature map with width dimension priority and a perspective feature map with height dimension priority;
[0021] Perform spatial feature extraction on the high-dimensional motion feature map, the perspective feature map with width dimension priority, and the perspective feature map with height dimension priority respectively through the ConvLSTM convolutional long short-term memory network to obtain basic spatial features, horizontal spatial features, and vertical spatial features;
[0022] Input the basic spatial features, horizontal spatial features, and vertical spatial features into the feature fusion network to obtain multi-view fusion features;
[0023] Input the multi-view fusion features into the pose network decoder composed of a multi-layer convolutional neural network to extract the pose transformation between video frames, and obtain a plurality of pose transformation matrices.
[0024] The beneficial effects of the above further solution are as follows: Extract motion features from three dimensions respectively through the ConvLSTM network and capture information through the feature fusion network, thereby realizing multi-scale spatio-temporal feature extraction, local detail enhancement, and global motion consistency optimization, and overall improving the accuracy and robustness of pose estimation, and effectively coping with problems such as motion blur and occlusion.
[0025] Furthermore, the ConvLSTM (Convolutional Long Short-term Memory) network specifically replaces the fully connected layer in the LSTM (Long Short-term Memory) network with a convolutional layer; the expression for the ConvLSTM network to extract spatial features is as follows:
[0026]
[0027] where i t is the output of the input gate; σ is the Sigmoid activation function; W xi and W hi are both convolutional kernels of the input gate; X t is the feature map input at the current time step; H t-1 is the hidden state of the previous time step; W ci is the cell state weight of the input gate; C t-1 is the cell state of the previous time step; b i is the bias of the input gate; f t is the output of the forget gate; W xf and W hf are both convolutional kernel weights of the forget gate; W cf is the cell state weight of the forget gate; b f is the bias of the forget gate; C t is the cell state of the current time step; tanh is the activation function; W xc is the convolutional kernel input to the candidate state; W hc is the convolutional kernel from the hidden state to the candidate state; b c is the bias term of the candidate state; o t is the output of the output gate; W xo and W ho are both convolutional kernels of the output gate; W co is the cell state weight of the output gate; b o is the bias of the output gate; h t is the hidden state of the current time step; is the Hadamard product; * is the convolution operation.
[0028] The beneficial effect of the above further solution is that the ConvLSTM replaces the fully connected structure through convolution operations, retains the spatial structure of the data, and at the same time captures the temporal dynamics using the gating mechanism of the LSTM, thereby more accurately obtaining the spatio-temporal features of the camera movement.
[0029] Furthermore, the feature fusion network includes M layers of feature fusion units and a fourth cross-feature enhancement CFA module;
[0030] Each feature fusion unit includes a first self-context feature enhancement ECA module, a second self-context feature enhancement ECA module, a third self-context feature enhancement ECA module, and a first cross-feature enhancement CFA module, a second cross-feature enhancement CFA module, and a third cross-feature enhancement CFA module that are respectively connected to the first self-context feature enhancement ECA module, the second self-context feature enhancement ECA module, and the third self-context feature enhancement ECA module;
[0031] The first cross-feature enhancement CFA module, the second cross-feature enhancement CFA module, and the third cross-feature enhancement CFA module of the current layer are connected to the first self-context feature enhancement ECA module, the second self-context feature enhancement ECA module, and the third self-context feature enhancement ECA module of the next layer in one-to-one correspondence;
[0032] The first cross-feature enhancement CFA module, the second cross-feature enhancement CFA module, and the third cross-feature enhancement CFA module of the M-th layer are all connected to the fourth cross-feature enhancement CFA module;
[0033] The fourth cross-feature enhancement CFA module outputs multi-view fusion features.
[0034] The beneficial effect of the above further solution is that through M-layer gradual iteration, the feature saliency within the view (ECA module) is gradually enhanced and the cross-view feature alignment (CFA module) is deepened at each layer. Finally, the output motion features can simultaneously retain local details (such as edge displacement) and global motion consistency (such as camera trajectory smoothness), effectively alleviating the monocular scale ambiguity problem.
[0035] Further, the second cross-feature enhancement CFA module and the third cross-feature enhancement CFA module of the M-th layer respectively output k feature vectors and v feature vectors, and the k feature vectors and v feature vectors of the second cross-feature enhancement CFA module and the third cross-feature enhancement CFA module of the M-th layer are fused to obtain k feature fusion vectors and v feature fusion vectors;
[0036] The k feature fusion vectors, the v feature fusion vectors, and the q feature vectors output by the first cross-feature enhancement CFA module of the M-th layer are input into the fourth cross-feature enhancement CFA module to obtain multi-view fusion features.
[0037] The beneficial effect of the above further solution is that the k / v fusion of the second and third CFA modules can eliminate redundant view noise and retain the key motion patterns shared across views (such as common depth change cues); finally, the q-k-v interaction of the fourth CFA module adaptively fuses multi-directional features through attention weights, significantly improving the pose estimation robustness under complex motions (such as rotation + translation).
[0038] Furthermore, the inputs of the first self-context feature enhancement ECA module, the second self-context feature enhancement ECA module, and the third self-context feature enhancement ECA module in the first layer are the basic spatial features, the spatial features in the horizontal direction, and the spatial features in the vertical direction, respectively.
[0039] The beneficial effects of the above further solution are as follows: Through the channel attention mechanism, the ECA module adaptively enhances the key local features (such as edge displacement and texture change) of each motion dimension (horizontal / vertical / basic), suppresses irrelevant noise, and thus improves the accuracy of pose prediction - for example, accurately decoupling the camera rotation and translation motions. At the same time, the differential weighting of different channels by ECA can handle view-related interferences (such as single-view occlusion or lighting changes). By retaining the robust features unique to each dimension (such as vertical features being more sensitive to height changes), the generalization ability of the model in complex scenarios is enhanced, and the collapse of the overall pose estimation due to the failure of a certain view is avoided.
[0040] Furthermore, the first cross-feature enhancement CFA module in each layer receives the q feature vector output by the corresponding self-context feature enhancement ECA module and the k feature vector and v feature vector output by the other two self-context feature enhancement ECA modules; the self-context feature enhancement ECA modules corresponding to the first cross-feature enhancement CFA module, the second cross-feature enhancement CFA module, and the third cross-feature enhancement CFA module are the first self-context feature enhancement ECA module, the second self-context feature enhancement ECA module, and the third self-context feature enhancement ECA module, respectively.
[0041] The beneficial effects of the above further solution are as follows: By directionally transmitting the q / k / v feature vectors (such as the first CFA module receiving the q of its own ECA and the k / v of other ECAs), the model is forced to explicitly model the non-linear coupling relationship between different motion dimensions (horizontal / vertical / basic) (such as camera rotation affecting both horizontal and vertical displacements simultaneously), thereby improving the accuracy of complex motion decoupling; each ECA module focuses on optimizing its own view features (such as the first ECA strengthening the basic motion), while the CFA module selectively fuses complementary information through cross-attention (such as the k / v of the horizontal view assisting in correcting the q of the vertical view), avoiding the degradation of single-view features, and significantly enhancing the robustness to interferences such as occlusion and motion blur.
[0042] Furthermore, the input of each cross-feature enhancement CFA module is the received q feature vector and the first k feature fusion vector and the first v feature fusion vector obtained by respectively fusing the two groups of received k feature vectors and v feature vectors.
[0043] The beneficial effects of the above further solution are as follows: By fusing two groups of k / v feature vectors and then performing attention interaction with the q feature, the model can more comprehensively integrate multi-perspective motion cues - the k / v fusion process can automatically filter out conflicting or noisy features between perspectives (such as a perspective becoming invalid due to occlusion), retain key motion patterns with high consistency, and then dynamically align with the q feature through the attention mechanism to enhance the robustness to local interferences (such as dynamic objects and illumination changes). BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 It is a flowchart of the method of the present invention.
[0045] Figure 2 It is a schematic diagram of the feature fusion network structure of the present invention.
[0046] Figure 3 It is a schematic diagram of the structures of the ECA module and the CFA module of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] The following describes the specific embodiments of the present invention to facilitate those skilled in the art of the present technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.
[0048] As Figure 1 shown, in an embodiment of the present invention, an unsupervised monocular visual pose estimation method for multi-perspective spatio-temporal feature fusion includes:
[0049] Obtain an image sequence, and according to the image sequence, splice adjacent frames pairwise to obtain a number of pairs of adjacent frame sequence data;
[0050] Process each pair of adjacent frame sequence data through the CADepth-Net channel attention depth estimation network to obtain a number of depth maps; and use the pose estimation network to perform convolutional layer processing on each pair of adjacent frame sequence data, and perform two rotations on the processed adjacent frame sequence data to obtain a high-dimensional motion feature map, a perspective feature map with width dimension prioritized, and a perspective feature map with height dimension prioritized, and fuse the features of the high-dimensional motion feature map, the perspective feature map with width dimension prioritized, and the perspective feature map with height dimension prioritized, and pass the result of the feature fusion through the pose network decoder to obtain a number of pose transformation matrices;
[0051] Perform image registration on each depth map, each pose transformation matrix, and each pair of adjacent frame sequence data respectively to obtain a pose prediction result;
[0052] Calculate the loss based on the pose prediction result, and optimize the pose estimation network and the CADepth-Net channel attention depth estimation network to obtain the trained pose estimation network and the CADepth-Net channel attention depth estimation network;
[0053] Obtain the image to be measured, and use the trained pose estimation network and the CADepth-Net channel attention depth estimation network to complete pose estimation.
[0054] In this embodiment, the present invention aims to solve the problems of low accuracy and poor adaptability in pose prediction when the unsupervised monocular visual pose estimation network faces complex dynamic scenes. By adding a ConvLSTM network and a feature fusion network, high-dimensional motion features are extracted, enhanced, and fused from multiple perspectives. This improvement breaks through the convention of single-perspective convolution of traditional convolutional neural networks, improves the accuracy and robustness of the pose network model, and improves the ability and accuracy of the traditional pose estimation network to predict the pose change of an object in a complex dynamic scene. The improvement in the accuracy of pose change prediction also improves the accuracy of depth prediction, and generally improves the accuracy of visual odometry.
[0055] In this embodiment, based on the unsupervised monocular visual odometry pose estimation based on sequence learning, two new convolutional perspectives are added to extract the motion information of high-dimensional feature maps, and a "convolutional long short-term memory network" (hereinafter referred to as the ConvLSTM network) is added to extract the temporal and spatial associations of the image sequence, and a "feature fusion network based on the multi-head attention mechanism" (hereinafter referred to as the feature fusion network) is added to fuse the high-dimensional motion features of the new perspective, ultimately improving the ability of the model to predict the pose.
[0056] First, splice the video sequences in pairs and input them into the pose network, which can effectively capture the motion information (such as translation, rotation, etc.) between adjacent frames and help the network learn the pose change T between adjacent frames t→t+1。Next, the ConvLSTM network is used to extract the multi-view motion features of the object from three different dimensions of convolution: Convolution in the channel direction is beneficial to capturing the correlation between different channels and enhancing the model's understanding of complex motion patterns; Convolution in the width and height directions extracts information from the perspectives of feature distribution and spatial structure respectively, obtaining multi-view multi-dimensional motion information. The specific operation is to rotate the feature map twice in different directions to obtain three perspective feature maps with the same shape (prioritizing the channel number dimension), prioritizing the width dimension, and prioritizing the height dimension. The three perspective feature maps are respectively input into three ConvLSTM networks, and the three output feature maps containing temporal information respectively enhance the spatial relationship (geometry structure-related information) of the pixels in their respective perspectives, the motion features in the horizontal direction, and the motion features in the vertical direction. For the entire video frame sequence, the high-dimensional temporal features of the object motion information are enhanced. Then, the above three feature maps are respectively used as inputs and enter the next feature fusion network composed of a self-context feature enhancement module (ECA) and a cross-feature enhancement module (CFA). Based on the multi-head attention mechanism, the self-context feature enhancement module enhances the effective information of the feature map, and the cross-feature enhancement module fuses the effective features under the three enhanced perspectives. This feature fusion network enhances and fuses the multi-view high-dimensional motion information features, reduces the redundancy of the features, improves the understanding ability of complex dynamic scenes, and helps with subsequent pose estimation. The feature map after feature enhancement and fusion is input into a pose decoder composed of a multi-layer convolutional neural network to extract the pose transformation between video frames. The multi-view feature enhancement network and feature fusion network proposed by the present invention effectively improve the accuracy of visual odometry depth estimation and pose estimation.
[0057] The obtaining of several pose transformation matrices is specifically as follows:
[0058] The data of each adjacent frame sequence is processed through a convolutional layer to obtain several high-dimensional motion feature maps;
[0059] Each high-dimensional motion feature map is respectively rotated twice in different directions to obtain a perspective feature map prioritizing the width dimension and a perspective feature map prioritizing the height dimension;
[0060] The high-dimensional motion feature map, the perspective feature map prioritizing the width dimension, and the perspective feature map prioritizing the height dimension are respectively subjected to spatial feature extraction through a ConvLSTM (Convolutional Long Short-Term Memory) network to obtain the basic spatial feature, the spatial feature in the horizontal direction, and the spatial feature in the vertical direction;
[0061] The basic spatial feature, the spatial feature in the horizontal direction, and the spatial feature in the vertical direction are input into the feature fusion network to obtain multi-view fusion features;
[0062] Input the multi-view fusion features into the pose network decoder composed of a multi-layer convolutional neural network to extract the pose transformation between video frames, and obtain a number of pose transformation matrices.
[0063] In this embodiment, monocular image sequence data is obtained, segmented by a fixed-step sliding window according to the monocular image sequence data, and adjacent two-frame image pairs are spliced to obtain N pairs of adjacent frame sequence data. Adjust the input channel number of the pose network with ResNet as the architecture to 3(N + 1), input the sequence into convolutional layer 1, and initially obtain multiple high-dimensional motion feature maps, whose tensor structure can be expressed as where C represents the number of image channels, and H and W represent the height and width of the feature map respectively. Use the torch.cat function to stack the N pairs of feature maps containing motion information in the channel number dimension into a high-dimensional multi-channel tensor with the structure of N×C×H×W.
[0064] Rotate the obtained high-dimensional motion feature maps in the following ways respectively to prepare for preferentially convolving and extracting spatial features and horizontal and vertical direction features in different dimensions.
[0065] Prioritize the channel number dimension, and the feature map shape remains N×C×H×W: used to extract and enhance spatial features such as inter-frame motion information, topological relationships, and texture information.
[0066] Prioritize the width dimension, and the feature map shape is changed to N×W×C×H, with the original width dimension of the feature map as the new feature map channel number dimension: used to extract and enhance the motion features in the horizontal direction.
[0067] Prioritize the height dimension, and the feature map shape is changed to N×W×C×W, with the original height dimension of the feature map as the new feature map channel number dimension: used to extract and enhance the motion features in the vertical direction.
[0068] The ConvLSTM convolutional long short-term memory network specifically replaces the fully connected layer in the LSTM long short-term memory network with a convolutional layer; the expression for the ConvLSTM convolutional long short-term memory network to extract spatial features is:
[0069]
[0070] where, i t is the output of the input gate; σ is the Sigmoid activation function; W xi and W hi are both convolutional kernels of the input gate; X t is the feature map input at the current time step; H t-1 is the hidden state at the previous time step; W ci is the cell state weight of the input gate; C t-1 is the cell state at the previous time step; b iis the bias of the input gate; f t is the output of the forget gate; W xf and W hf are both the convolutional kernel weights of the forget gate; W cf is the cell state weight of the forget gate; b f is the bias of the forget gate; C t is the cell state at the current time step; tanh is the activation function; W xc is the convolutional kernel input to the candidate state; W hc is the convolutional kernel from the hidden state to the candidate state; b c is the bias term of the candidate state; o t is the output of the output gate; W xo and W ho are both the convolutional kernel weights of the output gate; W co is the cell state weight of the output gate; b o is the bias of the output gate; h t is the hidden state at the current time step; is the Hadamard product; * is the convolution operation.
[0071] In this embodiment, based on the traditional long short-term memory network (LSTM), the ConvLSTM network transforms the input data from one-dimensional to multi-dimensional feature maps, and the fully connected layer of its computing unit becomes a convolutional layer.
[0072] In this step, the three high-dimensional feature vectors in the rotation direction are respectively input into three ConvLSTM networks. Each high-dimensional feature vector is stacked into N pairs of adjacent frame feature vectors in the dimension of its number of channels, and in the form of X1, X2,... X N , X1′, X2′,... X N ′, X1″, X2″,... X N ″ are respectively input into the three ConvLSTM networks.
[0073] The output corresponding to the time step not only captures the immediate information input at the current time step, but also preserves the key information of the previous time step and context.
[0074] Finally, the outputs of the three ConvLSTM networks are stacked in the dimension of the number of channels respectively, forming three high-dimensional feature maps with temporal information in three perspectives. This spatial multi-perspective feature enhances the front-back correlation in the time dimension, which is beneficial to improving the accuracy and robustness of pose estimation.
[0075] The feature fusion network includes M layers of feature fusion units and a fourth cross-feature enhancement CFA module;
[0076] Each feature fusion unit includes a first self-context feature enhancement ECA module, a second self-context feature enhancement ECA module, a third self-context feature enhancement ECA module, and a first cross-feature enhancement CFA module, a second cross-feature enhancement CFA module, and a third cross-feature enhancement CFA module that are respectively connected to the first self-context feature enhancement ECA module, the second self-context feature enhancement ECA module, and the third self-context feature enhancement ECA module;
[0077] The first cross-feature enhancement CFA module, the second cross-feature enhancement CFA module, and the third cross-feature enhancement CFA module of the current layer are connected to the first self-context feature enhancement ECA module, the second self-context feature enhancement ECA module, and the third self-context feature enhancement ECA module of the next layer in one-to-one correspondence;
[0078] The first cross-feature enhancement CFA module, the second cross-feature enhancement CFA module, and the third cross-feature enhancement CFA module of the Mth layer are all connected to the fourth cross-feature enhancement CFA module;
[0079] The fourth cross-feature enhancement CFA module outputs multi-view fusion features.
[0080] The second cross-feature enhancement CFA module and the third cross-feature enhancement CFA module of the Mth layer respectively output k feature vectors and v feature vectors, and fuse the k feature vectors and v feature vectors of the second cross-feature enhancement CFA module and the third cross-feature enhancement CFA module of the Mth layer to obtain k feature fusion vectors and v feature fusion vectors;
[0081] Input the k feature fusion vectors, v feature fusion vectors, and q feature vectors output by the first cross-feature enhancement CFA module of the Mth layer into the fourth cross-feature enhancement CFA module to obtain multi-view fusion features.
[0082] The inputs of the first self-context feature enhancement ECA module, the second self-context feature enhancement ECA module, and the third self-context feature enhancement ECA module of the first layer are the basic spatial features, the spatial features in the horizontal direction, and the spatial features in the vertical direction respectively.
[0083] Each layer's first cross-feature enhancement CFA module receives the q feature vectors output by the corresponding self-context feature enhancement ECA module and the k feature vectors and v feature vectors output by the other two self-context feature enhancement ECA modules; the self-context feature enhancement ECA modules corresponding to the first cross-feature enhancement CFA module, the second cross-feature enhancement CFA module, and the third cross-feature enhancement CFA module are the first self-context feature enhancement ECA module, the second self-context feature enhancement ECA module, and the third self-context feature enhancement ECA module in sequence.
[0084] The input of each cross - feature enhanced CFA module is the received q feature vector, as well as the first k - feature fusion vector and the first v - feature fusion vector obtained by respectively fusing the received two groups of k feature vectors and v feature vectors.
[0085] In this embodiment, as Figure 2 and Figure 3 shown, the spatial features are pre - processed and input into a feature fusion network composed of a self - context feature enhancement module (ECA) and a cross - feature enhancement module (CFA). This feature fusion network consists of an M - layer fusion network and a decoder. Each layer of the fusion network includes an ECA module and a CFA module, and the internal structure is as follows:
[0086] First, three parallel ECA modules are close to the input, and then three parallel CFA modules are connected. Each ECA module is respectively connected to a CFA module. The CFA module that processes spatial features such as inter - frame motion information, topological relationship, and texture information simultaneously receives the q feature transmitted by the ECA module that processes spatial features such as inter - frame motion information, topological relationship, and texture information, and the k and v features transmitted by the other two ECAs; the CFA module that processes horizontal motion information simultaneously receives the q feature transmitted by the ECA module that processes horizontal motion information, and the k and v transmitted by the other two ECAs; the CFA module that processes vertical motion information is the same. This structure is connected M times. The network layer closer to the input extracts the shallow features of the feature map, such as features with less semantic information like edges and textures; the network layer closer to the output extracts deeper information, such as the shape of objects, motion trajectories, etc., which have richer context information and stronger generalization ability. As the depth of the feature fusion network increases, the attention mechanism will fuse more complex dynamic and static features, providing more comprehensive information for subsequent decoding and pose prediction. The last independent CFA module, as a decoder, fuses the k, v, and q values of the M - layer fusion network to generate the final feature fusion map. Among them, the outputs of the two CFA modules that fuse motion information in the last layer of the fusion network are output and concatenated as the fused k and v features; the output of the other CFA module is used as the q feature.
[0087] (1) Selection of the input of the feature map.
[0088] The feature maps that enhance motion features from three perspectives are processed as follows: The tensor of the unrotated spatial feature map is input into the ECA module that processes spatial features such as inter - frame motion information, topological relationship, and texture information; the tensor of the rotated spatial feature map that enhances motion features preferentially in the width dimension is input into the ECA module that processes horizontal - direction motion information; the tensor of the rotated spatial feature map that enhances motion features preferentially in the height dimension is input into the ECA module that processes vertical - direction motion information.
[0089] (2) The self-context feature enhancement module enhances the feature representation.
[0090] The self-context enhancement module corresponding to each time step is based on a multi-head adaptive mechanism and combines spatial position encoding to introduce different position information in the feature map.
[0091] The three ECA modules respectively extract the surrounding region features and the target motion features. The surrounding region features are such as static background, repeated image information, irrelevant objects, etc.; the target features are such as the same type of features under two perspectives (horizontal, vertical), such as object motion trajectories, depth information, optical flow change features, etc. The ECA modules as a whole further enhance the effective information in the feature map.
[0092] (3) The cross-feature enhancement module fuses local features and surrounding features.
[0093] X q represents the feature query map, which is used to determine the attention weights between channels; X kv represents the feature maps of keys and values, which are used to compare and weight with the feature query map. This cross-feature enhancement module is based on a multi-head cross-attention mechanism and combines spatial position encoding and a feed-forward neural network to further fuse local features and surrounding features.
[0094] In the current attention mechanism, X q can represent the key features or target poses in the current frame sequence to ensure that the attention focus of the model always focuses on the pose or feature changes in the current frame sequence. X kv provides context information.
[0095] After each CFA module calculates the attention weights, it weights and aggregates the corresponding v features and outputs the fused features. This aggregation helps to reduce redundant information and strengthen the influence of important features. By introducing k and v features from the ECA modules (i.e., other branches) that process different types of information, the CFA module can obtain cross-frame or cross-branch context information. These context information help the model to more accurately compare, match and fuse the current query features with the information in other feature branches in the cross-attention mechanism.
[0096] The feature map sequence after multi-view feature fusion continues to enter the pose network decoder to output the matrix representing the pose changes between frames. This pose matrix sequence, together with the original video frames and the predicted depth map, undergoes an image registration operation to predict s on the basis of the original video frame I images. This image sequence is more accurate than the prediction results of the original unsupervised model.
Claims
1. An unsupervised monocular visual pose estimation method for multi-view spatio-temporal feature fusion, characterized in that, Including: Obtain an image sequence, and according to the image sequence, splice adjacent frames pairwise to obtain a number of pairs of adjacent frame sequence data; Process each adjacent frame sequence data through a CADepth-Net channel attention depth estimation network to obtain a number of depth maps; and use a pose estimation network to perform convolutional layer processing on each adjacent frame sequence data, and perform two rotations on each adjacent frame sequence data after convolutional layer processing to obtain a high-dimensional motion feature map, a width-dimension-prior perspective feature map, and a height-dimension-prior perspective feature map, and fuse the features of the high-dimensional motion feature map, the width-dimension-prior perspective feature map, and the height-dimension-prior perspective feature map, and pass the result of the feature fusion through a pose network decoder to obtain a number of pose transformation matrices; Perform image registration on each depth map, each pose transformation matrix, and each adjacent frame sequence data respectively to obtain a pose prediction result; Calculate the loss according to the pose prediction result, and optimize the pose estimation network and the CADepth-Net channel attention depth estimation network to obtain a trained pose estimation network and a CADepth-Net channel attention depth estimation network; Obtain a to-be-tested image, and use the trained pose estimation network and CADepth-Net channel attention depth estimation network to complete pose estimation.
2. The unsupervised monocular visual pose estimation method for multi-view spatio-temporal feature fusion according to claim 1, wherein The obtaining of a number of pose transformation matrices is specifically as follows: Process each adjacent frame sequence data through a convolutional layer to obtain a number of high-dimensional motion feature maps; Rotate each high-dimensional motion feature map twice in different directions respectively to obtain a width-dimension-prior perspective feature map and a height-dimension-prior perspective feature map; Extract spatial features of the high-dimensional motion feature map, the width-dimension-prior perspective feature map, and the height-dimension-prior perspective feature map respectively through a ConvLSTM convolutional long short-term memory network to obtain a basic spatial feature, a horizontal-direction spatial feature, and a vertical-direction spatial feature; Input the basic spatial feature, the horizontal-direction spatial feature, and the vertical-direction spatial feature into a feature fusion network to obtain a multi-view fusion feature; Input the multi-view fusion feature into a pose network decoder composed of a multi-layer convolutional neural network to extract the pose transformation between video frames to obtain a number of pose transformation matrices.
3. The unsupervised monocular visual pose estimation method for multi-view spatio-temporal feature fusion according to claim 2, characterized in that, The ConvLSTM convolutional long short-term memory network specifically replaces the fully connected layer in the LSTM long short-term memory network with a convolutional layer; the expression for the ConvLSTM convolutional long short-term memory network to extract spatial features is: where, i t is the output of the input gate; σ is the Sigmoid activation function; W xi and W hi are both the convolutional kernels of the input gate; X t is the feature map input at the current time step; H t-1 is the hidden state at the previous time step; W ci is the cell state weight of the input gate; C t-1 is the cell state at the previous time step; b i is the bias of the input gate; f t is the output of the forget gate; W xf and W hf are both the convolutional kernel weights of the forget gate; W cf is the cell state weight of the forget gate; b f is the bias of the forget gate; C t is the cell state at the current time step; tanh is the activation function; W xc is the convolutional kernel input to the candidate state; W hc is the convolutional kernel from the hidden state to the candidate state; b c is the bias term of the candidate state; o t is the output of the output gate; W xo and W ho are both the convolutional kernels of the output gate; W co is the cell state weight of the output gate; b o is the bias of the output gate; h t is the hidden state at the current time step; is the Hadamard product; * is the convolution operation.
4. The unsupervised monocular visual pose estimation method for multi-view spatio-temporal feature fusion according to claim 2, characterized in that, The feature fusion network includes M feature fusion units and a fourth cross feature augmentation CFA module; Each feature fusion unit includes a first self-context feature augmentation ECA module, a second self-context feature augmentation ECA module, a third self-context feature augmentation ECA module, and a first cross feature augmentation CFA module, a second cross feature augmentation CFA module, and a third cross feature augmentation CFA module that are respectively connected to the first self-context feature augmentation ECA module, the second self-context feature augmentation ECA module, and the third self-context feature augmentation ECA module; The first cross - feature enhancement CFA module, the second cross - feature enhancement CFA module, and the third cross - feature enhancement CFA module of the current layer are connected one - to - one with the first self - context feature enhancement ECA module, the second self - context feature enhancement ECA module, and the third self - context feature enhancement ECA module of the next layer; The first cross - feature enhancement CFA module, the second cross - feature enhancement CFA module, and the third cross - feature enhancement CFA module of the M - th layer are all connected to the fourth cross - feature enhancement CFA module; The fourth cross - feature enhancement CFA module outputs multi - perspective fusion features.
5. The unsupervised monocular visual pose estimation method for multi-view spatio-temporal feature fusion according to claim 4, wherein, The second cross - feature enhancement CFA module and the third cross - feature enhancement CFA module of the M - th layer respectively output k feature vectors and v feature vectors, and fuse the k feature vectors and v feature vectors of the second cross - feature enhancement CFA module and the third cross - feature enhancement CFA module of the M - th layer to obtain k feature fusion vectors and v feature fusion vectors; Input the k feature fusion vectors, v feature fusion vectors, and the q feature vectors output by the first cross - feature enhancement CFA module of the M - th layer into the fourth cross - feature enhancement CFA module to obtain multi - perspective fusion features.
6. The unsupervised monocular visual pose estimation method for multi-view spatio-temporal feature fusion according to claim 4, wherein The inputs of the first self - context feature enhancement ECA module, the second self - context feature enhancement ECA module, and the third self - context feature enhancement ECA module of the first layer are the basic spatial features, the spatial features in the horizontal direction, and the spatial features in the vertical direction respectively.
7. The unsupervised monocular visual pose estimation method for multi-view spatio-temporal feature fusion according to claim 4, characterized in that, The first cross - feature enhancement CFA module of each layer receives the q feature vectors output by the corresponding self - context feature enhancement ECA module and the k feature vectors and v feature vectors output by the other two self - context feature enhancement ECA modules; the self - context feature enhancement ECA modules corresponding to the first cross - feature enhancement CFA module, the second cross - feature enhancement CFA module, and the third cross - feature enhancement CFA module are the first self - context feature enhancement ECA module, the second self - context feature enhancement ECA module, and the third self - context feature enhancement ECA module in sequence.
8. The unsupervised monocular visual pose estimation method for multi-view spatio-temporal feature fusion according to claim 7, characterized in that The input of each cross - feature enhancement CFA module is the received q feature vectors and the first k feature fusion vector and the first v feature fusion vector obtained by respectively fusing the received two groups of k feature vectors and v feature vectors.