Scene Flow Estimation Method and System Based on 3D Motion Feature Encoding

Through the scene flow estimation network model based on 3D motion feature encoding, the problems of large calculation complexity and parameter quantity in the prior art are solved, and the scene flow estimation effect with high precision, high speed and low parameter quantity is achieved.

CN117635650BActive Publication Date: 2025-07-01NAT UNIV OF DEFENSE TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311809121.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-26
Publication Date
2025-07-01
Estimated Expiration
2043-12-26

AI Technical Summary

Technical Problem

While improving the accuracy, the existing scenario flow estimation method has a large calculation complexity and parameter volume, which limits its wide application.

Method used

A scene flow estimation network model based on 3D motion feature encoding is adopted. Through feature encoding and context encoding, a pixel-by-pixel correlation tensor is calculated and a correlation pyramid is generated. Iterative prediction is performed in combination with 3D motion feature encoding, and finally upsampling is used to generate a full-resolution scene stream.

Benefits of technology

The scene flow estimation with high precision, high speed and low parameter quantity is realized, and the estimation performance and computing efficiency are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117635650B_ABST
    Figure CN117635650B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for scene flow estimation based on 3D motion feature encoding. The method includes: obtaining two adjacent RGB images and corresponding depth maps to obtain image features and context features; calculating the per-pixel correlation tensor of the image features and pooling to generate a four-layer correlation pyramid; based on the correlation pyramid, context features, and depth maps, obtaining the coarse-resolution image plane motion and coarse-resolution depth dimension motion, and performing upsampling to generate the full-resolution plane and depth dimension motion, and finally combining the camera internal parameters to obtain the scene flow. The present invention is applied to the field of scene flow estimation. Based on 3D motion feature encoding, it fully excavates the 3D correlation and motion clues of point pairs between adjacent frames, thereby greatly improving the accuracy of scene flow estimation, and having significant advantages in terms of inference speed and the number of parameters, with the characteristics of high precision, high speed, and low number of parameters, so as to improve the estimation performance and calculation efficiency in current scene flow estimation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of scene flow estimation, and specifically to a scene flow estimation method and system based on 3D motion feature encoding. Background Art

[0002] Scene Flow Estimation (SFE) aims to estimate the per-pixel / point three-dimensional motion between consecutive stereo, RGB-D, or point cloud data frames. Due to its wide range of applications, including obstacle avoidance and path planning, target tracking, gesture recognition, human pose estimation, moving object detection and segmentation, etc., SFE has been widely studied in recent years.

[0003] Traditional methods for estimating scene flow usually utilize variational techniques, which can be solved as an energy minimization process. Recently, deep learning has emerged as a powerful method, making it possible to directly extract feature representations and patterns from data. Existing methods generally rely on feature similarity as their basic principle. For example, FlowNet3D applies PointNet++ to directly operate on point clouds. It proposes to calculate a flow embedding at each layer to obtain the correlation between two point clouds, and then propagate it through a refinement layer to estimate the scene flow. The current state-of-the-art CamLiFlow contains two branches, 2D and 3D, along with multiple bidirectional connections at specific layers. Each of its branches adopts a pyramid structure to refine the optical flow and scene flow from the feature similarity tensors of two frames in a coarse-to-fine manner. Furthermore, there are some variants of feature similarity. For example, OpticalExp creatively uses optical flow dilation to recover the depth change between frames. By combining the initial depth value of the first frame, it can estimate the depth value of the second frame. This optical flow dilation essentially also stems from correct optical flow estimation. RAFT-3D uses semantic information to classify each pixel into a soft rigid body group, and then solves the 3D rigid body transformation of the group through a differentiable Gauss-Newton iteration. In order to improve the accuracy, these scene flow methods often design complex model structures or introduce additional semantic extraction networks, thereby increasing the computational complexity and the number of parameters, which limits their wide application. Summary of the Invention

[0004] Aiming at the deficiencies in the above-mentioned prior art, the present invention provides a scene flow estimation method and system based on 3D motion feature encoding, which can achieve high accuracy, high speed, and low number of parameters.

[0005] To achieve the above object, the present invention provides a scene flow estimation method based on 3D motion feature encoding, which estimates the scene flow by using a scene flow estimation network model based on 3D motion feature encoding, and includes the following steps:

[0006] Step 1, obtain two adjacent RGB images I t and I t+1 as well as the corresponding depth maps D t and D t+1 ;

[0007] Step 2, based on feature encoding and context encoding, obtain the image features t of image I and context features as well as the image features t+1 of image I

[0008] Step 3, calculate the per-pixel correlation tensor between the image features and the image features and pool to generate a four-layer correlation pyramid Y t ;

[0009] Step 4, based on the correlation pyramid Y t , context features and depth maps D t and D t+1 , combined with 3D motion feature encoding for iterative prediction, obtain the coarse-resolution image plane motion after n iterations and the coarse-resolution depth dimension motion

[0010] Step 5, upsample the coarse-resolution image plane motion and the coarse-resolution depth dimension motion to generate the full-resolution plane and depth dimension motion and based on the camera intrinsics and the full-resolution plane and depth dimension motion obtain the scene flow

[0011] To achieve the above object, the present invention also provides a scene flow estimation system based on 3D motion feature encoding. Based on the scene flow estimation network model, sample the above method to estimate the scene flow. The scene flow estimation system includes:

[0012] A feature encoder for extracting the image features of RGB images I t and I t+1 ;

[0013] A context encoder for extracting the context features of RGB image I t ;

[0014] A correlation pyramid calculation module for calculating the correlation between the image features and the image features Per-pixel correlation tensors and pooling to generate a four-layer correlation pyramid Y t ;

[0015] A 3D single-resolution iterative module for iteratively predicting based on the correlation pyramid Y t , context features and depth map D t , D t+1 , combined with 3D motion feature encoding for iterative prediction to obtain the coarse-resolution image plane motion after n iterations and the coarse-resolution depth dimension motion

[0016] An upsampling module for upsampling the coarse-resolution image plane motion and the coarse-resolution depth dimension motion to generate full-resolution plane and depth dimension motion

[0017] A scene flow estimation module for obtaining the scene flow based on the camera intrinsics and the full-resolution plane and depth dimension motion to obtain the scene flow

[0018] Compared with the prior art, the present invention has the following beneficial technical effects:

[0019] 1. The present invention designs a brand-new scene flow estimation network model based on 3D motion feature encoding. Through the encoding of 3D motion features, the scene flow estimation network model can uniformly process the three dimensions of the image motion plane and depth from a 2D perspective, and the extracted 3D motion features are highly compatible with the GRU iterative module in the model, capable of fully mining the 3D correlation and motion cues between adjacent frames, thereby greatly improving the accuracy of scene flow estimation;

[0020] 2. The present invention does not need to design a complex model structure or introduce an additional semantic extraction network, and has significant advantages in terms of inference speed and the number of parameters, combining the characteristics of high precision, high speed, and low number of parameters, thereby improving the estimation performance and computational efficiency in current scene flow estimation.

[0021] Thus, the scene flow estimation method in the present invention can combine the characteristics of high precision, high speed, and low number of parameters. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on the structures shown in these drawings without creative efforts.

[0023] Figure 1 This is the flowchart of the scene flow estimation method based on 3D motion feature encoding in Embodiment 1 of the present invention;

[0024] Figure 2 This is the iterative schematic diagram of the GRU predictor in Embodiment 1 of the present invention;

[0025] Figure 3 This is the schematic diagram of the generation principle of 3D motion features in Embodiment 1 of the present invention;

[0026] Figure 4 This is the structural block diagram of the scene flow estimation system based on 3D motion feature encoding in Embodiment 2 of the present invention.

[0027] The realization, functional features and advantages of the object of the present invention will be further described in conjunction with the embodiments with reference to the accompanying drawings. Detailed implementation manners

[0028] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts belong to the scope of protection of the present invention.

[0029] It should be noted that all directional indications (such as up, down, left, right, front, back...) in the embodiments of the present invention are only used to explain the relative positional relationship and movement conditions between components in a specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indications will also change accordingly.

[0030] In addition, the technical solutions between the various embodiments of the present invention can be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement it. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

[0031] Embodiment 1

[0032] This embodiment discloses a scene flow estimation method based on 3D motion feature encoding. This method estimates the scene flow by sampling a scene flow estimation network model based on 3D motion feature encoding.

[0033] The scene flow estimation network model includes a feature encoder, a context encoder, a correlation pyramid calculation module, a 3D single-resolution iterative module, an upsampling module, and a scene flow estimation module. Among them, the feature encoder and the context encoder are networks with the same structure and independent parameters, which are used to extract the image features and context features of RGB images respectively. The correlation pyramid calculation module is mainly used to perform a pixel-by-pixel dot product operation on the image features of two adjacent frames of RGB images, and then pool the correlation tensor to generate a four-layer correlation pyramid. The 3D single-resolution iterative module is essentially a GRU iterative module, that is, the data enters the same GRU prediction network multiple times in an iterative manner, so as to obtain the rough-resolution image plane motion after n iterations and the rough-resolution depth dimension motion The upsampling module then performs an upsampling operation on the rough-resolution image plane motion and the rough-resolution depth dimension motion to generate the full-resolution plane and depth dimension motion Finally, the scene flow estimation module combines the camera internal parameters to convert the full-resolution plane and depth dimension motion into the final scene flow

[0034] Reference Figure 1 , in this embodiment, the scene flow estimation method based on 3D motion feature encoding specifically includes the following steps:

[0035] Step 1, obtain two adjacent frames of RGB images I t , I t+1 and the corresponding depth maps D t , D t+1 ;

[0036] Step 2, based on feature encoding and context encoding, obtain the image features t and context features of the image I and the image features t+1 of the image I

[0037] Step 3, calculate the pixel-by-pixel correlation tensor of the image features and the image features , and pool to generate a four-layer correlation pyramid Y t ;

[0038] Step 4, based on the correlation pyramid Y t , context features and depth maps D t , D t+1 , obtain the rough-resolution image plane motion after n iterations With the coarse-resolution depth dimension motion Perform iterative prediction in combination with 3D motion feature encoding;

[0039] Step 5, upsample the coarse-resolution image plane motion With the coarse-resolution depth dimension motion To generate full-resolution plane and depth dimension motion And based on the camera intrinsics, obtain the scene flow from the full-resolution plane and depth dimension motion Get the scene flow

[0040] In this embodiment, a brand-new scene flow estimation network model is designed based on 3D motion feature encoding. Through the encoding of 3D motion features, the scene flow estimation network model can uniformly process the three dimensions of the image motion plane and depth from a 2D perspective, and the extracted 3D motion features are highly compatible with the GRU iterative module in the model, capable of fully mining the 3D correlation and motion cues between adjacent frame pairs, thereby greatly improving the accuracy of scene flow estimation. There is no need to design a complex model structure or introduce an additional semantic extraction network, which has significant advantages in terms of inference speed and the number of parameters, combining the characteristics of high precision, high speed, and low number of parameters, so as to improve the estimation performance and computational efficiency in current scene flow estimation.

[0041] In step 4, when calculating the coarse-resolution image plane motion With the coarse-resolution depth dimension motion The data will enter the same GRU predictor network multiple times in an iterative manner. Refer to Figure 2 , taking the j-th iteration as an example, that is, the coarse-resolution image plane motion after the j-th iteration With the coarse-resolution depth dimension motion The calculation process is as follows:

[0042] Step 4.1, based on the coarse-resolution image plane motion after the (j - 1)-th iteration Index the relevant pyramid Y t And the depth map D t+1 To obtain the appearance correlation feature And the depth correlation feature where j = 1 to n, and the initial values of the iteration of the coarse-resolution image plane motion and the coarse-resolution depth dimension motion are both 0, that is

[0043] Step 4.2, based on the appearance correlation feature The depth correlation feature The coarse-resolution image plane motion The coarse-resolution depth dimension motion Obtain the 3D motion feature

[0044] Step 4.3, based on the 3D motion features and the hidden state feature h of the (j - 1)-th iteration t,j-1 , predict the coarse-resolution optical flow residual and the depth motion residual Then the coarse-resolution image plane motion and the coarse-resolution depth dimension motion can be updated to the coarse-resolution image plane motion after the j-th iteration and the coarse-resolution depth dimension motion

[0045] In the specific implementation process of Step 4.1, the appearance-related features are obtained by bilinearly sampling each layer of the correlation pyramid Y t and then concatenating them. Among them, the grid size for sampling in each layer is (2×r + 1) 2 , where r is the neighborhood radius during sampling. The process of obtaining the depth-related features is as follows:

[0046] Bilinearly sample the inverse depth map of the image I t+1 , and the grid size for sampling is also (2×r + 1) 2 ;

[0047] Find the inverse depth of the predicted depth after the (j - 1)-th iteration to obtain the predicted inverse depth after the (j - 1)-th iteration, which is

[0048] Find the residual between the predicted inverse depth after the (j - 1)-th iteration and the bilinearly sampled result and concatenate them to obtain the depth-related features

[0049] Reference Figure 3 is the detailed generation process of the 3D motion features in Step 4.2. Figure 3 In it, conv@n×n,l represents a convolution operation with a kernel size of n and an input channel number of l, C represents the concatenation operation, and R represents a Relu operation.

[0050] In the GRU predictor network, the network updates the hidden state feature by means of convolution operations in combination with the input. For example, the update process of the hidden state feature h

[0051] of the j-th iteration t,j is as follows:

[0052] z t,j = σ(Conv([h t,j-1 ,xt,j ,W z ))

[0053] r t,j = σ(Conv([h t,j-1 ,x t,j ,W r ))

[0054]

[0055]

[0056] where σ is the activation function, x t,j is the concatenation of the context feature and the 3D motion feature , W z , W r , W h are the convolutional weights, ⊙ is the Hadamard product, z t,j is the update gate state, r t,j is the reset gate state, is the current memory state.

[0057] In the training phase of the scene flow estimation network model, two adjacent frames of RGB images are used as a training unit, and at the same time, the optical flow and the inverse depth change of the first frame of RGB image are used as the supervision values. The loss function in the training phase is:[[]]

[0058]

[0059] where L is the loss function, γ is the learning weight used to control the results of early iterations, α is the learning weight used to balance the optical flow and the depth change, is the ground truth of the optical flow, is the coarse-resolution image plane motion at the k-th iteration, is the ground truth of the inverse depth change, is the predicted value of the inverse depth change at the k-th iteration, that is

[0060] The following further illustrates the scene flow estimation method based on 3D motion feature encoding in this embodiment with specific examples.

[0061] The datasets used in the examples are FlyingThings3D and Spring. Among them, FlyingThings3D can provide 80,604 data units for the training of the scene flow estimation network model. Spring provides online evaluation and ranking of the test set on its benchmark dataset website.

[0062] The performance of the algorithm is comprehensively evaluated using the first-frame depth anomaly rate D1, the second-frame matching depth anomaly rate D2, the optical flow anomaly rate Fl, and the scene flow anomaly rate SF. The smaller the values of these metrics, the more accurate the precision. At the same time, the inference speed and the number of network parameters of each method are statistically counted for comparison.

[0063] The Python language and the PyTorch deep learning framework are used to implement the scene flow estimation method based on 3D motion feature encoding in this embodiment. Training and inference are performed using an NVIDIA TITAN RTX GPU, and the number of GRU iterations evaluated on Spring is 24. The C+T training process is adopted, and the C+T process means that the network is trained on the FlyingChairs or FlyingThings3D dataset. The batch size is 8, and the learning rate is 1.25×10 -4 , the number of training steps is 100K times, and γ = 0.85 and α = 1000 are set.

[0064] Table 1 lists the comparison experiment results of the method in this embodiment after the C+T training process and two currently state-of-the-art scene flow methods on the Spring test set. LEAStereo is used to provide the depth values of binocular stereo matching. To ensure fairness, the weights of all baseline methods after the C+T training process are taken for testing.

[0065] Table 1

[0066] Scene flow method D1 D2 F1 SF Inference time Number of parameters RAFT-3D 23.2 73.4 48.1 66.9 311ms 44.5M CamLiFlow 23.1 44.1 24.0 34.2 100ms 7.7M Proposed method 19.9 15.0 9.2 11.4 97ms 7.1M

[0067] It can be seen from the experimental results that compared with other scene flow methods, the method proposed in this embodiment obtains lower values in all dataset metrics, which means that the method proposed in this embodiment has significant performance superiority compared with other methods.

[0068] In addition, a binocular stereo matching algorithm with performance similar to RAFT-3D and CamLiFlow is also used to provide the estimation of depth values, but it has achieved a significant lead in terms of matching depth change, optical flow, and scene flow metrics. Especially in terms of the scene flow SF metric, the method proposed in this embodiment has a 66.7% reduction in error compared with the current best CamLiFlow. This fully demonstrates the advanced performance of the method proposed in this embodiment.

[0069] On the other hand, the inference time and the number of parameters of several methods are also compared. Among them, the inference time is the time when running on a 540P image using an NVIDIA TITAN RTX GPU. It can be seen from the results that the method proposed in this embodiment is superior to other advanced methods in terms of inference speed and the number of parameters.

[0070] Embodiment 2

[0071] Based on the scene flow estimation method based on 3D motion feature encoding in Embodiment 1, this embodiment discloses a scene flow estimation system based on 3D motion feature encoding. Refer to Figure 4 , the scene flow estimation system includes a feature encoder, a context encoder, a correlation pyramid calculation module, a 3D single-resolution iteration module, an upsampling module, and a scene flow estimation module. The scene flow estimation system is used to execute some or all of the steps of the scene flow estimation method in Embodiment 1, thereby realizing scene flow estimation. Specifically:

[0072] The feature encoder is used to extract the image features of RGB images I t 、I t+1 ;

[0073] The context encoder is used to extract the context features of RGB image I t ;

[0074] The correlation pyramid calculation module is used to calculate the per-pixel correlation tensor of the image features and the image features , and pool to generate a four-layer correlation pyramid Y t ;

[0075] The 3D single-resolution iteration module is used to obtain the coarse-resolution image plane motion t after n iterations and the coarse-resolution depth dimension motion according to the correlation pyramid Y t , D t+1 , the context features and the depth map D

[0076] The upsampling module is used to upsample the coarse-resolution image plane motion and the coarse-resolution depth dimension motion to generate the full-resolution plane and depth dimension motion

[0077] The scene flow estimation module is used to obtain the scene flow according to the camera intrinsics and the full-resolution plane and depth dimension motion

[0078] In this embodiment, the specific working processes and working principles of the feature encoder, the context encoder, the correlation pyramid calculation module, the 3D single-resolution iteration module, the upsampling module, and the scene flow estimation module are the same as those of the method in Embodiment 1, so they will not be elaborated in this embodiment.

[0079] The above are only the preferred embodiments of the present invention, and do not thereby limit the patent scope of the present invention. Any equivalent structural transformation made under the inventive concept of the present invention by using the content of the specification and drawings of the present invention, or any direct / indirect application in other related technical fields shall be included within the patent protection scope of the present invention.

Claims

1. A scene flow estimation method based on 3D motion feature encoding, characterized in that, Estimate the scene flow using a scene flow estimation network model based on 3D motion feature encoding, including the following steps: Step 1, obtain two adjacent RGB images , and the corresponding depth maps , ; Step 2, based on the feature encoding and context encoding, obtain the image image features and context features , as well as the image image features ; Step 3, calculate the image features and the image features to generate a per-pixel correlation tensor, and pool it to generate a four-layer correlation pyramid ; Step 4, based on the relevant pyramid and context features as well as the depth map , , combined with 3D motion feature encoding for iterative prediction to obtain the planar motion of the coarse-resolution image after iterations and the motion in the depth dimension of the coarse resolution ; Step 5, perform upsampling on the planar motion of the coarse-resolution image and the depth-dimension motion of the coarse-resolution to generate the full-resolution planar and depth-dimension motion , and based on the camera intrinsics and the full-resolution planar and depth-dimension motion , obtain the scene flow ; In step 4, enter the same GRU prediction network multiple times in an iterative manner. After the th iteration, the motion of the coarse-resolution image plane and the motion of the coarse-resolution depth dimension are calculated as follows: Step 4.1, based on the rough-resolution image plane motion after the th iteration, index the relevant pyramid and the depth map respectively to obtain the appearance correlation features and the depth correlation features , where , and the initial iteration values of the rough-resolution image plane motion and the rough-resolution depth dimension motion are both 0, that is , , and the acquisition process of the depth correlation features is as follows: Perform bilinear sampling on the inverse depth map of the image , where the grid size for sampling is , being the neighborhood radius during sampling; For the predicted depth after the inverse depth is calculated to obtain the predicted inverse depth after the Obtain the residual between the predicted inverse depth after the -th iteration and the bilinear sampling result and splice them, then the depth-related feature is obtained ; Step 4.2, based on the appearance-related features , depth-related features , coarse-resolution image plane motion , coarse-resolution depth dimension motion , to obtain 3D motion features ; Step 4.3, based on the 3D motion features and the hidden state features of the n-th iteration, predict the coarse-resolution optical flow residual and the depth motion residual , and then the coarse-resolution image plane motion , the coarse-resolution depth dimension motion can be updated to the coarse-resolution image plane motion and the coarse-resolution depth dimension motion after the n-th iteration.

2. The scene flow estimation method based on 3D motion feature encoding according to claim 1, characterized in that In step 4.1, the apparent correlation feature is obtained by bilinearly sampling and splicing each layer of the relevant pyramid , where the grid size of each layer of sampling is , being the neighborhood radius during sampling.

3. The method for scene flow estimation based on 3D motion feature encoding according to claim 1 or 2, characterized in that, The update process of the hidden state features in the th iteration is as follows: Among them, is the activation function, is the context feature concatenated with the 3D motion feature ; , , are the convolution weights, is the Hadamard product, is the update gate state, is the reset gate state, is the current memory state.

4. The method for scene flow estimation based on 3D motion feature encoding according to claim 1 or 2, wherein In the training stage of the scene flow estimation network model, use two adjacent frames of RGB images as a training unit, and at the same time use the optical flow and inverse depth change of the first frame of RGB image as the supervision value. The loss function in the training stage is: wherein, is the loss function, is the learning weight for controlling the early iteration results, is the learning weight for balancing the optical flow and depth change, is the ground truth of optical flow, is the coarse-resolution image plane motion at the th iteration, is the ground truth of inverse depth change, is the predicted value of inverse depth change at the th iteration, i.e., .

5. A scene flow estimation system based on 3D motion feature encoding, characterized in that, Based on the scene flow estimation network model, sample the method described in any one of claims 1 to 4 to estimate the scene flow. The scene flow estimation system includes: A feature encoder for extracting the image features of RGB images , ; , ; Context encoder for extracting context features of RGB images of ; A related pyramid calculation module for calculating image features With the image features Of the per-pixel correlation tensor, and pooling to generate a four-layer correlation pyramid ; 3D single-resolution iterative module, used to obtain, according to the relevant pyramid , context features and depth map , , the planar motion of the coarse-resolution image and the depth-dimensional motion of the coarse-resolution after iterations ; An upsampling module for upsampling the planar motion of a low-resolution image and the depth-dimensional motion of the low-resolution image to generate full-resolution planar and depth-dimensional motion ; A scene flow estimation module, which is used to obtain scene flow according to the motion in the full-resolution plane and depth dimension involved in the camera , and obtain the scene flow .