Monocular 3D scene flow estimation method and system combined with object information, and storage medium
By combining object information, using high-quality object masks and feature aggregation modules, monocular 3D scene flow estimation is optimized, the problems of occlusion and boundary impact are solved, and the estimation accuracy and robustness are improved.
Patent Information
- Application Number
- CN202510268360.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-27
AI Technical Summary
The existing monocular 3D scene flow estimation method is susceptible to occlusion and motion boundaries, and fails to make full use of object-level information, resulting in low estimation accuracy and blurred boundaries.
A monocular 3D scene flow estimation method combining object information is adopted to generate high-quality object masks through segmentation models, and the mask feature aggregation module and GRU update network is combined to optimize the estimation of scene flow and disparity.
It effectively solves the fragmentation and boundary blurring problems in scene flow estimation, and improves the accuracy and robustness of scene flow estimation, especially in the processing of dynamic scenes and occlusion areas.
Smart Images

Figure CN120219429A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision, and particularly relates to a monocular 3D scene flow estimation method, system and storage medium that combines object information. Background Art
[0002] Monocular scene flow estimation aims to infer the three-dimensional motion of each pixel from consecutive frames captured by a monocular camera, including optical flow and depth changes, which is a three-dimensional extension of optical flow. Scene flow estimation has very broad application prospects in fields such as autonomous driving, intelligent robots, and target tracking. Early methods were based on energy minimization and geometric optimization, combining optical flow and disparity for joint inference, while in recent years, deep learning methods have greatly improved performance through recursive updates, self-supervised learning, and geometric constraints. Modern methods usually combine an optical flow estimation network (such as RAFT) with a monocular depth estimation network (such as Monodepth2), and simultaneously complete the estimation of depth and motion through an end-to-end learning method. In addition, self-supervised learning methods significantly reduce the dependence on labeled data by combining geometric constraints (such as image reconstruction error, disparity consistency), while improving the estimation accuracy of occluded regions. To meet the real-time requirements, a monocular 3D scene flow estimation model based on a feature pyramid (Self-Mono-SF) uses a joint decoder to solve the monocular scene flow task, which actually simplifies the method of jointly estimating depth and optical flow and directly outputs the scene flow. However, currently in the field of monocular 3D scene flow estimation, these methods are easily affected by occlusions and motion boundaries and do not fully utilize object-level information. Summary of the Invention
[0003] The purpose of the present invention is to provide a monocular 3D scene flow estimation method, system and storage medium that combines object information, solve the problems of fragmentation and blurred boundaries in scene flow estimation, and improve the accuracy of scene flow estimation at the same time.
[0004] The purpose of the present invention is achieved through the following technical solutions:
[0005] A monocular 3D scene flow estimation method that combines object information, the specific steps are as follows:
[0006] Step 1: Prepare a monocular image sequence for training and testing of the monocular scene flow network, including images at times t and t + 1 of the monocular camera; construct a feature encoder, a disparity initialization layer, and a 4D correlation pyramid for extracting matching features and context features, input two consecutive monocular images, and the output is the initialized disparity, 4D correlation volume, and context features corresponding to the first frame image;
[0007] Step 2: Construct a mask feature aggregation module, input two consecutive monocular images, and the output is mask-aligned features;
[0008] Step 3: Construct a GRU update network combined with object masks, aggregate the features of Step 1 and Step 2, encode them as motion features as the input, and synchronously output the scene flow and disparity through two task heads;
[0009] Step 4: Construct the overall network loss function; input two consecutive frames of images at adjacent times and train the network in a self-supervised manner;
[0010] Step 5: Input two consecutive frames of images into the trained model for testing, and the visual output is the estimated scene flow and the depth map of the first frame.
[0011] Furthermore, the images of the monocular camera at times t and t + 1 in Step 1 are two consecutive frames of images I t and I t+1 , and input them into two feature encoders to obtain matching features f1 and f2. The first-frame image I t is input into a context encoder with the same architecture as the feature encoder to obtain context feature f context , where H and W respectively represent the height and width of the input image, and C represents the number of channels of the input image;
[0012] The disparity initialization layer consists of two convolutional layers, and maps the matching feature f1 of the first-frame image to a single-channel disparity map through two convolutional operations:
[0013] d0 = ConV w (f1)
[0014] where d0 represents the initialized disparity, and ConV w represents two convolutional layers;
[0015] According to the matching features f1 and f2, calculate the 4D correlation quantity C ijkh :
[0016] C ijkh (I t , I t+1 ) = <(f1) ij , (f2) kh >
[0017] where < > represents the dot product operation, and ij and kh respectively represent the pixel indices of images I t and I t+1 ;
[0018] Take the 4D correlation quantity C ijkhDesigned to be 4 - dimensional, corresponding to the resolution sizes of the matching features f1 and f2 respectively. Then, the last two dimensions are pooled using pooling layers with kernel sizes of 1, 2, 4, and 8 to obtain a pyramid of relevant quantities {C0, C1, C2, C3} with 4 layers.
[0019] Furthermore, in step 2, using the high - quality object masks provided by the segmentation model, first, the object masks generated by the segmentation model are converted into a full - segmentation representation, where each pixel is uniquely assigned to a mask. For each object mask region, the feature values within the region are aggregated through a max - pooling operation to generate object - level aggregated features. Then, these object features are backfilled into the corresponding mask regions to form an enhanced feature map, which is concatenated with the original feature map in the channel dimension to generate the final mask features.
[0020] Furthermore, step 2 is specifically as follows:
[0021] Use the Segment - Anything Model (SAM) to segment and analyze two consecutive input images I t and I t+1 to generate masks for the corresponding two segmented images, wherein, it contains binary masks of n t objects in image I t ; given masks M1, M2 and matching features f1, f2, corresponding mask features g1 and g2 are generated through a mask feature aggregation module;
[0022] Sort all current object masks according to the size of the mask regions; for pixels belonging to multiple masks, assign them to the mask with the smallest area; for pixels not belonging to any mask, create a new background object mask to cover these pixels; then downsample masks M1 and M2 to reduce the resolution to 1 / 8 to obtain new masks Next, after passing the matching features f1 and f2 through a 1×1 convolutional layer, then perform max - pooling with the segmented regions and the new masks to obtain pooled features Its operation is defined as follows:
[0023]
[0024] where maxpool() represents the max - pooling operation;
[0025] Then, concatenate the features before pooling and the pooled features through a 1×1 convolutional layer to obtain the mask feature g t (t ∈ {1, 2}), and its operation is defined as follows:
[0026]
[0027] Among them, Cat() represents the concatenation operation; then, the masked feature g2 and the optical flow output by the update network are warped and the correlation is calculated with the masked feature g1 corresponding to the first frame image to obtain the masked alignment feature c m 。
[0028] Further, in step 3, the masked alignment feature obtained by the masked aggregation module is concatenated with the context feature and the motion feature in the channel dimension, and then the concatenated feature is used as the input and sent into the GRU update network. The hidden state is updated recursively and the scene flow and the disparity residual are predicted
[0029] Further, step 3 is specifically as follows
[0030] The update network of the output scene flow and depth based on GRU encodes the relevant pyramid C k 、the currently estimated disparity, the optical flow F and the scene flow, and the masked alignment feature c m into the motion feature; in each iteration, the relevant features for residual update are retrieved from the pre-computed 4D correlation pyramid C k ; based on the currently estimated scene flow and disparity, the corresponding pixel position p' of each pixel in the target frame is calculated, and then a group of pixels of the adjacent corresponding pixel p' within the range of [-r, r] is retrieved of the relevant features
[0031] Then GRU takes the concatenation of the motion feature and the context feature as the input; the input hidden state is initialized by the context encoder, where the tanh function is used as the activation function; the currently estimated disparity and scene flow are decoded through the disparity and scene flow heads, and the mask head is used to generate a convex mask for upsampling; three convolutional layers are used to predict the scene flow and the disparity residual respectively; the outputs of the update network are the residual update values of the scene flow and the disparity respectively. The final prediction is the sum of all the residual outputs with the initial values, specifically as follows
[0032] s k+1 = Δs + s k
[0033] d k+1 = Δd + d k
[0034] where s k and d k are the scene flow and the disparity obtained in the current k-th iteration; s k+1 and d k+1 are the estimated values of the scene flow and the disparity after the (k + 1)-th iteration update; Δs and Δd are the residual update amounts of the scene flow and the disparity calculated by the update network in the current iteration step
[0035] Further, the overall network loss function in step 4 is composed of the weighted sum of the depth loss and the scene flow loss, and its specific definition formula is:
[0036] L total = L d + λ sf L sf
[0037] where, L sf is the scene flow loss; L d is the depth loss; λ sf is the weight constant.
[0038] Further, the depth loss L d includes the photometric loss and the smoothness loss, and the regularization constant is 0.1. The specific expression is:
[0039] L d = L d,ph + λ d,sm L d,sm
[0040] where, λ d,sm is the weight constant; the specific expression of the smoothness loss L d,sm is:
[0041]
[0042] where, N is the total number of pixels; p is a pixel coordinate in the image; i∈{x,y} represents the horizontal and vertical directions; is the second-order gradient of the depth map d t in the direction i, measuring the depth change rate of the pixel in this direction; is the weight value based on the gradient of the image I t , represents the edge intensity of the image object, and the parameter β controls the attenuation degree of the weight; when the edge intensity is large, the exponential value of this term is small, reducing the constraint on depth smoothing;
[0043] The specific expression of the photometric loss L d,ph is:
[0044]
[0045] where, I t is the input original image; is the target image reconstructed by the disparity; is the disparity occlusion mask; ρ census () is the occlusion-aware loss function; the given right view is warped backward with the output depth d t to obtain the synthesized left view and calculating the given left view I t and the synthesized left view to calculate the photometric loss based on the photometric difference;
[0046] The parallax occlusion mask is obtained by inputting the right view into the network and forward warping it; the occlusion-aware loss function ρ census is used to calculate the Hamming distance of visible pixels, and the specific expression is:
[0047]
[0048] where I is the reference image; is the reconstructed image; O is the weight mask used to mark the valid pixel region; T(I, p, y) is the structural descriptor of image I at pixel p and its neighborhood y, used to describe the intensity difference between pixel p and its neighborhood p + y, and is normalized to avoid excessive numerical values, σ t is a smoothing term; f G (t1, t2) is used to compare the difference between two structural descriptors t1 and t2. The numerator is the squared difference between the two descriptors, used to measure the local difference, and the denominator introduces the smoothing term σ G to prevent numerical instability;
[0049] The scene flow loss is composed of the photometric loss, the 3D reconstruction loss, and the smoothing loss, and the specific expression is:
[0050] L sf = L sf,ph + λ sf,pt L sf,pt + λ sf,sm L sf,sm
[0051] where λ sf,pt and λ sf,sm are their respective loss constants;
[0052] The specific expression of the scene flow photometric loss L sf,ph is:
[0053]
[0054] where I t is the input original image; is the target image reconstructed through the scene flow; is the scene flow occlusion mask; ρ census () is the occlusion-aware loss function;
[0055] The 3D reconstruction loss L sf,ptIt is to calculate the Euclidean distance for the three-dimensional points of visible pixels, and the specific expression is:
[0056]
[0057] where p′ is the corresponding pixel of the given scene flow and depth estimation; is the depth map at time t; is the depth map at time t + 1; K -1 is the inverse of the camera internal parameter matrix. Using the given camera focal length f focal and the depth estimation of the stereo baseline to convert the depth Assume that the camera focal length and the stereo baseline are given. At this time, the network outputs the depth at a fixed ratio; the 3D distance loss from each point to the camera is normalized to penalize the relative distance to the camera.
[0058] Smoothing loss L sf,sm The specific expression is:
[0059]
[0060] where N is the total number of pixels; p is a pixel coordinate in the image; i∈{x,y} represents the horizontal and vertical directions; is the scene flow The second-order gradient in the direction i is used to measure the smoothness of the scene flow in this direction; is the weight value based on the gradient of the image I t The gradient, represents the edge intensity of the image object, and the parameter β controls the attenuation degree of the weight; ||P t ||2 is the 3D distance normalization term, which is used to normalize the smoothing loss of the scene flow according to the object distance to ensure the consistency of smoothness in 3D space.
[0061] A computer device / system includes a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of a monocular 3D scene flow estimation method combined with object information.
[0062] A computer-readable storage medium stores a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of a monocular 3D scene flow estimation method combined with object information are implemented.
[0063] The beneficial effects of the present invention are as follows:
[0064] In the present invention, the segmentation model selects the latest Segment Anything Model (SAM). The segmentation model in the present invention can generate high-quality object masks, provide accurate boundary and region information, and effectively solve the problems of boundary blur and fragmentation in scene flow estimation. Secondly, through the fine-grained segmentation of object masks, the independent movements of different parts of objects in dynamic scenes can be distinguished, and the motion estimation can be refined. In addition, high-quality masks can also mark occluded areas and untrusted pixels, enhancing the robustness of the model in occluded scenes and dynamic regions. The mask feature aggregation module proposed in the present invention improves the context awareness ability of the model and reduces the estimation error in low-texture or blurred areas. The updated network that simultaneously outputs scene flow and disparity can not only mark occluded areas and blurred areas through residual updates and weight masks, but also iteratively optimize to reduce errors, making the self-supervised loss easier to converge and improving the stability of the training process. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 is a flowchart of the present invention;
[0066] Figure 2 is a structural diagram of a monocular 3D scene flow estimation network;
[0067] Figure 3 is a structural diagram of a mask feature aggregation module;
[0068] Figure 4 is a structural diagram of a GRU update network combined with an object mask. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0069] The present invention will be further described below with reference to the accompanying drawings.
[0070] Embodiment 1:
[0071] The overall process of the present invention is as Figure 1 shown and is specifically implemented through the following steps:
[0072] Step 1: Prepare a monocular image sequence for training and testing a monocular scene flow network, including images at times t and t + 1 of a monocular camera. Construct a feature encoder, a disparity initialization layer, and a 4D correlation pyramid for extracting matching features and context features. Input two consecutive frames of monocular images, and the outputs are the initialized disparity, the 4D correlation volume, and the context features corresponding to the first frame image.
[0073] As Figure 2As shown, the feature encoder network for extracting matching features and context features consists of a two-path weight-sharing feature encoder based on a deep convolutional network and a context feature encoder. The feature encoder consists of 6 residual blocks, 2 with a resolution of 1 / 2, 2 with a resolution of 1 / 4, and 2 with a resolution of 1 / 8. Given two input images, the feature encoder network extracts two types of features: matching features and context features. For two consecutive frames of images I t and I t+1 First, they are input into the two-path feature encoder to obtain matching features f1 and f2. The first frame of image I t is input into the context encoder with the same architecture as the feature encoder to obtain context feature f context , where H and W represent the height and width of the input image respectively, and C represents the number of channels of the input image.
[0074] The disparity initialization layer consists of two convolutional layers. Through two convolutional operations, the matching feature f1 of the first frame of image is mapped into a single-channel disparity map. The size of the convolutional kernel of the first layer is 3×3, the number of input channels is C, and the number of output channels is Using the ReLU activation function, the high-dimensional matching features are compressed to low dimensions, and the feature size is The size of the convolutional kernel of the second layer is 3×3, the number of input channels is The number of output channels is 1, and the prediction value is directly output without using the ReLU activation function, mapping the features of each pixel point into a single-channel disparity map with an unchanged resolution. Its operation is defined as follows:
[0075] d0 = ConV w (f1)
[0076] where d0 represents the initialized disparity, with a size of ConV w represents two convolutional layers.
[0077] The 4D correlation pyramid represents the visual similarity between two frames of images by calculating the complete correlation quantity between the matching features f1 and f2. Given the matching features f1 and f2, the correlation quantity is obtained by calculating the dot product between all vector pairs of the two feature maps. The calculation process is as follows:
[0078] C ijkh (I t ,I t+1 ) = <(f1) ij ,(f2) kh >
[0079] where C ijkh is the 4D correlation quantity, and its size is < > represents the dot product operation, and ij and kh represent the pixel indices of images I t and I t+1 respectively.
[0080] Then, the 4D correlation quantity is pooled to construct a correlation pyramid. The 4D correlation quantity C ijkh is designed to be 4-dimensional, corresponding to the resolution sizes of the matching features f1 and f2 respectively. Then, the last two dimensions are pooled with pooling layers with kernel sizes of 1, 2, 4, and 8 to obtain a correlation quantity pyramid {C0, C1, C2, C3} with 4 layers. The size of each layer of the correlation quantity C k is
[0081] Step 2: Construct a mask feature aggregation module. Input two consecutive frames of monocular images, and the output is mask-aligned features.
[0082] Two consecutive frames of images I t and I t+1 are first input into a feature encoder composed of convolutions to obtain matching features f1, f2, and context features f of the same size, where H and W represent the height and width of the input image respectively, and C represents the number of channels of the input image. For the two consecutive input images I context and I t and I t+1 the Segment Anything Model (SAM) is used to segment and analyze the images, generating masks for the corresponding two frames of segmented images which contain binary masks of n t objects in image I t . These masks are generated by prompting with approximately 1000 points arranged on the image and combined with non-maximum suppression to remove redundant masks. As Figure 3 shown, given masks M1, M2 and matching features f1, f2, the corresponding mask features g1 and g2 are generated through the mask feature aggregation module. To better use the masks generated by the segmentation model in the network, masks M1 and M2 are converted into a complete segmentation representation form, that is, each pixel is precisely assigned to a mask. First, all current object masks are sorted according to the size of the mask area. For pixels belonging to multiple masks, they are assigned to the mask with the smallest area. For pixels not belonging to any mask, a new "background" object mask is created to cover these pixels. Then, masks M1 and M2 are downsampled to reduce the resolution to 1 / 8 to obtain new masks Next, the matching features f1 and f2 pass through a 1×1 convolutional layer, and then the pooling features are obtained through max pooling with the segmentation region and the new masks whose operation is defined as follows:
[0083]
[0084] Among them, maxpool() represents the max pooling operation.
[0085] Then, the features before pooling and the pooled features are concatenated through a 1×1 convolutional layer to obtain the masked feature g t (t ∈ {1, 2}), and its operation is defined as follows:
[0086]
[0087] Among them, Cat() represents the concatenation operation; then, the masked feature g2 and the optical flow output by the update network are warped and the correlation with the masked feature g1 corresponding to the first frame image is calculated to obtain the masked alignment feature c m .
[0088] Step 3: Construct a GRU update network that combines object masks, aggregate the features of Step 1 and Step 2, encode them into motion features as inputs, and synchronously output the scene flow and disparity through two task heads.
[0089] Calculate the per-pixel correlation tensor of the matching features f1 and f2 of two frames of images, and pool to generate a multi-layer 4D correlation pyramid C k . The matching feature f1 obtains the initial disparity d0 through the disparity initialization layer. As Figure 4 shown, the update network for the output scene flow and depth based on GRU encodes the correlation pyramid C k , the current estimated disparity, optical flow and scene flow, and the masked alignment feature c m into the motion features. In each iteration, the correlation features for residual update are retrieved from the pre-computed 4D correlation pyramid C k . Then, based on the current estimated scene flow and disparity, calculate the corresponding pixel position of each pixel in the target frame, and the specific operation is as follows:
[0090] p′ = K(d t (p) · K -1 p + s t→t+1 (p))
[0091] where p′ is the coordinate of the pixel in the target frame; p is the coordinate of the pixel in the current frame; K is the internal parameter matrix of the camera, which is used to convert the pixel coordinates to the camera coordinate system; K -1 is the inverse of the camera internal parameter matrix, which is used to calculate the camera coordinates from the pixel coordinates; d t (p) is the disparity value of the pixel p in the current frame; s t→t+1 (p) is the estimated scene flow value of the pixel p in the current frame, indicating the three-dimensional displacement of the pixel from the current frame to the target frame.
[0092] Then, a set of pixels of the adjacent corresponding pixel p' is retrieved within the range of [-r, r]. The relevant features are as follows:
[0093]
[0094] Among them, and are offsets, representing displacements in the x and y directions respectively.
[0095] The input optical flow can be obtained by taking the difference between the pixel coordinates of the current frame and the target frame. The specific operations are as follows:
[0096] F = p' - p
[0097] Among them, F is the estimated optical flow, and the size is
[0098] Given the current estimate and the retrieved relevant features, they are passed through the motion feature encoder to obtain motion features. Then GRU takes the concatenation of the motion features and the context features (from the context encoder) as input. The input hidden state is initialized by the context encoder, where the tanh function is used as the activation function. The current estimated disparity and scene flow are decoded through the disparity and scene flow heads, and the mask head is used to generate a convex mask for upsampling. Three convolutional layers are used to predict the scene flow and the disparity residual respectively. The outputs of the updated network are the residual update values of the scene flow and the disparity respectively. The final prediction is the sum of all the residual outputs with the initial values. The specific operations are as follows:
[0099] s k+1 = Δs + s k
[0100] d k+1 = Δd + d k
[0101] Among them, s k and d k are the scene flow and the disparity obtained in the current k-th iteration; s k+1 and d k+1 are the estimated values of the scene flow and the disparity after the (k + 1)-th iteration update; Δs and Δd are the residual update amounts of the scene flow and the disparity calculated by the updated network in the current iteration step.
[0102] Step 4: Input two consecutive frames of images at the input end of the network, and perform supervised training on the network using the overall network loss function.
[0103] The overall network loss function is composed of the weighted sum of the depth loss and the scene flow loss. Its specific definition formula is:
[0104] Ltotal = L d + λ sf L sf
[0105] where L sf is the scene flow loss; L d is the depth loss; λ sf is the weight constant.
[0106] For the depth loss, this network uses the perspective of the stereo image pair as the guidance for depth estimation during the training phase. To achieve accurate depth estimation, the depth loss adopted by this network includes photometric loss and smoothness loss, and the regularization constant is 0.1. The specific expression is:
[0107] L d = L d,ph + λ d,sm L d,sm
[0108] where L d,sm is the smoothness loss; L d,ph is the photometric loss; λ d,sm is the weight constant.
[0109] The specific expression of the smoothness loss is:
[0110]
[0111] where N is the total number of pixels; p is the coordinate of a pixel in the image; i∈{x,y} denotes the horizontal and vertical directions; is the second-order gradient of the depth map d t in the direction i, measuring the depth change rate of the pixel in this direction; is the weight value based on the gradient of the image I t denotes the edge intensity of the image object, and the parameter β controls the attenuation degree of the weight. When the edge intensity is large (such as the object boundary), the exponential value of this term is small, reducing the constraint on depth smoothness.
[0112] The specific expression of the photometric loss is:
[0113]
[0114] where I t is the input original image; is the target image reconstructed by the disparity; is the disparity occlusion mask; ρ census () is the occlusion-aware loss function. Given the right view and the output depth d tBackward warping yields the synthesized left view And by computing the given left view I t and the synthesized left view to calculate the photometric loss by their photometric difference. The disparity occlusion mask is obtained by inputting the right view into the network and forward warping it. The occlusion-aware loss function ρ census is used to calculate the Hamming distance of visible pixels, and the specific expression is:
[0115]
[0116] where I is the reference image; is the reconstructed image; O is the weight mask used to mark the valid pixel region; T(I, p, y) is the structural descriptor of image I at pixel p and its neighborhood y, used to describe the intensity difference between pixel p and its neighborhood p + y, and is normalized to avoid excessive numerical values, σ t is a smoothing term; f G (t1, t2) is used to compare the difference between two structural descriptors t1 and t2. The numerator is the squared difference between the two descriptors, used to measure the local difference, and the denominator introduces the smoothing term σ G to prevent numerical instability. The occlusion-aware loss function ρ census The overall idea is to calculate the local structural descriptor T(I, p, y) for the 7×7 neighborhood (y ∈ [-3, 3] 2 ) around pixel p, and then compare the structural similarity between I and G using the f function. The denominator uses the weight mask O to perform weighted summation on the valid region, and finally is constrained by normalization.
[0117] The scene flow loss adopted by this network is composed of photometric loss, 3D reconstruction loss, and smoothing loss, and the specific expression is:
[0118] L sf = L sf,ph + λ sf,pt L sf,pt + λ sf,sm L sf,sm
[0119] where L sf,ph is the scene flow photometric loss; L sf,pt is the 3D reconstruction loss; L sf,sm is the smoothing loss; λ sf,pt and λ sf,sm are the respective loss constants.
[0120] The specific expression of the scene flow photometric loss is:
[0121]
[0122] Among them, I t is the input original image; is the target image reconstructed by scene flow; is the scene flow occlusion mask; ρ census () is the occlusion-aware loss function. First, using the internal matrix K of the camera and the estimated depth d t and the scene flow to obtain the synthetic reference image Then, given the reference image I t and the synthetic reference image compensate for the scene flow smoothness loss. The scene flow occlusion mask O t sf is obtained by de-occluding the backward scene flow Finally, use the occlusion-aware loss function ρ census to calculate the scene flow photometric loss.
[0123] The 3D reconstruction loss is the calculation of the Euclidean distance applied to the 3D points of the visible pixels, and the specific expression is:
[0124]
[0125]
[0126] Among them, p′ is the corresponding pixel given the scene flow and depth estimation; is the depth map at time t; is the depth map at time t + 1; K -1 is the inverse of the camera internal parameter matrix. Using the given camera focal length f focal and the depth estimation of the stereo baseline to convert the depth Assume that the camera focal length and the stereo baseline are given. At this time, the network can output the depth at a fixed ratio. Normalize the 3D distance loss of each point to the camera to penalize the relative distance to the camera.
[0127] The specific expression of the smoothness loss is:
[0128]
[0129] Among them, N is the total number of pixels; p is a pixel coordinate in the image; i∈{x,y} represents the horizontal and vertical directions; is the scene flow at the second-order gradient in the direction i, used to measure the smoothness of the scene flow in this direction; is based on the image I tThe weight value of the gradient represents the edge intensity of the image object, and the parameter β controls the attenuation degree of the weight; ||P t ||2 is a 3D distance normalization term, which is used to normalize the smoothness loss of the scene flow according to the object distance to ensure the consistency of smoothness in 3D space.
[0130] Step Five: Input two consecutive frames of images into the trained model for testing, and visually output the corresponding scene flow and the depth map of the first frame.
[0131] Embodiment 2:
[0132] A monocular 3D scene flow estimation method combining object information according to the present invention, the network architecture includes: a feature encoder for extracting matching features and context features, a disparity initialization layer and a 4D correlation pyramid, a mask feature aggregation module, and an update network based on GRU for outputting scene flow and depth.
[0133] The Segment Anything Model (SAM) adopted in the present invention is a general image segmentation model, which is pre-trained on a very large and diverse dataset. It can distinguish different instances and exhibits impressive zero-shot performance on objects not seen during training. In addition, SAM can detect objects at different scales and levels, including segmenting small object parts (such as hands and arms). This can reduce complexity and help distinguish the movements of different parts of the object separately. High-quality object masks are not limited to being generated using the SAM model, and object masks generated by other segmentation models are also applicable to the present invention and are within the protection scope of the present invention.
[0134] In the present invention, the segmentation model can provide accurate object boundary and region information, and can distinguish different object instances and parts even in complex scenes, significantly improving the accuracy and robustness of scene segmentation. For this reason, the present invention designs a mask feature aggregation module and an update network that simultaneously outputs scene flow and disparity to make full use of the advantages of object masks.
[0135] The core of the mask feature aggregation module is to closely combine the object mask provided by the segmentation model with the matching features. By integrating and optimizing the features in the mask region, it solves the common fragmentation and boundary blur problems in traditional scene flow estimation. Especially when dealing with dynamic scenes, object occlusion, or low-texture regions, this module can more accurately express the correlation relationship and boundary information between pixels.
[0136] The update network that simultaneously outputs the scene flow and the disparity iteratively updates the scene flow and the disparity. By fully utilizing the relationship between the previous frame and the current frame, it gradually optimizes complex scenes and finally converges to a more accurate estimated value. In this way, the accuracy of scene flow estimation and the robustness to complex scenes are effectively improved, providing a novel and effective solution for the monocular scene flow task.
[0137] In summary, the present invention uses the high-quality object masks provided by a segmentation model (such as the Segment Anything Model SAM) to provide accurate object boundary and region information, and proposes a mask feature aggregation module and an update network with a mask head. The mask feature aggregation module combines the object mask with the matching features to solve the fragmentation and boundary blur problems in scene flow estimation, and then the update network that outputs the scene flow and the disparity iteratively updates the scene flow and the disparity to improve the accuracy of scene flow estimation.
[0138] Embodiment 3: The present invention provides a computer-readable storage medium storing one or more programs;
[0139] The one or more programs include instructions;
[0140] Preferably, when the instructions are executed by a computing device, the computing device is used to execute the steps of implementing the above data classification and grading method.
[0141] Embodiment 4: The present invention provides a computer device, which is characterized by including: one or more processors and a memory;
[0142] One or more computer programs are stored in the memory and are configured to be executed by the one or more processors;
[0143] Preferably, the one or more computer programs are used to execute the steps of implementing the above data classification and grading method.
[0144] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A monocular 3D scene flow estimation method combining object information, characterized by: The specific steps are as follows: Step 1: Prepare a monocular image sequence for training and testing the monocular scene flow network, including images of the monocular camera at time t and t+1; construct a feature encoder, disparity initialization layer, and 4D correlation pyramid to extract matching features and context features, input two consecutive frames of monocular images, and output the initialization disparity, 4D correlation volume, and context features corresponding to the first frame of the image; Step 2: Construct a mask feature aggregation module, input two consecutive monocular images, and output mask alignment features; Step 3: Construct a GRU update network combined with object mask, aggregate the features of step 1 and step 2, encode them into motion features as input, and synchronously output scene flow and disparity through two task heads; Step 4: Construct the overall network loss function; input two consecutive frames of images at adjacent moments to train the network in a self-supervised manner; Step 5: Input two consecutive frames of images into the trained model for testing, and the visual output is the estimated scene flow and the depth map of the first frame.
2. The monocular 3D scene flow estimation method combined with object information according to claim 1, characterized in that: The image of the monocular camera at time t and t+1 in step 1 is two consecutive frames of image I t and I t+1 , input two feature encoders to get, Matching features f1 and f2, the first frame image I t Input the context encoder with the same architecture as the feature encoder to get Context feature f context , where H and W represent the height and width of the input image respectively, and C represents the number of channels of the input image; The disparity initialization layer consists of two convolutional layers, which map the matching feature f1 of the first frame image into a single-channel disparity map through two convolution operations: d0=ConV w (f1) Among them, d0 represents the initialization disparity, ConV w Represents two convolutional layers; According to the matching features f1 and f2, calculate the 4D correlation C ijkh : C ijkh (I t ,I t+1 )=<(f1) ij ,(f2) kh > Among them, <> represents the dot product operation, ij and kh represent the image I respectively. t and I t+1 The pixel index of The 4D correlation quantity C ijkh It is designed to be 4-dimensional, corresponding to the resolution size of the matching features f1 and f2, and then the last two dimensions are pooled with pooling layers with kernel sizes of 1, 2, 4 and 8 to obtain a correlation pyramid {C0, C1, C2, C3} with 4 layers.
3. The monocular 3D scene flow estimation method in combination with object information according to claim 1, characterized in that: The step 2 utilizes the high-quality object mask provided by the segmentation model. First, the object mask generated by the segmentation model is converted into a full segmentation representation, and each pixel is uniquely assigned to a mask; for each object mask area, the feature values in the area are aggregated through a maximum pooling operation to generate object-level aggregated features; then these object features are backfilled into the corresponding mask area to form an enhanced feature map, and are spliced with the original feature map in the channel dimension to generate the final mask feature.
4. The monocular 3D scene flow estimation method in combination with object information according to claim 3, characterized in that: The step 2 is specifically as follows: Use the segmentation model SAM to input two consecutive frames of image I t and I t+1 Perform segmentation and analysis to generate masks corresponding to the two frames of images after segmentation It contains the image I t Medium t Binary mask of an object; given masks M1, M2 and matching features f1, f2, the corresponding mask features g1 and g2 are generated through the mask feature aggregation module; Sort all current object masks by the size of the mask area; for pixels that belong to multiple masks, assign them to the mask with the smallest area; for pixels that do not belong to any mask, create a new background object mask to cover these pixels; then downsample masks M1 and M2 to reduce the resolution to 1 / 8 to obtain a new mask Then, the matching features f1 and f2 are passed through a 1×1 convolutional layer, and then the pooled features are obtained by performing maximum pooling on the segmented area and the new mask. Its operation is defined as follows: Among them, maxpool() represents the maximum pooling operation; Then the features before pooling are concatenated with the pooling features through a 1×1 convolution layer to obtain the mask feature g t (t∈{1,2}), its operation is defined as follows: Among them, Cat() represents the splicing operation; then the mask feature g2 and the optical flow output by the update network are warped and the correlation is calculated with the mask feature g1 corresponding to the first frame image to obtain the mask alignment feature c m .
5. The monocular 3D scene flow estimation method in combination with object information according to claim 1, characterized in that: The step 3 concatenates the mask alignment features obtained by the mask aggregation module with the context features and the motion features in the channel dimension, and then sends the concatenated features as input to the GRU update network, updates the hidden state through recursive optimization, and predicts the scene flow and disparity residual.
6. The monocular 3D scene flow estimation method in combination with object information according to claim 5, characterized in that: The step 3 is specifically as follows: The GRU-based output scene flow and depth update network transforms the relevant pyramid C through the motion feature encoder k , the current estimated disparity, optical flow F and scene flow, mask alignment feature c m Encoded into motion features; In each iteration, from the precomputed 4D correlation pyramid C k Retrieve the relevant features of the residual update from ; Based on the currently estimated scene flow and disparity, calculate the corresponding pixel position p′ of each pixel in the target frame, and then retrieve a set of pixels of the adjacent corresponding pixel p′ in the range of [-r, r] relevant characteristics of Then the GRU takes the concatenation of motion features and context features as input; the input hidden state is initialized by the context encoder, with the tanh function as the activation function; the currently estimated disparity and scene flow are decoded by the disparity and scene flow heads, and the mask header is used to generate a convex mask for upsampling; three convolutional layers are used to predict the scene flow and disparity residuals respectively; the outputs of the update network are the residual update values of the scene flow and disparity respectively. The final prediction is the sum of all residual outputs with initial values, as follows: s k+1 =Δs+s k d k+1 =Δd+d k Among them, s k and d k is the scene flow and disparity obtained at the current k-th iteration; s k+1 and d k+1 are the scene flow and disparity estimates after the k+1th iteration update; Δs and Δd are the residual updates of the scene flow and disparity calculated by the update network in the current iteration step.
7. The monocular 3D scene flow estimation method in combination with object information according to claim 1, characterized in that: The overall network loss function in step 4 is composed of the weighted combination of depth loss and scene flow loss, and its specific definition is: L total =L d +λ sf L sf Among them, L sf is the scene flow loss; L d is the depth loss; sf is the weight constant.
8. The monocular 3D scene flow estimation method in combination with object information according to claim 7, characterized in that: The depth loss L d Including photometric loss and smoothness loss, the regularization constant is 0.1, and the specific expression is: L d =L d,ph +λ d,sm L d,sm Among them, λ d,sm is the weight constant; smoothing loss L d,sm The specific expression is: Where N is the total number of pixels; p is the coordinate of a pixel in the image; i∈{x,y} Indicates horizontal and vertical directions; is the depth map d t The second-order gradient in direction i measures the rate of change of the depth of the pixel in that direction; Based on image I t The weight value of the gradient, ▽ i I t (p) represents the edge strength of the image object, and the parameter β controls the attenuation degree of the weight; when the edge strength is large, the exponential value of this item is small, reducing the constraints on deep smoothing; Luminosity loss L d,ph The specific expression is: Among them, I t is the original input image; is the target image reconstructed by disparity; is the parallax occlusion mask; ρ census () is the occlusion perception loss function; the given right view With the output depth d t Reverse warping to get the composite left view And by calculating the given left view I t With the composite left view The luminosity loss is calculated by the luminosity difference; The parallax occlusion mask is by placing the right view Input into the network and forward warp it to obtain the occlusion-aware loss function ρ census Used to calculate the Hamming distance of visible pixels. The specific expression is: Where, I is the reference image; is the reconstructed image; O is the weight mask used to mark the valid pixel area; T(I,p,y) is the structural descriptor of image I at pixel p and its neighborhood y, which is used to describe the intensity difference between pixel p and neighborhood p+y and is standardized to avoid excessive values, σ t is a smooth term; f G (t1, t2) is the difference between two structural descriptors t1 and t2. The numerator is the square difference between the two descriptors, which is used to measure the local difference. The denominator is obtained by introducing a smoothing term σ G , to prevent numerical instability; The scene flow loss is composed of photometric loss, 3D reconstruction loss and smoothness loss. The specific expression is: L sf =L sf,ph +λ sf,pt L sf,pt +λ sf,sm L sf,sm Among them, λ sf,pt and λ sf,sm are the respective loss constants; Scene flow loss L sf,ph The specific expression is: Among them, I t is the original input image; is the target image reconstructed by scene flow; is the scene flow occlusion mask; ρ census () is the occlusion perception loss function; 3D reconstruction loss L sf,pt It is the calculation of the Euclidean distance of the three-dimensional points of the visible pixels. The specific expression is: Where p′ is the corresponding pixel of the given scene flow and depth estimation; is the depth map at time t; is the depth map at time t+1; K -1 is the inverse of the camera intrinsic parameter matrix. Using the given camera focal length f focal and depth estimation of stereo baseline to convert depth Assuming the camera focal length and stereo baseline are given, the network outputs depth at a fixed scale; the 3D distance loss from each point to the camera is normalized to penalize the relative distance to the camera. Smoothing loss L sf,sm The specific expression is: Where N is the total number of pixels; p is the coordinate of a pixel in the image; i∈{x,y} Indicates horizontal and vertical directions; For scene flow The second-order gradient in direction i is used to measure the smoothness of the scene flow in that direction; Based on image I t The weight value of the gradient, ▽ i I t (p) represents the edge strength of the image object, and the parameter β controls the attenuation degree of the weight; ||P t ||2 is the 3D distance normalization term, which is used to normalize the smoothness loss of the scene flow according to the object distance to ensure consistent smoothness in the 3D space.
9. A computer device / equipment / system comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Cited By
Dynamic scene depth estimation method and device based on scene flow
CN120976284A
A method and apparatus for dynamic scene depth estimation based on scene flow
CN120976284B