Methods, apparatus, equipment, and storage media for estimating depth and self-motion trajectory

By combining depth estimation networks, motion estimation networks, and implicit cue networks, static and dynamic features between video frames are extracted, solving the problems of moving object artifacts and pose transformation errors in existing methods, and achieving more accurate camera self-motion trajectory and depth estimation.

CN115953468BActive Publication Date: 2026-06-02AGRICULTURAL BANK OF CHINA

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
AGRICULTURAL BANK OF CHINA
Filing Date
2022-12-09
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing depth and self-motion trajectory estimation methods do not fully utilize the dynamic features between video frames, leading to artifacts of moving objects. Furthermore, motion estimation networks struggle to predict pose changes when the camera is stationary, failing to effectively constrain scene depth consistency and resulting in pose change errors and unclear predictions of moving object edges.

Method used

By combining depth estimation networks, motion estimation networks, and implicit cue networks, geometric constraints are enhanced by extracting static and dynamic features between the source and target views. Further constraints are then applied using 3D reconstruction loss, resulting in more accurate camera trajectory estimation.

Benefits of technology

It effectively alleviates the artifact problem of moving objects, improves the quality of depth estimation, reduces pose transformation error, and achieves more accurate camera self-motion trajectory estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115953468B_ABST
    Figure CN115953468B_ABST
Patent Text Reader

Abstract

This invention discloses a method, apparatus, device, and storage medium for estimating depth and self-motion trajectory. The method includes: acquiring a preset model, a source view, and a target view; the preset model includes a depth estimation network, a motion estimation network, and an implicit cue network; the implicit cue network is used to extract static and dynamic features between the source view and the target view from the motion estimation network and identity-map them to the depth estimation network; the source view and the target view are two adjacent color images; inputting the source view and the target view into the preset model; estimating the camera's self-motion trajectory based on the motion estimation network; and estimating the depth of the source view and / or the target view based on the depth estimation network and the static and dynamic features between the source view and the target view. The solution provided by this invention can effectively alleviate the artifact problem of moving objects, improve the estimation quality of monocular image depth, and reduce pose transformation errors, thereby achieving more accurate estimation of the camera's self-motion trajectory.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and in particular to methods, apparatus, devices, and storage media for estimating depth and self-motion trajectories. Background Technology

[0002] The depth of the view and the camera's self-motion trajectory are crucial for understanding geometric scenes from videos or images, and are widely used in fields such as robot vision navigation, autonomous driving, and intelligent transportation applications.

[0003] Existing methods for estimating depth and self-motion trajectories are typically based on depth estimation networks and motion estimation networks. The depth estimation network aims to estimate the depth map of the target view, while the motion estimation network aims to estimate the pose transformation matrix of the source view relative to the target view. Based on the camera model, camera parameters, the depth map of the target view, and the pose transformation matrix of the source view relative to the target view, a reconstructed source view is formed. Then, the structural similarity, L1 loss function, and smoothness loss function between the source view and the reconstructed view are calculated, thereby simultaneously optimizing the depth estimation network and the motion estimation network.

[0004] However, existing methods only perform depth prediction on a single image, which is static and singular, failing to fully utilize the rich dynamic and static features between video frames for depth estimation. Furthermore, according to camera imaging principles, multiple real-world scenes can be projected onto the same pixel plane, resulting in different depth values. The aforementioned methods, using only pixel-level view reconstruction loss, cannot effectively constrain consistent scene depth. Due to the joint training of motion estimation and depth estimation networks, inconsistent scene depths will cause inconsistent transformation vectors, leading to shifts in video frame pose transformations. Additionally, existing methods only use camera pose transformation information between video frames, without fully considering other implicit cues: if the source and target views are the same view (i.e., the camera is stationary), the motion estimation network will struggle to predict camera pose transformations; conversely, if a moving object exists alongside the moving camera capturing it, the moving object violates the scene assumptions of static scene reconstruction, and the edges of the moving object are not effectively constrained, resulting in a single view failing to predict a clear outline of the moving object. Summary of the Invention

[0005] This invention provides a method, apparatus, device, and storage medium for estimating depth and self-motion trajectory, which can effectively alleviate the artifact problem of moving objects, improve the estimation quality of monocular image depth, reduce pose transformation error, and achieve more accurate estimation of camera self-motion trajectory.

[0006] According to one aspect of the present invention, a method for estimating depth and self-motion trajectory is provided, comprising:

[0007] Obtain a preset model, a source view, and a target view. The preset model includes a depth estimation network, a motion estimation network, and an implicit cue network. The implicit cue network is used to extract static and dynamic features between the source view and the target view from the motion estimation network and identity map them to the depth estimation network. The source view and the target view are two color images at adjacent time points.

[0008] Input the source view and target view into the preset model;

[0009] Based on a motion estimation network, the camera's self-motion trajectory is estimated; and based on a depth estimation network and static and dynamic features between the source view and the target view, the depth of the source view and / or the target view is estimated.

[0010] Optionally, before obtaining the preset model, the following steps are also included:

[0011] Obtain the first training view and the second training view;

[0012] The preset model is trained based on the first training view and the second training view.

[0013] Optionally, the preset model is trained based on the first training view and the second training view, including:

[0014] Input the first training view and the second training view into the preset model to obtain the first reconstructed view and the second reconstructed view;

[0015] Based on the first training view, the second training view, the first reconstruction view, and the second reconstruction view, determine the reprojection loss and the smoothness loss;

[0016] Determine the loss in 3D reconstruction;

[0017] Based on the reprojection loss, smoothness loss, and 3D reconstruction loss, the preset model is trained using backpropagation and gradient descent principles.

[0018] Optionally, a first reconstructed view and a second reconstructed view are obtained, including:

[0019] Determine the first training view I respectively s Relative to the second training view I t The first pose transformation matrix T t->s Second training view I t Relative to the first training view I s The second pose transformation matrix T s->t Depth map D of the first training view s Depth map D of the second training view t ;

[0020] Based on camera parameters K and the depth map D of the second training viewt Determine the coordinates P of the first spatial point cloud of the second training view in the camera coordinate system. t ;

[0021] Based on the coordinates P of the first spatial point cloud t and the first pose transformation matrix T t->s Based on the principle of Euclidean transformation of coordinate systems, the coordinates P of the second spatial point cloud in the coordinate system where the first training view is located are determined. s ';

[0022] Based on the camera model, the second spatial point cloud is projected and sampled to obtain the first reconstructed view;

[0023] Based on camera parameters K and the depth map D of the first training view s Determine the coordinates P of the first training view in the third spatial point cloud within the camera coordinate system. s ;

[0024] Based on the coordinates P of the third space point cloud s Second pose transformation matrix T s->t Based on the principle of Euclidean transformation of coordinate systems, the coordinates P of the fourth spatial point cloud in the coordinate system where the first training view and the second training view are located are determined. t ';

[0025] Based on the camera model, the fourth spatial point cloud is projected and sampled to obtain the second reconstructed view.

[0026] Optionally, determine the 3D reconstruction loss, including:

[0027] Based on the coordinates P of the first spatial point cloud t The coordinates P of the second spatial point cloud s ', Coordinates P of the third-space point cloud s The coordinates P of the fourth space point cloud t ', Determine the 3D reconstruction loss L=|P t '-P t |+|P s '-P s |

[0028] Optional, the first pose transformation matrix T t->s =PoseNet(I s ,I t );

[0029] Second pose transformation matrix T s->t =PoseNet(I t ,I s );

[0030] Depth map D of the first training views =DepthNet_Decoder(ICNet(PoseNet_Encoder(I t ,I s ))+DepthNet_Encoder(I s )));

[0031] Depth map D of the second training view t =DepthNet_Decoder(ICNet(PoseNet_Encoder(I s ,I t ))+DepthNet_Encoder(I t )));

[0032] The coordinates P of the first spatial point cloud t =K -1 D t I t ;

[0033] The coordinates P of the second spatial point cloud s '=T t->s P t ;

[0034] The coordinates P of the third-space point cloud s =K -1 D s I s ;

[0035] The coordinates P of the fourth space point cloud t '=T s->t P s ;

[0036] Wherein, ICNet represents the Implicit Cue Network, PoseNet represents the Motion Estimation Network, PoseNet_Encoder represents the encoder of the Motion Estimation Network, DepthNet_Encoder represents the encoder of the Depth Estimation Network, and DepthNet_Decoder represents the decoder of the Depth Estimation Network.

[0037] Optionally, the first training view and the second training view are two color images at adjacent time points; or, the second training view is a color image of an adjacent frame generated based on the first training view to simulate the first training view.

[0038] According to another aspect of the present invention, a depth and self-motion trajectory estimation apparatus is provided, comprising: an acquisition module and an estimation module; wherein,

[0039] The acquisition module is used to acquire a preset model, a source view, and a target view. The preset model includes a depth estimation network, a motion estimation network, and an implicit cue network. The implicit cue network is used to extract static and dynamic features between the source view and the target view from the motion estimation network and identity map them to the depth estimation network. The source view and the target view are two color images at adjacent time points.

[0040] The estimation module is used to input the source view and the target view into a preset model; estimate the camera's self-motion trajectory based on a motion estimation network; and estimate the depth of the source view and / or the target view based on a depth estimation network and static and dynamic features between the source view and the target view.

[0041] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0042] At least one processor; and

[0043] A memory that is communicatively connected to at least one processor; wherein,

[0044] The memory stores a computer program that can be executed by at least one processor, such that the at least one processor is able to perform the depth and self-motion trajectory estimation method according to any embodiment of the present invention.

[0045] According to another aspect of the present invention, a computer-readable storage medium is provided, which stores computer instructions for causing a processor to execute and implement a method for estimating depth and self-motion trajectory according to any embodiment of the present invention.

[0046] The technical solution of this invention, through the design of a preset model, includes a depth estimation network, a motion estimation network, and an implicit cue network. The implicit cue network extracts static and dynamic features between the source view and the target view from the motion estimation network and maps them identically to the depth estimation network. This supplements the depth information of a single static view, enhances the geometric constraints of static objects, and compensates for the dynamic features of moving objects, thereby effectively alleviating the artifact problem of moving objects and improving the depth estimation quality of the source view and / or the target view. At the same time, the 3D reconstruction loss of this invention further constrains the view reconstruction from the perspective of spatial point cloud, so that the camera transformation process has consistent depth and pose transformation, effectively reducing pose transformation error, and making the cumulative offset of the motion trajectory predicted in long videos smaller, so as to achieve more accurate estimation of the camera's self-motion trajectory.

[0047] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0048] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0049] Figure 1 This is a flowchart illustrating a method for estimating depth and self-motion trajectory provided in Embodiment 1 of the present invention;

[0050] Figure 2 This is a flowchart illustrating a method for estimating depth and self-motion trajectory provided in Embodiment 2 of the present invention;

[0051] Figure 3 This is a schematic diagram of the structure of a depth and self-motion trajectory estimation device provided in Embodiment 3 of the present invention;

[0052] Figure 4 This is a schematic diagram of another depth and self-motion trajectory estimation device provided in Embodiment 3 of the present invention;

[0053] Figure 5 This is a schematic diagram of the structure of an electronic device provided in Embodiment 4 of the present invention. Detailed Implementation

[0054] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0055] It should be noted that the terms "first," "second," "third," "fourth," "source," "target," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0056] Example 1

[0057] Figure 1 This is a flowchart illustrating a method for estimating depth and self-motion trajectory according to Embodiment 1 of the present invention. This embodiment is applicable to situations where the self-motion trajectory of a camera and the depth of a view need to be estimated. This method can be executed by a depth and self-motion trajectory estimation device, which can be implemented in hardware and / or software and can be configured in an electronic device (such as a computer or server). Figure 1 As shown, the method includes:

[0058] S110. Obtain a preset model, a source view, and a target view. The preset model includes a depth estimation network, a motion estimation network, and an implicit cue network. The implicit cue network is used to extract static and dynamic features between the source view and the target view from the motion estimation network and to map them onto the depth estimation network. The source view and the target view are two color images at adjacent time points.

[0059] The preset model is a pre-trained model stored in the depth and self-motion trajectory estimation device, which can be used to estimate the camera's self-motion trajectory and the depth of the view. Estimating the camera's self-motion trajectory means estimating the camera's motion relative to a fixed scene, usually represented by a pose transformation vector or pose transformation matrix; estimating the depth of the view means estimating the depth map of the view, which is a single-channel two-dimensional image composed of the vertical distances from pixels in the scene to the camera's imaging plane.

[0060] The pre-defined model includes a depth estimation network, a motion estimation network, and an implicit cue network. The implicit cue network extracts static and dynamic features between the source and target views from the motion estimation network and maps them identically to the depth estimation network. In other words, the implicit cue network can acquire implicit video cues, which are data features extracted from video frames based on specific convolutional neural networks. For example, monocular motion disparity cues show that nearby objects move faster than distant objects.

[0061] The source view and the target view are two adjacent frames of color images. For monocular video training scenarios, the source view is usually a color image at time t-1 or t+1, and the target view is usually a color image at time t.

[0062] It should be noted that, before step S110 is executed, the present invention can also train the preset model. Specifically, a first training view and a second training view can be obtained; and the preset model can be trained based on the first training view and the second training view.

[0063] S120. Input the source view and target view into the preset model.

[0064] S130. Based on the motion estimation network, estimate the camera's self-motion trajectory; and based on the depth estimation network and the static and dynamic features between the source view and the target view, estimate the depth of the source view and / or the target view.

[0065] Because implicit cue networks can extract static and dynamic features between the source and target views from the motion estimation network and identity map them to the depth estimation network, they can supplement the depth information of a single static view, enhance the geometric constraints of static objects, and compensate for the dynamic features of moving objects, thereby effectively alleviating the artifact problem of moving objects and improving the depth estimation quality of the source and / or target views.

[0066] Example 2

[0067] Figure 2 This is a flowchart illustrating a method for estimating depth and self-motion trajectory according to Embodiment 2 of the present invention. This embodiment provides a detailed pre-set model training method based on Embodiment 1. For example... Figure 2 As shown, the method includes:

[0068] S201. Obtain the first training view and the second training view.

[0069] The first training view and the second training view are two adjacent frames of color images; or, the second training view is a color image of an adjacent frame generated based on the first training view to simulate the first training view.

[0070] S202. Input the first training view and the second training view into the preset model to obtain the first reconstructed view and the second reconstructed view.

[0071] Specifically, the method for "obtaining the first reconstructed view and the second reconstructed view" in step S202 may include the following 7 steps:

[0072] Step 1: Determine the first training view I respectively s Relative to the second training view I t The first pose transformation matrix T t->s Second training view I t Relative to the first training view I s The second pose transformation matrix T s->t Depth map D of the first training view s Depth map D of the second training view t .

[0073] The preset models include DepthNet (depth estimation network), PoseNet (motion estimation network), and ICNet (implicit cue network).

[0074] The DepthNet encoding part of the depth estimation network uses ResNet-18 as the basic framework. The decoder uses 4 layers of convolutional upsampling blocks to predict disparity maps with resolutions of 1 / 8, 1 / 4, 1 / 2 and the same as the input image, respectively. Skip connections are used to sum the feature pixels of the encoding layer to the features of the decoding layer, thereby achieving multi-scale feature fusion.

[0075] For the first training view I s Second training view I t The motion estimation network PoseNet aims to obtain the first training view from it. s Relative to the second training view I t The first pose transformation vector is obtained and converted into the first pose transformation matrix T. t->s And obtain the second training view I t Relative to the first training view I s The second pose transformation vector is then converted into the second pose transformation matrix T. s->t .

[0076] The PoseNet encoding part of the motion estimation network is similar to the DepthNet encoding structure. The first layer expands the number of channels from 3 to 6 to accept two color image inputs. The decoding layer downsamples the 512-dimensional features to generate a 6-dimensional pose transformation vector. Since the pose transformation matrix is ​​not differentiable, while the transformation vector can be converted into a differentiable transformation matrix, the output of PoseNet is a 6-dimensional pose transformation vector.

[0077] Specifically, the first pose transformation matrix T t->s =PoseNet(I s ,I t The second pose transformation matrix T s->t =PoseNet(I t ,I s ).

[0078] The motion estimation network PoseNet extracts dynamic information between adjacent frames, which is complex and redundant. To obtain effective depth cues, the implicit cue network ICNet is used to further abstract effective features and apply them to the decoder of the depth estimation network, thereby obtaining the depth map D of the first training view. s Depth map D of the second training view t .

[0079] The Implicit Cue Network (ICNet) employs a three-bottleneck layer, which includes 1×1, 3×3, and 1×1 convolutions. The input feature size of the bottleneck layer is consistent with the output size. The output features of two bottleneck layers are used to calculate the similarity of the output features in high-dimensional space through a Gaussian kernel function. Finally, this similarity is multiplied pixel-by-pixel with the output features of the third bottleneck layer and then connected to the depth estimation network.

[0080] Specifically, the depth map D of the first training view s =DepthNet_Decoder(ICNet(PoseNet_Encoder(I t ,I s ))+DepthNet_Encoder(I s )));Depth map D of the second training view t =DepthNet_Decoder(ICNet(PoseNet_Encoder(I s ,I t ))+DepthNet_Encoder(I t PoseNet_Encoder represents the encoder of the motion estimation network, DepthNet_Encoder represents the encoder of the depth estimation network, and DepthNet_Decoder represents the decoder of the depth estimation network.

[0081] Step 2: Based on camera parameters K and the depth map D of the second training view t Determine the coordinates P of the first spatial point cloud of the second training view in the camera coordinate system. t .

[0082] After obtaining the depth map D of the second training view tThen, combining the camera imaging principle, based on the camera parameters K and the depth map D of the second training view... t Determine the coordinates P of the first spatial point cloud of the second training view in the camera coordinate system. t .

[0083] Specifically, the coordinates P of the first spatial point cloud t =K -1 D t I t The camera coordinate system can be understood as coordinate system O. t -Z t -X t -Y t .

[0084] Step 3: Based on the coordinates P of the first spatial point cloud t and the first pose transformation matrix T t->s Based on the principle of Euclidean transformation of coordinate systems, the coordinates P of the second spatial point cloud in the coordinate system where the first training view is located are determined. s '.

[0085] Specifically, the coordinates P of the second spatial point cloud s '=T t->s P t The second training view can be understood as coordinate system O in the coordinate system where the first training view is located. s -Z s -X s -Y s .

[0086] Step 4: Based on the camera model, project and sample the second spatial point cloud to obtain the first reconstructed view.

[0087] Step 5: Based on the camera parameters K and the depth map D of the first training view s Determine the coordinates P of the first training view in the third spatial point cloud within the camera coordinate system. s .

[0088] Similarly, after obtaining the depth map D of the first training view... s Then, combining the camera imaging principle, based on the camera parameters K and the depth map D of the first training view... s Determine the coordinates P of the first training view in the third spatial point cloud within the camera coordinate system. s .

[0089] Specifically, the coordinates P of the third-space point cloud s =K -1 D s I s The camera coordinate system can be understood as coordinate system O. t -Zt -X t -Y t .

[0090] Step 6: Based on the coordinates P of the third-space point cloud s Second pose transformation matrix T s->t Based on the principle of Euclidean transformation of coordinate systems, the coordinates P of the fourth spatial point cloud in the coordinate system where the first training view and the second training view are located are determined. t '.

[0091] Specifically, the coordinates P of the fourth-space point cloud t '=T s->t P s The coordinate system in which the first training view and the second training view reside can be understood as coordinate system O. s -Z s -X s -Y s .

[0092] Step 7: Based on the camera model, project and sample the fourth spatial point cloud to obtain the second reconstructed view.

[0093] S203. Determine the reprojection loss and smoothness loss based on the first training view, the second training view, the first reconstruction view, and the second reconstruction view.

[0094] S204. Determine the 3D reconstruction loss.

[0095] Specifically, it can be based on the coordinates P of the first spatial point cloud. t The coordinates P of the second spatial point cloud s ', Coordinates P of the third-space point cloud s The coordinates P of the fourth space point cloud t ', Determine the 3D reconstruction loss L=|P t '-P t |+|P s '-P s |

[0096] S205. Based on the reprojection loss, smoothness loss, and 3D reconstruction loss, the preset model is trained using the principles of backpropagation and gradient descent.

[0097] During training, the parameters of the depth estimation network, motion estimation network, and implicit cue network are updated synchronously.

[0098] S206. Obtain the preset model, source view, and target view.

[0099] S207. Input the source view and target view into the preset model.

[0100] S208. Based on the motion estimation network, estimate the camera's self-motion trajectory; and based on the depth estimation network and the static and dynamic features between the source view and the target view, estimate the depth of the source view and / or the target view.

[0101] This effectively alleviates the artifact problem of moving objects, improves the depth estimation quality of monocular images, reduces pose transformation errors, and enables more accurate estimation of the camera's self-motion trajectory.

[0102] This invention provides a method for estimating depth and self-motion trajectory, comprising: acquiring a preset model, a source view, and a target view; the preset model including a depth estimation network, a motion estimation network, and an implicit cue network; the implicit cue network being used to extract static and dynamic features between the source view and the target view from the motion estimation network and identity-map them to the depth estimation network; the source view and the target view being two adjacent frames of color images; inputting the source view and the target view into the preset model; estimating the camera's self-motion trajectory based on the motion estimation network; and estimating the depth of the source view and / or the target view based on the depth estimation network and the static and dynamic features between the source view and the target view. By designing a preset model that includes a depth estimation network, a motion estimation network, and an implicit cue network, the implicit cue network extracts static and dynamic features between the source and target views from the motion estimation network and maps them identically to the depth estimation network. This supplements the depth information of a single static view, enhances the geometric constraints of static objects, and compensates for the dynamic features of moving objects, thereby effectively mitigating the artifact problem of moving objects and improving the depth estimation quality of the source and / or target views. Simultaneously, the 3D reconstruction loss of this invention further constrains the view reconstruction from the perspective of spatial point clouds, ensuring consistent depth and pose transformations during camera transformation, effectively reducing pose transformation errors, and minimizing the cumulative offset of the predicted motion trajectory in long videos, thus achieving more accurate estimation of the camera's self-motion trajectory.

[0103] Example 3

[0104] Figure 3 This is a schematic diagram of the structure of a depth and self-motion trajectory estimation device provided in Embodiment 3 of the present invention. Figure 3 As shown, the device includes an acquisition module 301 and an estimation module 302.

[0105] The acquisition module 301 is used to acquire a preset model, a source view, and a target view. The preset model includes a depth estimation network, a motion estimation network, and an implicit cue network. The implicit cue network is used to extract static and dynamic features between the source view and the target view from the motion estimation network and identity map them to the depth estimation network. The source view and the target view are two color images at adjacent time points.

[0106] The estimation module 302 is used to input the source view and the target view into a preset model; estimate the camera's self-motion trajectory based on a motion estimation network; and estimate the depth of the source view and / or the target view based on a depth estimation network and static and dynamic features between the source view and the target view.

[0107] Combination Figure 3 , Figure 4 This is a schematic diagram of another depth and self-motion trajectory estimation device provided in Embodiment 3 of the present invention. Figure 4 As shown, it also includes: training module 303.

[0108] The training module 303 is used to acquire a first training view and a second training view before the acquisition module 301 acquires the preset model; and to train the preset model based on the first training view and the second training view.

[0109] Optionally, the training module 303 is specifically used to input the first training view and the second training view into the preset model to obtain the first reconstructed view and the second reconstructed view; determine the reprojection loss and the smoothness loss based on the first training view, the second training view, the first reconstructed view and the second reconstructed view; determine the 3D reconstruction loss; and train the preset model based on the principles of backpropagation and gradient descent based on the reprojection loss, the smoothness loss and the 3D reconstruction loss.

[0110] Optionally, the training module 303 is specifically used to determine the first training view I. s Relative to the second training view I t The first pose transformation matrix T t->s Second training view I t Relative to the first training view I s The second pose transformation matrix T s->t Depth map D of the first training view s Depth map D of the second training view t Based on camera parameters K and the depth map D of the second training view. t Determine the coordinates P of the first spatial point cloud of the second training view in the camera coordinate system. t Based on the coordinates P of the first spatial point cloud t and the first pose transformation matrix T t->s Based on the principle of Euclidean transformation of coordinate systems, the coordinates P of the second spatial point cloud in the coordinate system where the first training view is located are determined. s Based on the camera model, the second spatial point cloud is projected and sampled to obtain the first reconstructed view; according to the camera parameters K and the depth map D of the first training view... s Determine the coordinates P of the first training view in the third spatial point cloud within the camera coordinate system. sBased on the coordinates P of the third-space point cloud s Second pose transformation matrix T s->t Based on the principle of Euclidean transformation of coordinate systems, the coordinates P of the fourth spatial point cloud in the coordinate system where the first training view and the second training view are located are determined. t Based on the camera model, the fourth spatial point cloud is projected and sampled to obtain the second reconstructed view.

[0111] Optionally, training module 303 is specifically used to train the coordinates P of the first spatial point cloud. t The coordinates P of the second spatial point cloud s ', Coordinates P of the third-space point cloud s The coordinates P of the fourth space point cloud t ', Determine the 3D reconstruction loss L=|P t '-P t |+|P s '-P s |

[0112] Optional, the first pose transformation matrix T t->s =PoseNet(I s ,I t );

[0113] Second pose transformation matrix T s->t =PoseNet(I t ,I s );

[0114] Depth map D of the first training view s =DepthNet_Decoder(ICNet(PoseNet_Encoder(I t ,I s ))+DepthNet_Encoder(I s )));

[0115] Depth map D of the second training view t =DepthNet_Decoder(ICNet(PoseNet_Encoder(I s ,I t ))+DepthNet_Encoder(I t )));

[0116] The coordinates P of the first spatial point cloud t =K -1 D t I t ;

[0117] The coordinates P of the second spatial point cloud s '=Tt->s P t ;

[0118] The coordinates P of the third-space point cloud s =K -1 D s I s ;

[0119] The coordinates P of the fourth space point cloud t '=T s->t P s ;

[0120] Wherein, ICNet represents the Implicit Cue Network, PoseNet represents the Motion Estimation Network, PoseNet_Encoder represents the encoder of the Motion Estimation Network, DepthNet_Encoder represents the encoder of the Depth Estimation Network, and DepthNet_Decoder represents the decoder of the Depth Estimation Network.

[0121] Optionally, the first training view and the second training view are two color images at adjacent time points; or, the second training view is a color image of an adjacent frame generated based on the first training view to simulate the first training view.

[0122] The depth and self-motion trajectory estimation device provided in the embodiments of the present invention can execute the depth and self-motion trajectory estimation method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method execution.

[0123] Example 4

[0124] Figure 5 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0125] like Figure 5As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0126] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0127] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as methods for estimating depth and self-motion trajectories.

[0128] In some embodiments, the depth and self-motion trajectory estimation method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the depth and self-motion trajectory estimation method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the depth and self-motion trajectory estimation method by any other suitable means (e.g., by means of firmware).

[0129] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0130] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0131] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0132] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0133] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0134] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0135] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0136] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for estimating depth and self-motion trajectory, characterized in that, include: Obtain a preset model, a source view, and a target view. The preset model includes a depth estimation network, a motion estimation network, and an implicit cue network. The implicit cue network is used to extract static and dynamic features between the source view and the target view from the motion estimation network and identity map them to the depth estimation network. The source view and the target view are two color images at adjacent time points. Input the source view and the target view into the preset model; Based on the motion estimation network, the camera's self-motion trajectory is estimated; and based on the depth estimation network and the static and dynamic features between the source view and the target view, the depth of the source view and / or the target view is estimated.

2. The method according to claim 1, characterized in that, Before obtaining the preset model, the following is also included: Obtain the first training view and the second training view; The preset model is trained based on the first training view and the second training view.

3. The method according to claim 2, characterized in that, The step of training the preset model based on the first training view and the second training view includes: The first training view and the second training view are input into the preset model to obtain the first reconstructed view and the second reconstructed view; Based on the first training view, the second training view, the first reconstruction view, and the second reconstruction view, determine the reprojection loss and the smoothness loss; Determine the loss in 3D reconstruction; Based on the reprojection loss, the smoothness loss, and the 3D reconstruction loss, the preset model is trained using backpropagation and gradient descent principles.

4. The method according to claim 3, characterized in that, Obtaining the first reconstructed view and the second reconstructed view includes: Determine the first training view I respectively s Relative to the second training view I t The first pose transformation matrix T t->s The second training view I t Relative to the first training view I s The second pose transformation matrix T s->t The depth map D of the first training view s and the depth map D of the second training view t ; Based on camera parameters K and the depth map D of the second training view t Determine the coordinates P of the second training view in the first spatial point cloud in the camera coordinate system. t ; Based on the coordinates P of the first spatial point cloud t and the first pose transformation matrix T t->s Based on the principle of Euclidean transformation of coordinate systems, the coordinates P of the second spatial point cloud in the coordinate system where the second training view is located in the first training view are determined. s '; Based on the camera model, the second spatial point cloud is projected and sampled to obtain the first reconstructed view; Based on camera parameters K and the depth map D of the first training view s Determine the coordinates P of the first training view in the third spatial point cloud of the camera coordinate system. s ; Based on the coordinates P of the third spatial point cloud s and the second pose transformation matrix T s->t Based on the principle of Euclidean transformation of coordinate systems, the coordinates P of the fourth spatial point cloud in the coordinate system where the first training view and the second training view are located are determined. t '; Based on the camera model, the fourth spatial point cloud is projected and sampled to obtain the second reconstructed view.

5. The method according to claim 4, characterized in that, The determination of the 3D reconstruction loss includes: Based on the coordinates P of the first spatial point cloud t The coordinates P of the second spatial point cloud s The coordinates P of the third spatial point cloud s and the coordinates P of the fourth spatial point cloud t ', Determine the 3D reconstruction loss L = |P t '-P t | + | P s '-P s | 6. The method according to claim 4, characterized in that, The first pose transformation matrix T t->s = PoseNet(I s , I t ); The second pose transformation matrix T s->t = PoseNet(I t , I s ); Depth map D of the first training view s = DepthNet_Decoder(ICNet(PoseNet_Encoder(I t ,I s )) + DepthNet_Encoder(I s )); Depth map D of the second training view t = DepthNet_Decoder(ICNet(PoseNet_Encoder(I s ,I t )) + DepthNet_Encoder(I t )); The coordinates P of the first spatial point cloud t = K -1 D t I t ; The coordinates P of the second spatial point cloud s ' = T t->s P t ; The coordinates P of the third spatial point cloud s = K -1 D s I s ; The coordinates P of the fourth spatial point cloud t ' = T s->t P s ; Wherein, ICNet represents the implicit cue network, PoseNet represents the motion estimation network, PoseNet_Encoder represents the encoder of the motion estimation network, DepthNet_Encoder represents the encoder of the depth estimation network, and DepthNet_Decoder represents the decoder of the depth estimation network.

7. The method according to claim 2, characterized in that, The first training view and the second training view are two adjacent frames of color images; or, the second training view is a color image generated based on the first training view to simulate adjacent frames of the first training view.

8. A device for estimating depth and self-motion trajectory, characterized in that, include: Acquisition module and estimation module; among which, The acquisition module is used to acquire a preset model, a source view, and a target view. The preset model includes a depth estimation network, a motion estimation network, and an implicit cue network. The implicit cue network is used to extract static and dynamic features between the source view and the target view from the motion estimation network and map them identically to the depth estimation network. The source view and the target view are two color images at adjacent time points. The estimation module is used to input the source view and the target view into the preset model; estimate the camera's self-motion trajectory based on the motion estimation network; and estimate the depth of the source view and / or the target view based on the depth estimation network and the static and dynamic features between the source view and the target view.

9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, which enables the at least one processor to perform the depth and self-motion trajectory estimation method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the method for estimating depth and self-motion trajectory as described in any one of claims 1-7.