Method for Establishing Monocular Depth Estimation Network Based on Dynamic View Selection and Its Application

By dynamically selecting source images with high time consistency and maximum photometric effective areas and pairing them with the target images for joint training, the problem of difficult position alignment and large depth estimation errors in the prior art is solved, and a higher accuracy of monocular endoscopic image depth estimation is achieved.

CN119313716BActive Publication Date: 2025-05-30HUAZHONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411469165.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-21
Publication Date
2025-05-30
Estimated Expiration
2044-10-21

AI Technical Summary

Technical Problem

In the prior art, when using adjacent frames as source image-target image pairs for monocular depth estimation, the position poses are difficult to align, resulting in large depth estimation errors.

Method used

By constructing the training dataset, the source image is dynamically selected, and frames with high temporal consistency and maximum luminosity effective area are selected from multiple historical frames in front of the target image as the source image, and paired with the target image for joint training of PoseNet and DepthNet.

Benefits of technology

It effectively improves the accuracy of monocular endoscopic image depth estimation, reduces the interference of static use of adjacent frames to pose alignment, and improves PoseNet's pose estimation accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119313716B_ABST
    Figure CN119313716B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for establishing a monocular depth estimation network based on dynamic view selection and its application, belonging to the field of monocular endoscope depth estimation, including: constructing a training data set, and using the photometric loss as a constraint, jointly training the pose estimation network PoseNet and the depth estimation network DepthNet with the training data set, and taking the DepthNet after joint training as the monocular depth estimation network; the construction method of the training data includes: respectively taking the N historical frames before the target image as source images, and calculating the time consistency scores corresponding to each source image; the time consistency score reflects the difference between the depth map and the ideal reference depth map; selecting the part of the historical frames with the highest time consistency score as candidate source images, and selecting the candidate source image with the largest photometrically valid region as the final source image, and together with the target image, constituting a piece of training data. The present invention can effectively improve the accuracy of monocular endoscope image depth estimation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of monocular endoscope depth estimation, and more specifically, relates to a method for establishing a monocular depth estimation network based on dynamic view selection and an application thereof. Background Art

[0002] Monocular endoscopes play an important role in gastrointestinal diagnosis and surgical operations. Automatic 3D depth estimation is the key for endoscopic robots to fully understand the surgical scene and expand the limited field of view through 3D scene reconstruction. Since the field of view of videos captured by endoscopes is usually very narrow, doctors need to repeatedly observe the entire tissue to form a judgment. 3D scene reconstruction helps to expand the field of view and assist other downstream tasks, such as surgical navigation through model alignment with preoperative computed tomography (CT). Depth estimation of monocular endoscopic images is a prerequisite for reconstructing 3D scenes. In a living environment, accurate monocular endoscope depth estimation is extremely challenging due to the lack of depth labels.

[0003] At present, monocular depth estimation based on deep learning mainly relies on self-supervised learning, and its core idea is the photometric constraint between the real image and the distorted image. The goal of self-supervised monocular depth estimation is to achieve photometric consistency, which is defined as the pixel similarity between the target image and the aligned source image. Specifically, two convolutional neural networks (CNNs) need to be constructed, one called the depth estimation network (DepthNet), which can recover the depth information, based on which the source image can be projected into the 3D space. The other is called the pose estimation network (PoseNet), which can obtain the relative pose from the source view to the target view. The relative pose is used to transform the 3D spatial coordinates of the source image into the 3D space of the target view, and the distorted image is reconstructed by backprojection and bilinear interpolation. In this process, DepthNet and PoseNet achieve joint training of DepthNet and PoseNet by minimizing the photometric loss (essentially the pixel-level difference between the distorted image and the target image), thereby achieving joint optimization of the parameters in DepthNet and PoseNet. After that, the optimized DepthNet can be used to achieve depth estimation of monocular endoscopic images.

[0004] To achieve the joint optimization of DepthNet and PoseNet, it is necessary to use source image-target image pairs as training data. In the stereo depth estimation task of binocular images, the left and right views are fixedly used as the source image-target image pair, and their relative poses are obtained in advance through calibration. In the monocular scenario, there is only a single camera sensor, and its depth estimation is more challenging. Most researchers use adjacent frames in the video data captured by a monocular endoscope as the source image-target image pair. The time interval between adjacent image frames is often in the order of dozens of milliseconds, so there are overlapping field-of-view regions in the image frames, which can be used to assist network training. However, this common training method of statically using adjacent frames as the source image-target image pair inevitably brings a problem, that is, the camera motion captured in these image frames in a very short time is small, and the provided visual cues are very limited, which brings great difficulties to pose estimation. That is to say, using the source image-target image pair constructed by adjacent image frames is not conducive to the pose alignment of the pose estimation network, which also leads to a large error in the depth of the monocular endoscope image estimated by the depth estimation network obtained through joint training. Summary of the Invention

[0005] In view of the defects and improvement requirements of the prior art, the present invention provides a method for establishing a monocular depth estimation network based on dynamic view selection and its application. The purpose is to effectively solve the problems of difficult pose alignment and large depth estimation error when statically using adjacent frames as image pairs, thereby effectively improving the accuracy of monocular endoscope image depth estimation.

[0006] To achieve the above object, according to one aspect of the present invention, a method for establishing a monocular depth estimation network based on dynamic view selection is provided, including:

[0007] Step S1: Construct a training data set; in the training data set, each piece of training data is a source image-target image pair, and the construction method of each piece of training data includes:

[0008] Select a frame in the monocular video frame sequence as the target image. For the N historical frames before the target image, calculate the cost cube from each historical frame to the target image respectively, and convert each cost cube into the depth map corresponding to each historical frame to obtain a depth map set;

[0009] Based on the depth map set, count the depth values that have appeared at each pixel position and the number of occurrences of each depth value, and use the depth value with the most occurrences at each position as the depth value at the corresponding position to obtain the reference depth image D mv , and calculate the temporal consistency score of each historical frame; the depth map corresponding to the historical frame and the reference depth image D mvThe smaller the difference is, the higher the temporal consistency score is;

[0010] Select the part of historical frames with the highest temporal consistency score as the candidate source image, and select the candidate source image with the largest photometric valid region as the final source image, which together with the target image constitutes a piece of training data; the photometric valid region of the image is the intersection of the non-occluded region of the candidate source image and the common field of view region between this image and the target image;

[0011] Step S2: Using the photometric loss as a constraint, jointly train the pose estimation network PoseNet and the depth estimation network DepthNet with the constructed training dataset, and use the jointly trained depth estimation network DepthNet as the monocular depth estimation network;

[0012] where N is an integer greater than 1; the pose estimation network PoseNet estimates the relative pose from the target image to the source image, and the depth estimation network DepthNet is used to estimate the depth map of the image.

[0013] Further, let I (0) represent the target image, and let I (-n) represent the nth historical frame image before the target image. Then the determination method of the common field of view region between the image I (-n) and the target image I (0) includes:

[0014] Use the pre-trained pose estimation network PoseNet to estimate the relative pose t (0) from the target image I (-n) to the image I (-n) ;

[0015] According to the relative pose T (-n) and the reference depth image D mv align the viewpoints from the image I (-n) to the target image I (0) to obtain the corresponding relationship between the pixel p (-n) in the warped image and the pixel p (0) in I (0) ;

[0016] Determine the region composed of the pixels p (-n) in the warped image that are within the range of [0, W - 1] × [0, H - 1] as the common field of view region between the image I (-n) and the target image I (0) ;

[0017] where W and H are the width and height of the image respectively, and n ∈ {1, 2..., N}.

[0018] Further, for the image I (-n)The determination method of the non-occluded region of

[0019] Use the pre-trained depth estimation network DepthNet to estimate the image I (-n) 's depth map D (-n) ;

[0020] Use the pre-trained pose estimation network PoseNet to estimate the image I (-n) to the target image i (0) 's relative pose t (-n) ';

[0021] According to the relative pose t (-n) ' and the depth map d (-n) Align the viewing pose of the target image i (0) to the image i (-n) to obtain the warped image i (-n) ';

[0022] Use the pre-trained pose estimation network PoseNet to estimate the target image i (0) to the warped image I (-n) 's relative pose T (-n) ";

[0023] According to the relative pose T (-n) " and the reference depth image D mv Align the viewing pose of the warped image I (-n) ' to the target image I (0) to obtain the warped image I (0) ";

[0024] Determine the region in the warped image I (0) " whose difference from the target image I (0) is not greater than the preset threshold as the non-occluded region of the image I (-n) .

[0025] Furthermore, the calculation method of the cost cube from the image I (-n) to the target image I (0) includes:

[0026] Perform mean sampling within the disparity range [d min , d max to obtain N d depth planes d min and d max represent the minimum depth value and the maximum depth value respectively;

[0027] For each depth plane D i , according to the relative pose T (-n) and the depth plane Di , align the pose of the viewpoint of image I (-n) to the target image I (0) to obtain a warped image and calculate as the matching cost at this depth value; i = 0, 1, 2…N d -1;

[0028] Concatenate the matching costs at each depth value to obtain the cost cube C (-n) from image I (0) to the target image I (-n) .

[0029] Furthermore, let represent the depth map corresponding to image I (-n) , then in the depth map , the depth value at any position (x, y) is:

[0030]

[0031] where C (-n) (x, y, i) represents the matching cost corresponding to the i-th depth value at position (x, y) in the cost cube; represents finding the depth value with the minimum matching cost at position (x, y).

[0032] Furthermore, let S (-n) represent the temporal consistency score of image I (-n) , then

[0033]

[0034] where D mv (x, y) represents the depth value at position (x, y) in the reference depth image D mv .

[0035] Furthermore, let I s and I t represent the source image and the target image respectively, let T t→s represent the relative pose of the target image I t to the source image I s , let D t represent the depth map of the target image I t , then according to the relative pose T t→s and the depth map D t , after aligning the pose of the viewpoint of the source image I s to the target image I t , the obtained warped image I s→t is:

[0036] Is→t (p s→t ) = I s <proj(D t , T t→s , K, p s→t )>

[0037] where p s→t represents the homogeneous coordinates of the pixel, K represents the given camera intrinsic matrix, proj(·) is the projection function, and <> represents differentiable bilinear sampling.

[0038] Furthermore, the expression of the photometric loss is as follows:

[0039]

[0040] where represents the photometric loss, SSIM represents the structural similarity index measure, and α is the weight parameter.

[0041] According to another aspect of the present invention, a monocular endoscope depth estimation method is provided, including:

[0042] Inputting the monocular endoscope image to be estimated into the monocular depth estimation network to obtain the depth map of the monocular image to be estimated;

[0043] where the monocular depth estimation network is established by the above-mentioned monocular depth estimation network establishment method based on dynamic view selection provided by the present invention.

[0044] According to another aspect of the present invention, a computer-readable storage medium is provided, including a stored computer program; when the computer program is executed by a processor, it controls the device where the computer-readable storage medium is located to execute the above-mentioned monocular depth estimation network establishment method based on dynamic view selection provided by the present invention, and / or the above-mentioned monocular depth estimation method provided by the present invention.

[0045] Based on the well-trained pose estimation network PoseNet and depth estimation network DepthNet conceived by the present invention, a higher-quality training dataset is further constructed to jointly train and optimize the two; when constructing the training dataset of the present invention, for the target image, adjacent frames are not directly used as the source image, but the most suitable historical frame is selected from multiple historical frames before it as the source image, realizing dynamic matching of the most suitable source image within a larger time span, greatly reducing the interference of static use of adjacent frames for pose alignment; when selecting the source image, first calculate the temporal consistency score of each historical frame, which reflects the difference degree between the estimated depth map of the target image and the ideal depth map when using the corresponding historical frame as the source image, and then select some frames with high consistency scores as candidate source images, which can ensure the accuracy of depth estimation to a certain extent; in the photometric optimization loss process, there needs to be enough effective regions. Therefore, based on the selected candidate source images, the present invention further selects the candidate source image with the largest photometric effective region as the finally matched source image to be paired with the target image, so as to ensure that when jointly training PoseNet and DepthNet based on the photometric loss, PoseNet can be effectively optimized, and finally the accuracy of PoseNet for monocular endoscopic depth estimation can be effectively improved.

[0046] Generally speaking, the present invention selects frames with high temporal consistency from multiple historical frames of the target image as candidate source images, and selects the candidate source image with the largest photometric effective region from them as the final source image to be paired with the target image as the training data for the joint training of PoseNet and DepthNet, which can effectively improve the depth estimation accuracy of the trained depth estimation network. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 Schematic diagram of the method for establishing a monocular depth estimation network based on dynamic view selection provided by an embodiment of the present invention;

[0048] Figure 2 Schematic diagram of the synthesis of existing distorted images;

[0049] Figure 3 Schematic diagram of calculating the temporal consistency score provided by an embodiment of the present invention;

[0050] Figure 4 Schematic diagram of calculating the photometric effective region provided by an embodiment of the present invention;

[0051] Figure 5 Schematic diagram of the comparison of estimation errors between the embodiment of the present invention and other depth estimation methods. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0053] In the present invention, terms such as "first" and "second" in the present invention and the accompanying drawings (if any) are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.

[0054] In order to solve the technical problem that when the existing monocular endoscope image depth estimation method directly uses adjacent frames as the source image-target image pair to jointly train the pose estimation network PoseNet and the depth estimation network DepthNet, the pose alignment cannot be effectively achieved, resulting in low depth estimation accuracy of the optimized depth estimation network DepthNet, the present invention provides a method for establishing a monocular depth estimation network based on dynamic view selection and its application. The overall concept is to improve the construction method of the training data set. For a given target image, the most suitable source image is dynamically selected based on temporal consistency and effective photometric regions within a larger time span and paired with the target image, thereby reducing the interference of statically using adjacent frames for pose alignment and adapting to the characteristics of the joint training of the pose estimation network PoseNet and the depth estimation network DepthNet under photometric loss constraints, effectively improving the estimation accuracy of the finally obtained depth estimation network.

[0055] Based on the above concept, in order to effectively improve the training efficiency and ensure the training effect, the present invention will first obtain a pre-trained pose estimation network PoseNet and a depth estimation network DepthNet with certain estimation capabilities. The pre-trained pose estimation network PoseNet and depth estimation network DepthNet can either directly select a trained network or be trained using existing training methods. Without loss of generality, the pre-training processes of the pose estimation network PoseNet and the depth estimation network DepthNet are included in the following embodiments.

[0056] The following are the embodiments.

[0057] Embodiment 1:

[0058] A method for establishing a monocular depth estimation network based on dynamic view selection, as Figure 1 shown, this embodiment includes the following steps S0 to S2:

[0059] Step S0: Use the adjacent-frame pre-trained pose estimation network PoseNet and depth estimation network DepthNet;

[0060] Step S1: Construct a training dataset based on the proposed method of dynamically constructing source image-target image pairs;

[0061] Step S2: Using the photometric loss as a constraint, jointly train the pose estimation network PoseNet and depth estimation network DepthNet with the constructed training dataset, and use the jointly trained depth estimation network DepthNet as a monocular depth estimation network.

[0062] The following is a detailed explanation of each step.

[0063] In this embodiment, step S0 specifically includes:

[0064] Step S01: Preprocess the original images: First, considering the computing resources and the required data size of the network, scale the sizes of all video frames to a unified size (320×156 in this embodiment); second, uniformly perform the following data augmentation on the input images:

[0065] 1) Horizontal flipping;

[0066] 2) Color transformation, that is, the brightness, contrast, and saturation are multiplied by a multiple within the range of [0.8, 1.2], and the hue is transformed within the range of [-0.1, 0.1];

[0067] Step S02: Use adjacent frames to construct source image-target image pairs, and use the constructed source image-target image pairs as training data to complete the pre-training of PoseNet and DepthNet.

[0068] In order to obtain the pose information between different views and get more reliable aligned views during the final self-supervised training, it is first necessary to fixedly use two adjacent endoscopic images to calculate the photometric error between the target image and the warped image after warping.

[0069] The depth estimation network DepthNet for predicting the depth of each pixel and the pose estimation network PoseNet for predicting the relative camera pose between two frames of images. Optionally, in this embodiment, both DepthNet and PoseNet are fully convolutional networks. PoseNet contains a lightweight encoder based on ResNet18, which has four convolutional layers. By simultaneously inputting two images, the relative pose from one view to another view can be predicted. The process of synthesizing the warped image is as Figure 2 shown, the target image is I t , and the adjacent image frame is used as the source image I s. First, DepthNet estimates the depth map D of the target frame I t of I t , and PoseNet estimates the relative camera motion T from the target view to the source view t→s . Using D t and T t→s , the pose alignment of the viewpoints from the source view I s to the target view I t can be performed, and a new warped image I s→t can be synthesized. The matching pixels of the obtained warped image I s→t and its source image satisfy the following relationship:

[0070] I s→t (p s→t ) = I s <proj(D t , T t→s , K, p s→t )>

[0071] where p s→t represents the pixel homogeneous coordinates, K represents the given camera intrinsic matrix, and proj(·) is the projection function. With the above pixel matching relationship, I s→t can obtain the value at each pixel position from I s using the differentiable bilinear sampling represented by <·>. After the transformation, DepthNet and PoseNet can be constrained by the photometric loss, and the calculation method is as follows:

[0072]

[0073] where represents the photometric loss, SSIM represents the structural similarity index measure, and α is the weight parameter; optionally, in this embodiment, the value of α is 0.85.

[0074] It is easy to understand that the above structural descriptions of DepthNet and PoseNet are only optional implementation manners of the present invention and should not be construed as the sole limitation to the present invention. In other embodiments of the present invention, other structures or even other types of network structures may also be used. In addition, in some other embodiments of the present invention, if pre-trained DepthNet and PoseNet are directly adopted, step S0 does not need to be executed.

[0075] In this embodiment, in step S1, for each source image-target image pair in the training dataset, the construction method is as follows:

[0076] To solve the problem that it is difficult to align the camera poses of adjacent image frames, a cost cube calculated using the frozen PoseNet is constructed to select reliable views, such as Figure 3 as shown.

[0077] First, taking the current frame I (0) as the target image, for the N past historical frames of I (0) the relative pose T (0) from I (-n) to its nth previous frame I (-n) is predicted using the frozen pre-trained PoseNet.

[0078] To calculate the matching cost at different depth probabilities, in this embodiment, some depths D min are defined by mean sampling within the interval from the minimum depth d max to the maximum depth d i Each D i represents a depth plane, and N d is the number of depth planes determined by mean sampling, specifically determined by the sampling interval. Then, pose alignment is performed using each depth plane D i and the warped images corresponding to each depth plane D i are synthesized in the same way as the synthesized warped image shown in Figure 2 The coordinates of the synthesized warped image and the coordinates of I (-n) satisfy the following correspondence:

[0079]

[0080] By exhaustively defining all depth planes, a set of reconstructed images The three-dimensional cost cube C (-n) can be expressed as:

[0081]

[0082] where (x, y) represents the two-dimensional spatial coordinates of the image, and I is the new dimension of the cost cube, which represents the matching cost under different depth planes. The cost volume temporal flow can be further calculated from the N past historical frames, that is, the set of three-dimensional cost cubes C (-n) calculated from the N past historical frames. These cost cubes represent the depth probability distribution of the target image viewpoint. Regardless of how the poses change between different frames, the depth probability distribution of the target frame viewpoint should remain consistent.

[0083] To obtain the temporal consistency score, in this embodiment, the three-dimensional cost cube C is first converted into a depth map through the minimum matching cost (-n) Convert it into a depth map

[0084]

[0085] Use the method of minimum matching cost to obtain depth maps under different historical frames Set Perform majority voting in the time dimension to select the reference depth image D mv The depth values at each position in, that is, based on the set of depth maps Count the depth values that have appeared at each pixel position and the number of occurrences of each depth value. The depth value with the most occurrences at each position is used as the depth value at the corresponding position. The determined reference depth image represents a typical stable and temporally consistent depth distribution.

[0086] Based on the determined reference depth image D mv , this embodiment will further calculate the temporal consistency scores of each historical frame; the smaller the difference between the depth map corresponding to the historical frame and the reference depth image D mv , the higher the temporal consistency score. Let S (-n) represent the temporal consistency score of I (-n) , then in this embodiment, the expression of S (-n) is as follows:

[0087]

[0088] When the temporal consistency score calculated in this embodiment uses the corresponding historical frame as the source image, the difference degree between the estimated depth map of the target image and the ideal depth map (i.e., the reference depth image D mv ). After obtaining the temporal consistency scores of each historical frame in this embodiment, the largest K frames will be selected as candidate source images. These candidate source images will be further screened to determine the final source image that matches the target image. This embodiment uses the reference depth image as a benchmark to determine the candidate source images, ensuring that the training data formed by pairing the finally determined source image with the target image can effectively improve the depth estimation accuracy of DepthNet.

[0089] In practical applications, the number K of the selected candidate source images can be flexibly set to a positive integer greater than 1. Optionally, in this embodiment, K is set to 0.5*N.

[0090] After selecting candidate source images with high temporal consistency, considering that sufficient effective regions are required during the photometric optimization loss process, this embodiment will further select the candidate source image with the largest photometric effective region from the candidate source images as the finally matched source image. The calculation of the common field of view region and the non-occluded region is involved in this process, as Figure 4 shown.

[0091] Specifically, there are usually two reasons for the generation of invalid regions: 1) the non-common field of view caused by camera movement, that is, the region within the viewing angle of the target image I (0) but outside the viewing angle of the candidate source image i (-n) ; 2) the region occluded by tissues and instruments.

[0092] First, in order to calculate the common field of view region between the target image and each candidate source image, according to the reference depth image D mv and the relative pose T (-n) , the viewing angles from I (-n) to I (0) can be pose-aligned to obtain the coordinate matching relationship between the pixel p (-n) in the warped image and the pixel p (0) in I (0) . Referring to the warped image synthesis relationship shown in Figure 2 , the expression of p (-n) is as follows:

[0093] p (-n) =proj(D mv ,T (-n) ,K,P (0) )

[0094] For the region in I (0) but not in I (-n) , the value of the homogeneous coordinate p (-n) of the corresponding point will exceed the boundary of I (-n) . By checking whether the coordinates of the points corresponding to I (0) in I (-n) are within the range of [0, W - 1] × [0, H - 1], it can be determined whether the corresponding points are in the common field of view region, where W and H are the width and height of the image respectively, and are uniformly set to 320 and 256 according to step S01. Finally, the region composed of the points within the range of [0, W - 1] × [0, H - 1] in the warped image p (-n) is determined as the common field of view region between I (-n) and I (0) , and the common field of view region mask

[0095] In addition, for the depth estimation at the same position in the occluded area being different due to different geometric structures, in this embodiment, a cyclic distortion method is adopted to filter out inconsistent areas and locate non-occluded areas. Specifically, project I (0) re-projection to obtain I (-n) ′, and then use I (-n) ′ for re-projection to obtain I (0) ″, that is, align the poses of the viewpoints from I (0) to I (-n) to obtain the distorted image I (-n) ′; then align the poses from I (-n) ′ to I (-n) ′ to obtain the distorted image I (0) ″. The depth information and relative pose required in this process are respectively estimated by the pre-trained DepthNet and PoseNet. Since the depth in the occluded area is inconsistent in the two images, the areas with large differences between the two distorted I (0) ″ and I (0) are occluded areas. On the contrary, the areas composed of points with small differences between I (0) ″ and I (0) are non-occluded areas. Optionally, in this embodiment, the following method is used to generate the non-occluded area mask

[0096]

[0097] where γ is a preset percentage threshold. Optionally, the percentage threshold γ is set to 0.2.

[0098] Thus, by combining and the intersection of the common field of view area and the non-occluded area, that is, the photometric effective area, can be obtained. The relevant expression is as follows:

[0099]

[0100] For the selected K candidate source images, calculate the photometric effective area respectively according to the above method, and select the candidate source image with the largest photometric effective area as the most suitable source image for photometric optimization, and pair it with the target image I (0) to form training data.

[0101] Step S2 of this embodiment is similar to the above step S02, the difference being that in step S2, the training data is the source image-target image pair constructed through the above step S1.

[0102] Generally speaking, the depth estimation network trained by the method of dynamic view selection proposed in this embodiment matches the most suitable source image frame within a large time span, greatly reducing the interference of statically using adjacent frames for pose alignment. Moreover, the proposed method for calculating the photometric effective area simultaneously detects the mutual view area and the unoccluded area, ensuring that the selected source image frame has a reliable size of the photometric effective area. Compared with other methods that statically use adjacent frames, this embodiment dynamically uses historical frames to train the network. Finally, the constructed monocular depth estimation network has higher accuracy than the networks trained by other methods on the depth label-free dataset. In addition, while improving the estimation accuracy of the depth estimation network through joint training, this embodiment can also improve the estimation accuracy of the pose estimation network for relative poses.

[0103] Embodiment 2:

[0104] A monocular endoscopic depth estimation method, including:

[0105] Input the monocular endoscopic image to be estimated into the monocular depth estimation network to obtain the depth map of the monocular image to be estimated;

[0106] Among them, the monocular depth estimation network is established by the method for establishing a monocular depth estimation network based on dynamic view selection provided in Embodiment 1 above.

[0107] Due to the relatively high depth estimation accuracy of the monocular depth estimation network established in Embodiment 1 above, based on this monocular depth estimation network, this embodiment has relatively high estimation accuracy.

[0108] The following further analyzes and verifies the beneficial effects that can be achieved in this embodiment by combining specific comparative experimental data.

[0109] This embodiment selects 8 mainstream depth estimation methods (Monodepth2, HR-Depth, DIFFNet, Endo-SfMLearner, AF-SfMLearner, MonoViT, GasMono, and Lite-Mono) as the comparative methods in this embodiment. Among them, Endo-SfMLearner and AF-SfMLearner are depth estimation methods customized for endoscopes, and the rest are depth estimation methods in natural scenes. These methods all use static adjacent frames for training.

[0110] This embodiment and each comparative method are respectively tested on the laparoscopic dataset (SCARED), and at the same time, a comparison is also made with other source frame selection methods. The evaluation metrics selected are Abs Rel, Sq Rel, RMSE, RMSE log, and δ, and the calculation formulas for each evaluation metric are as follows:

[0111]

[0112] Among them, d i represents the predicted depth value, and represents the true depth value.

[0113] Among the evaluation metrics, the smaller the Abs Rel, Sq Rel, RMSE, and RMSE log, the better the corresponding effect; δ is the accuracy rate, and the larger the value, the better the corresponding effect.

[0114] Before measurement, the predicted depth map needs to be scaled first because its scale is unknown. The scaling factor is the median of the gold standard depth divided by the median of the predicted depth.

[0115] Table 1 shows the evaluation metrics of this embodiment and other methods. According to the results shown in Table 1, this embodiment exceeds other methods in all five metrics. Among them, in the Sq Rel metric, this method reduces by 15.89% compared to the sub-optimal method AF-SfMLearner.

[0116] Evaluation Metrics of Test Results of Each Method in Table 1

[0117]

[0118] Figure 5 shows the error comparison diagrams of various methods. Among them, each row is a case in the SCARED dataset, and each column from left to right is divided into the RGB image and the local error diagrams of Monodepth2, HR-Depth, DIFFNet, Endo-SfMLearner, AF-SfMLearner, MonoViT, GasMono, Lite-Mono, and this embodiment; the error diagrams are obtained by mapping the Abs Rel values of pixels to different colors. According to Figure 5 the results shown, the error of this embodiment is significantly smaller than that of other methods.

[0119] Embodiment 3:

[0120] A computer-readable storage medium includes a stored computer program; when the computer program is executed by a processor, it controls the device where the computer-readable storage medium is located to execute the monocular depth estimation network establishment method based on dynamic view selection provided in the above Embodiment 1, and / or, the monocular depth estimation method provided in the above Embodiment 2.

[0121] Those skilled in the art can easily understand that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for establishing a monocular depth estimation network based on dynamic view selection, characterized in that: include: Step S1: construct a training data set; In the training data set, each piece of training data is a source image-target image pair, and the construction method of each training data includes: A frame in a monocular video frame sequence is selected as the target image. For the N historical frames before the target image, each historical frame is used as the source image, and the cost cube from the source image to the target image is calculated. Each cost cube is converted into a depth map corresponding to each historical frame to obtain a depth map set. Based on the depth map set, the depth values ​​that appear at each pixel position and the number of occurrences of each depth value are counted, and the depth value with the largest number of occurrences at each position is used as the depth value at the corresponding position to obtain a reference depth image D mv , and calculate the temporal consistency score of each historical frame; the depth map corresponding to the historical frame and the reference depth image D mv The smaller the difference, the higher the temporal consistency score; Select some historical frames with the highest temporal consistency score as candidate source images, and select the candidate source image with the largest photometric effective area as the final source image, which together with the target image constitute a piece of training data; the photometric effective area of ​​the image is the intersection of the non-occluded area of ​​the candidate source image and the common field of view area between the image and the target image; Step S2: Taking photometric loss as a constraint, jointly training the pose estimation network PoseNet and the depth estimation network DepthNet using the constructed training data set, and using the jointly trained depth estimation network DepthNet as the monocular depth estimation network; Among them, N is an integer greater than 1; the pose estimation network PoseNet is used to estimate the relative pose of the target image to the source image, and the depth estimation network DepthNet is used to estimate the depth map of the image.

2. The method for establishing a monocular depth estimation network based on dynamic view selection according to claim 1, characterized in that: Take I (0) Represents the target image, with I (-n) represents the nth historical frame image before the target image, then image I (-n) With the target image I (0) Methods for determining the common visual field between the two include: Use the pre-trained pose estimation network PoseNet to estimate the target image I (0) To Image I (-n) The relative position T (-n) ; According to the relative posture T (-n) and the reference depth image D mv Image I (-n) To the target image I (0) The viewpoints are aligned to obtain the pixel p in the distorted image. (-n) and I (0) Medium pixel p (0) The corresponding relationship; The pixel p in the distorted image that is in the range [0,W-1]×[0,H-1] (-n) The region formed is determined as image I (-n) With the target image I (0) The common visual field between Where W and H are the width and height of the image, respectively, n∈{1,2…,N}.

3. The method for establishing a monocular depth estimation network based on dynamic view selection according to claim 2, characterized in that: Image I (-n) The non-occluded area is determined by: Use the pre-trained depth estimation network DepthNet to estimate image I (-n) The depth map D (-n) ; Use the pre-trained pose estimation network PoseNet to estimate image I (-n) To the target image I (0) The relative position T (-n) ′; According to the relative posture T (-n) ′ and the depth map D (-n) The target image I (0) To Image I (-n) The viewpoints are aligned to obtain the distorted image I (-n) ′; Use the pre-trained pose estimation network PoseNet to estimate the target image I (0) To the distorted image I (-n) The relative position T of (-n) ''; According to the relative posture T (-n) '' and the reference depth image D mv The distorted image I (-n) ′ to the target image I (0) The viewpoints are aligned to obtain the distorted image I (0) ''; The distorted image I (0) ′′ and the target image I (0) The area where the difference is not greater than the preset threshold is determined as image I (-n) non-occluded area.

4. The method for establishing a monocular depth estimation network based on dynamic view selection according to claim 2 or 3, characterized in that: Image I (-n) To the target image I (0) The cost cube is calculated by: In the parallax range [d min ,d max ] to obtain N d Depth Plane d min and d max Respectively represent the minimum depth value and the maximum depth value; For each depth plane D i , according to the relative posture T (-n) and the depth plane D i , image I (-n) To the target image I (0) The viewpoints are aligned to obtain the distorted image And calculate As the matching cost at this depth value; i = 0, 1, 2...N d -1; The matching costs at each depth value are concatenated to obtain image I (-n) To the target image I (0) The cost cube C (-n) .

5. The method for establishing a monocular depth estimation network based on dynamic view selection according to claim 4, characterized in that: by Represents image I (-n) The corresponding depth map, then the depth map In the example, the depth value at any position (x, y) is: Among them, C (-n) (x, y, i) represents the matching cost corresponding to the i-th depth value at position (x, y) in the cost cube; Indicates the depth value with the minimum matching cost at the position (x, y).

6. The method for establishing a monocular depth estimation network based on dynamic view selection according to claim 5, characterized in that: S (-n) Represents image I (-n) The temporal consistency score of Among them, D mv (x, y) represents the reference depth image D mv The depth value at position (x,y) in the image.

7. The method for establishing a monocular depth estimation network based on dynamic view selection according to claim 1, characterized in that: Take I s and I t Denote the source image and the target image respectively, with T t→s Represents the target image I t To source image I s The relative position of t Represents the target image I t The depth map of t→s and the depth map D t The source image I s To the target image I t After the viewpoints are aligned, the distorted image I s→t for: I s→t (p s→t )=I s <proj(D t ,T t→s ,K,p s→t )> Among them, p s→t represents pixel homogeneous coordinates, K represents the given camera intrinsic parameter matrix, proj(·) is the projection function, and <> represents differentiable bilinear sampling.

8. The method for establishing a monocular depth estimation network based on dynamic view selection according to claim 7, characterized in that: The expression of the luminosity loss is as follows: in, represents photometric loss, SSIM represents structural similarity index metric, and α is the weight parameter.

9. A monocular endoscope depth estimation method, characterized in that: include: Inputting the monocular endoscopic image to be estimated into a monocular depth estimation network to obtain a depth map of the monocular endoscopic image to be estimated; Wherein, the monocular depth estimation network is established by the method for establishing a monocular depth estimation network based on dynamic view selection described in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that: Including a stored computer program; when the computer program is executed by a processor, it controls the device where the computer-readable storage medium is located to execute the method for establishing a monocular depth estimation network based on dynamic view selection as described in any one of claims 1 to 8, and / or the monocular endoscope depth estimation method as described in claim 9.