The invention discloses an unsupervised
monocular endoscope depth
estimation method combined with a segmentation
large model. A network framework comprises a depth prediction network phi D, a
pose estimation network phi T, an intrinsic
decomposition module phi I and a synthesis reconstruction module phi L. Wherein the depth
estimation network adopts a deep
large model fine tuning technology, is combined with rich semantic
prior information in a segmentation
large model, and is comprehensively trained by using
decomposition synthesis loss,
reflection loss, mapping synthesis loss and an edge
perception depth
smoothing loss function guided by a segmentation
mask. The
semantic enhancement branch network adopts a
network architecture of a segmented large model, image
semantic information is accurately extracted, structures such as different tissues and organs are distinguished, and a semantic
cognition system for an operation image is established. And obtaining semantic features of a target image through a
fine tuning encoder, and after the semantic features are fused with depth features, enabling a depth estimation model to be combined with context
semantics to optimize depth prediction. Meanwhile, an
automatic segmentation mask graph generated by a segmentation large model is utilized to calculate depth smooth loss guided by a
mask, and the rationality of a depth
estimation result is further constrained. According to the scheme, the
encoder and the
automatic segmentation mask in the large model are segmented, semantic
prior information in the large model is fully combined, the depth estimation effect is improved, and the operation precision and the safety performance of the
endoscope minimally
invasive surgery are further enhanced.