A self-supervised monocular depth estimation method based on visual slam algorithm

By employing a self-supervised monocular depth estimation method, and utilizing depth estimation and pose estimation networks, combined with multi-scale feature fusion and displacement operations, the problems of difficult image spatial information acquisition and slow computation speed in monocular visual SLAM algorithms for mobile robots are solved, achieving high-precision depth estimation with low computational cost.

CN122415705APending Publication Date: 2026-07-17UNIV OF SCI & TECH BEIJING

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
UNIV OF SCI & TECH BEIJING
Filing Date
2026-05-09
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing monocular vision SLAM algorithms for mobile robots suffer from problems such as difficulty in acquiring spatial information from images, inaccurate localization in dynamic environments, slow computation speed, low prediction accuracy, and large computational load in depth estimation algorithms.

Method used

A self-supervised monocular depth estimation method is adopted, and an algorithm framework is constructed using a depth estimation network and a pose estimation network. The image is reconstructed by using the depth map and the pose transformation matrix. Multi-scale feature fusion and translation operations are combined. The network model is trained and optimized using the KITTI, Make3D and AirSim datasets.

Benefits of technology

It improves the prediction accuracy and computation speed of depth estimation, reduces the number of model parameters, and is suitable for integration into visual SLAM systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122415705A_ABST
    Figure CN122415705A_ABST
Patent Text Reader

Abstract

本发明公开了一种基于视觉SLAM算法的自监督单目深度估计方法,涉及一般的图像数据处理或产生技术领域,包括以下步骤:利用深度估计网络和位姿估计网络,构建自监督单目深度估计算法框架,应用于视觉SLAM算法;通过深度估计网络以源图像作为输入提取深度图;通过位姿估计网络以源图像和源图像的关联图像作为输入,提取两个图像之间的位姿变换矩阵;利用源图像、深度图和位姿变换矩阵,对关联图像进行重建。本发明所提出的方法具有推理速度快、预测精度高等优点,适合集成在视觉SLAM系统上。
Need to check novelty before this filing date? Find Prior Art