一种基于视觉大模型的稀疏点云引导视频深度预测方法

By using visual large model-guided depth completion and high-resolution depth inference, combined with sparse point clouds and visible light video, the problems of temporal instability and lack of scale information in depth prediction are solved, and high-precision and stable video depth generation is achieved.

CN120997272BActive Publication Date: 2026-07-17ZHEJIANG UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHEJIANG UNIV
Filing Date
2025-08-13
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing depth estimation methods based on large visual models cannot handle sparse point clouds and visible light videos, resulting in a lack of temporal stability and accurate scale information in depth prediction, which limits their application in fields such as autonomous driving and 3D reconstruction.

Method used

By using visual large model-guided depth completion and high-resolution depth inference, combined with sparse point clouds and visible light video, and utilizing the local least squares method of spatiotemporal neighborhood and the temporal alignment module, temporally stable high-resolution video depth is generated.

Benefits of technology

It improves the accuracy and temporal stability of depth prediction, and can extract robust scene features from sparse point clouds and visible light videos to generate detailed high-resolution video depth.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997272B_ABST
    Figure CN120997272B_ABST
Patent Text Reader

Abstract

本发明公开了一种基于视觉大模型的稀疏点云引导视频深度预测方法,该方法的目的是借助高精度雷达设备采集的稀疏点云序列引导可见光视频进行深度预测。给定一对可见光视频和雷达设备采集的稀疏点云序列,该方法首先通过视觉大模型引导的深度补全模块,生成准确的低分辨率视频深度;然后通过基于视觉大模型的高分辨率深度推理模块,生成时域稳定与细节丰富的高分辨率视频深度。本发明通过视觉大模型的泛化能力提取场景鲁棒的可见光视频特征,并通过设计基于时空邻域的局部最小二乘方法和时域对齐的特征融合模块,从雷达设备采集的稀疏点云序列中针对性地将尺度信息注入可见光视频特征,实现时域稳定的高精度视频深度预测结果。
Need to check novelty before this filing date? Find Prior Art