An end-to-end pose-aware 3d diffusion generation method

By performing end-to-end 3D generation within the observation space and utilizing monocular depth estimation and multi-head cross-modal fusion mechanisms, the problems of spatial misalignment and rotational transformation ambiguity in existing methods are solved, achieving accurate alignment of 3D objects with input images and high-quality composite scene generation.

CN122415863APending Publication Date: 2026-07-17RENMIN UNIVERSITY OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
RENMIN UNIVERSITY OF CHINA
Filing Date
2026-04-13
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing methods suffer from spatial misalignment and rotational transformation ambiguity in pose-aware 3D generation tasks, making it difficult to accurately align the generated 3D objects with the input image, especially in poor handling of occlusion relationships in combined multi-object scenes.

Method used

Employing a pose-aware diffusion model (PAD), end-to-end 3D generation is performed directly in the observation space. Local point clouds are extracted and injected into the latent space through a monocular depth estimation model. Combined with a multi-head cross-modal fusion mechanism and a depth noise enhancement strategy, the generated results are ensured to be geometrically aligned with the input image at the pixel level. Instance segmentation is used to process multi-object scenes.

Benefits of technology

It achieves precise alignment of the generated 3D object with the input image in pose and space, improves the spatial consistency and robustness of the generated result, supports high-quality composite 3D scene generation, and effectively handles the occlusion relationship between objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122415863A_ABST
    Figure CN122415863A_ABST
Patent Text Reader

Abstract

本公开提供一种端到端位姿感知的三维扩散生成方法。针对为图像生成位姿感知的任务,设计位姿感知扩散模型,将输入图片和条件点云信息输入所述位姿感知扩散模型,并通过解码器进行解码,得到带位姿的物体三维模型;所述位姿感知扩散模型采用流匹配框架进行三维隐变量的去噪生成,并通过隐空间点云条件注入机制,直接在观测坐标系中进行整个三维生成过程。该方案能够直接隐空间点云条件注入机制在生成的三维内容与输入图像之间建立严格的像素级几何对应关系,使模型在无需后置位姿估计的情况下即可输出位姿对齐的完整闭合三维网格。
Need to check novelty before this filing date? Find Prior Art