An end-to-end pose-aware 3d diffusion generation method
By performing end-to-end 3D generation within the observation space and utilizing monocular depth estimation and multi-head cross-modal fusion mechanisms, the problems of spatial misalignment and rotational transformation ambiguity in existing methods are solved, achieving accurate alignment of 3D objects with input images and high-quality composite scene generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- RENMIN UNIVERSITY OF CHINA
- Filing Date
- 2026-04-13
- Publication Date
- 2026-07-17
AI Technical Summary
Existing methods suffer from spatial misalignment and rotational transformation ambiguity in pose-aware 3D generation tasks, making it difficult to accurately align the generated 3D objects with the input image, especially in poor handling of occlusion relationships in combined multi-object scenes.
Employing a pose-aware diffusion model (PAD), end-to-end 3D generation is performed directly in the observation space. Local point clouds are extracted and injected into the latent space through a monocular depth estimation model. Combined with a multi-head cross-modal fusion mechanism and a depth noise enhancement strategy, the generated results are ensured to be geometrically aligned with the input image at the pixel level. Instance segmentation is used to process multi-object scenes.
It achieves precise alignment of the generated 3D object with the input image in pose and space, improves the spatial consistency and robustness of the generated result, supports high-quality composite 3D scene generation, and effectively handles the occlusion relationship between objects.
Smart Images

Figure CN122415863A_ABST