3D pose image conditioned hierarchical diffusion framework system

The hierarchical diffusion framework system guided by 3D pose images solves the problems of insufficient spatiotemporal continuity modeling of 3D pose and inconsistency in multimodal fusion in existing technologies, and achieves improved pose control accuracy and visual quality in high-fidelity image generation.

CN122176787APending Publication Date: 2026-06-09NANTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANTONG UNIV
Filing Date
2026-01-28
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Existing technologies suffer from insufficient 3D pose spatiotemporal continuity modeling, limited precision in controlling complex joint movements, and inconsistent multimodal fusion when generating high-fidelity images, especially those with specific poses and anatomical realism. This results in inadequate pose control precision in the generated images.

Method used

A hierarchical diffusion framework system guided by 3D pose images is adopted. Through 3D pose processing unit, PoseNet network unit, bidirectional cross-attention mechanism unit and decoupled cross-attention unit, combined with key point encoding, pose rendering, PoseNet network, bidirectional cross-attention and decoupled cross-attention strategies, it realizes comprehensive encoding of 3D pose information and deep interaction of multimodal features.

Benefits of technology

It significantly improves the pose fidelity and visual quality of generated images, simplifies the user interaction process, achieves precise control of complex joint movements and consistency of multimodal feature fusion, and enhances the semantic alignment and pose accuracy of generated images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122176787A_ABST
    Figure CN122176787A_ABST
Patent Text Reader

Abstract

This invention discloses a hierarchical diffusion framework system for conditional guidance based on 3D pose images, addressing the problems of insufficient spatiotemporal continuity modeling of pose, limited precision in controlling complex joint movements, and inconsistent multimodal feature fusion in existing 2D pose guidance methods for high-fidelity image generation. This invention innovatively designs a 3D pose processing unit, a PoseNet network, a bidirectional cross-attention mechanism, and a decoupled cross-attention unit, achieving deep interaction and multimodal fusion between text features and pose information, significantly improving the pose fidelity, cross-modal consistency, and visual quality of the generated images. This invention requires only text prompts during inference, eliminating the need for additional pose guidance, simplifying the user interaction process while maintaining pose accuracy, demonstrating significant technological innovation and application value.
Need to check novelty before this filing date? Find Prior Art