Controllable clothing image generation method for enhancing multi-modal complementary attention based on retrieval
By constructing a multimodal clothing dataset and introducing cross-modal retrieval and hierarchical guidance modules, the problem of discrepancies between clothing image generation results and text descriptions in existing technologies has been solved, thereby improving the detail reproduction and semantic consistency of clothing images.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XI'AN POLYTECHNIC UNIVERSITY
- Filing Date
- 2025-12-25
- Publication Date
- 2026-05-26
AI Technical Summary
Existing clothing image generation technologies struggle to accurately interpret complex clothing text descriptions, lack prior knowledge of clothing structure and style, produce results that deviate from the text descriptions, and lack controllable generation capabilities.
A multimodal clothing dataset is constructed, and a cross-modal retrieval module and a retrieval hierarchical guidance module are introduced. The cross-modal retrieval module maps text descriptions and clothing images to a unified feature space. Multimodal feature fusion is performed using an appearance context encoder and contextual semantic awareness attention to improve semantic understanding and visual detail reconstruction capabilities.
It improves the detail reproduction and semantic consistency in the clothing image generation process, making the generated clothing images and text descriptions more consistent with expectations and the appearance more realistic and natural.
Smart Images

Figure CN122087142A_ABST