Controllable clothing image generation method for enhancing multi-modal complementary attention based on retrieval

By constructing a multimodal clothing dataset and introducing cross-modal retrieval and hierarchical guidance modules, the problem of discrepancies between clothing image generation results and text descriptions in existing technologies has been solved, thereby improving the detail reproduction and semantic consistency of clothing images.

CN122087142APending Publication Date: 2026-05-26XI'AN POLYTECHNIC UNIVERSITY

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XI'AN POLYTECHNIC UNIVERSITY
Filing Date
2025-12-25
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing clothing image generation technologies struggle to accurately interpret complex clothing text descriptions, lack prior knowledge of clothing structure and style, produce results that deviate from the text descriptions, and lack controllable generation capabilities.

Method used

A multimodal clothing dataset is constructed, and a cross-modal retrieval module and a retrieval hierarchical guidance module are introduced. The cross-modal retrieval module maps text descriptions and clothing images to a unified feature space. Multimodal feature fusion is performed using an appearance context encoder and contextual semantic awareness attention to improve semantic understanding and visual detail reconstruction capabilities.

Benefits of technology

It improves the detail reproduction and semantic consistency in the clothing image generation process, making the generated clothing images and text descriptions more consistent with expectations and the appearance more realistic and natural.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122087142A_ABST
    Figure CN122087142A_ABST
Patent Text Reader

Abstract

The invention discloses a controllable clothing image generation method for enhancing multi-modal complementary attention based on retrieval, and the method specifically comprises the following steps: 1, constructing a multi-modal clothing data set which comprises a clothing image and a text description corresponding to the clothing image, and the data set comprises a training set and a test set; step 2, constructing a retrieval enhancement multi-modal complementary attention network used for generating a clothing image, wherein the network comprises a cross-modal retrieval module and a retrieval hierarchical guide module; 3, training a retrieval hierarchical guide module by adopting text description in a training set and a retrieval clothing image output by a cross-modal retrieval module to obtain a trained model weight; and 4, testing the weight of the model trained in the step 3 by using a test set, and outputting a generated clothing image. According to the invention, the problems that the generation result does not conform to the input text description and the appearance details are not clear in the existing clothing image generation method are solved.
Need to check novelty before this filing date? Find Prior Art