Visual Content Generation With Thinking-Stage Large Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing large models generate visual content with poor quality that fails to meet user requirements effectively.
Innovation Solution
Incorporate a thinking stage in the large model process to generate thinking process information based on user instructions, utilizing multimodal large models, and employ training methods like autoregressive training and reinforcement learning to enhance the accuracy of visual content generation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If traditional large models are used to generate visual content directly from instructions, then the generation process is simple and fast, but the quality and accuracy of generated visual content is poor
Solution Approach 1:
The patent segments the visual content generation process into distinct stages: instruction understanding, thinking process generation, and visual content generation. By introducing an intermediate thinking process stage that separates reasoning from generation, the model can systematically analyze instructions and plan before producing visual content, thereby improving quality without requiring complete process redesign
Solution Approach 2:
The patent applies preliminary action by generating thinking process information before actual visual content generation. The model first produces reasoning traces, analysis, and planning information that guide subsequent generation steps, ensuring that the final visual output is based on thorough preliminary consideration of the instruction
2Measurement precision
If thinking process information is generated before visual content, then the accuracy of visual content generation is improved, but the generation time and computational resources increase
Solution Approach 1:
The patent maintains continuity of useful action by making the thinking process generation and visual content generation an integrated, continuous process rather than separate discrete steps. The thinking process tokens are generated autoregressively and immediately fed into the visual generation pipeline, eliminating idle time and ensuring that the computational work flows continuously from reasoning to generation
Data Source
AI summary
Large model-based visual content generation and target large model training methods, relating to artificial intelligence fields such as deep learning, a large model, computer vision and natural language processing, are provided. A large model-based visual content generation method may include: obtaining target instruction information; inputting the target instruction information into a target large model to obtain and output corresponding target result information, where the target result information includes target visual content, the target result information is generated by the target large model according to target thinking information, and the target thinking information is thinking process information generated by the target large model for the target instruction information.


