Video Generation Using Text and Image Features
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for sharing life status on social platforms are limited, as they often rely on single pieces of information such as geographical location, text, or images, failing to simultaneously meet user requirements in vision and hearing.
Innovation Solution
A video generation method and apparatus that automatically generates a video based on text and image information, using a target generator network to create a rich and dynamic video that can be shared in real-time.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a video is shot manually by the user, then the video content can be customized, but it requires time consumption and is limited by user shooting technology and conditions
Solution Approach 1:
The patent uses pre-stored video templates as copies that can be automatically selected and applied based on user input (text or images), eliminating the need for manual video shooting while maintaining content customization. The system copies relevant elements from templates to generate personalized videos.
Solution Approach 2:
The system performs automatic video generation through self-service mechanisms by processing user input (text/images) and automatically selecting appropriate video templates and content elements, removing the need for user intervention in the complex video editing process.
2Productivity
If images are directly used to synthesize a video, then the video can be generated quickly, but it is limited to switching-type showing in slides and lacks content richness
Solution Approach 1:
The patent combines multiple types of content elements (video clips, images, text, audio) into a composite video structure, where different media types are integrated to create rich and diverse video content rather than simply sliding through images.
Solution Approach 2:
The video template system is designed to be multi-functional, supporting various content types (text, images, video clips, audio) and automatic selection based on user input, enabling the same template framework to generate diverse video content for different scenarios.
3Ease of operation
If single pieces of information (geographical location, text, or image) are shared, then the sharing process is simple, but user requirements in vision and hearing cannot be simultaneously met
Solution Approach 1:
The patent merges multiple information types (text, images, video clips, audio) into a single integrated video output that simultaneously satisfies visual and auditory requirements, while the user only needs to provide simple input (text or images).
Solution Approach 2:
The video template system acts as an intermediary that transforms simple user input (text/images) into rich multi-sensory video content, bridging the gap between ease of operation and information completeness.
Data Source
AI summary
A video generation method may be applied to the field of image processing and video generation in the field of artificial intelligence. The method includes: receiving a video generation instruction, and obtaining text information and image information in response to the video generation instruction, where the text information includes one or more keywords, and the image information includes N images; obtaining, based on the one or more keywords, one or more image features that is in each of the N images and corresponds to the one or more keywords; and inputting the one or more keywords and the one or more image features of the N images into a target generator network to generate a target video, where the target video includes M images, and the M images are generated based on the one or more image features of the N images and correspond to the one or more keywords.


