Video Generation Using Text and Image Features

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for sharing life status on social platforms are limited, as they often rely on single pieces of information such as geographical location, text, or images, failing to simultaneously meet user requirements in vision and hearing.

Innovation Solution

A video generation method and apparatus that automatically generates a video based on text and image information, using a target generator network to create a rich and dynamic video that can be shared in real-time.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a video is shot manually by the user, then the video content can be customized, but it requires time consumption and is limited by user shooting technology and conditions

Engineering Contradiction:
Improvevideo content customizationVSAvoidtime consumption for manual shooting
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent uses pre-stored video templates as copies that can be automatically selected and applied based on user input (text or images), eliminating the need for manual video shooting while maintaining content customization. The system copies relevant elements from templates to generate personalized videos.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs automatic video generation through self-service mechanisms by processing user input (text/images) and automatically selecting appropriate video templates and content elements, removing the need for user intervention in the complex video editing process.

Inventive Principle:
Principle #25Self-service

2Productivity

If images are directly used to synthesize a video, then the video can be generated quickly, but it is limited to switching-type showing in slides and lacks content richness

Engineering Contradiction:
Improvevideo generation speedVSAvoidvideo content richness
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent combines multiple types of content elements (video clips, images, text, audio) into a composite video structure, where different media types are integrated to create rich and diverse video content rather than simply sliding through images.

Inventive Principle:
Principle #40Composite materials

Solution Approach 2:

The video template system is designed to be multi-functional, supporting various content types (text, images, video clips, audio) and automatic selection based on user input, enabling the same template framework to generate diverse video content for different scenarios.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Ease of operation

If single pieces of information (geographical location, text, or image) are shared, then the sharing process is simple, but user requirements in vision and hearing cannot be simultaneously met

Engineering Contradiction:
Improvesharing process simplicityVSAvoidmulti-sensory information completeness
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent merges multiple information types (text, images, video clips, audio) into a single integrated video output that simultaneously satisfies visual and auditory requirements, while the user only needs to provide simple input (text or images).

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The video template system acts as an intermediary that transforms simple user input (text/images) into rich multi-sensory video content, bridging the gap between ease of operation and information completeness.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12225271B2Video generation method and related apparatus
Publication Date: 2025.02.11 HUAWEI TECH CO LTD
  • US12225271B2 patent drawing
  • US12225271B2 patent drawing
  • US12225271B2 patent drawing

AI summary

A video generation method may be applied to the field of image processing and video generation in the field of artificial intelligence. The method includes: receiving a video generation instruction, and obtaining text information and image information in response to the video generation instruction, where the text information includes one or more keywords, and the image information includes N images; obtaining, based on the one or more keywords, one or more image features that is in each of the N images and corresponds to the one or more keywords; and inputting the one or more keywords and the one or more image features of the N images into a target generator network to generate a target video, where the target video includes M images, and the M images are generated based on the one or more image features of the N images and correspond to the one or more keywords.