Image Animation via Semantic Region Motion Transfer

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional image animation methods generate dull and rigid videos by applying pre-obtained motion patterns to entire images without considering semantic differences, making it difficult for users to visualize and achieve desired output videos.

Innovation Solution

The method involves obtaining an input image and a reference video, determining the motion pattern of a reference object in the reference video, and generating an output video with the input image as a starting frame, where the motion of the target object in the output video follows the motion pattern of the reference object, allowing for intuitive application of motion patterns and high-quality video generation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If pre-obtained motion patterns are applied to entire images, then the animation generation process is simplified, but the video quality becomes dull and rigid

Engineering Contradiction:
Improveanimation generation processVSAvoidvideo quality
Core Design Contradiction:
Ease of manufactureVSManufacturing precision

Solution Approach 1:

The patent segments the image into multiple semantic regions (e.g., sky, water, land, objects) and applies different motion patterns to each region independently. This allows the system to maintain simplified automated generation while achieving high video quality through region-specific motion control, directly resolving the contradiction between ease of manufacture and manufacturing precision.

Inventive Principle:
Principle #1Segmentation

2Device complexity

If motion patterns are applied without considering semantic differences, then the processing complexity is reduced, but the realism and engagement of the output video deteriorates

Engineering Contradiction:
Improveprocessing complexityVSAvoidrealism and engagement
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent implements local quality by assigning different motion characteristics to different semantic regions within the image. Each region (sky, water, land, objects) receives motion patterns appropriate to its semantic type, creating realistic and engaging video output while maintaining manageable processing complexity through automated semantic segmentation and motion pattern selection.

Inventive Principle:
Principle #3Local quality

3Extent of automation

If uniform motion patterns are applied to all regions, then the automation level is high, but the user ability to visualize and achieve desired output is reduced

Engineering Contradiction:
Improveautomation levelVSAvoiduser visualization and control
Core Design Contradiction:
Extent of automationVSEase of operation

Solution Approach 1:

The patent incorporates feedback mechanisms that allow users to view the segmented semantic regions and their assigned motion patterns, and to adjust these assignments to achieve desired visual outcomes. This maintains high automation while improving ease of operation through user-friendly control interfaces that leverage the semantic segmentation results.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20240153189A1Image animation
Publication Date: 2024.05.09 MICROSOFT TECHNOLOGY LICENSING LLC
  • US20240153189A1 patent drawing
  • US20240153189A1 patent drawing
  • US20240153189A1 patent drawing

AI summary

According to implementations of the subject matter described herein, there is provided a solution for generating a video from an image. In this solution, an input image and a reference video are obtained; a motion pattern of a reference object in the reference video is determined based on the reference video. An output video with the input image as a starting frame is generated. Motion of a target object in the output video has the motion pattern of the reference object and the target object is in the input image. In this way, according to the solution, the motion pattern of the reference object in the reference video can be intuitively applied to the input image to generate the output video, and the motion of the target object in the output video has the motion pattern of the reference object.