Avatar Generation Using Invariant and Equivariant Feature Sets

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current 3D face reconstruction methods, such as those using point cloud technology, are inefficient due to the need for large amounts of annotated data and fail to accurately represent invariant features of a face, resulting in unsatisfactory avatar generation in applications like teleconferencing and entertainment.

Innovation Solution

A method that generates an avatar by identifying correlation among image, audio, and text information in a video to create feature sets representing invariant and equivariant features, allowing for more accurate and efficient avatar generation without relying on point cloud technology.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If point cloud technology is used for 3D face reconstruction, then avatar generation can be achieved, but large amounts of annotated data are required which reduces efficiency and increases processing cost

Engineering Contradiction:
Improveavatar generation accuracyVSAvoidface reconstruction efficiency
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent extracts only the essential features needed for avatar generation from the video input, rather than processing complete point cloud data. It identifies and extracts invariant features (identity characteristics) and equivariant features (expression and pose characteristics) separately, eliminating the need for large amounts of annotated training data while maintaining reconstruction accuracy

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments the face representation into two distinct feature sets: invariant features that remain consistent across different expressions and poses, and equivariant features that change with expressions and poses. This segmentation allows independent processing of each feature type, improving efficiency without sacrificing avatar generation quality

Inventive Principle:
Principle #1Segmentation

2Manufacturing precision

If point cloud technology is used for 3D face reconstruction, then avatar generation can be achieved, but processing cost increases due to large data requirements

Engineering Contradiction:
Improveavatar generation accuracyVSAvoiddata annotation volume
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential features needed for avatar generation from the video input, rather than processing complete point cloud data. It identifies and extracts invariant features (identity characteristics) and equivariant features (expression and pose characteristics) separately, eliminating the need for large amounts of annotated training data while maintaining reconstruction accuracy

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent uses lightweight feature representations instead of heavy point cloud data structures. By representing faces as compact feature vectors rather than dense point clouds, it reduces both storage requirements and processing costs while maintaining the necessary information for accurate avatar generation

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

3Productivity

If existing techniques are used for avatar generation, then processing can be performed, but invariant features of the face cannot be accurately represented

Engineering Contradiction:
Improveprocessing capabilityVSAvoidinvariant feature representation accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent segments the face representation into two distinct feature sets: invariant features that remain consistent across different expressions and poses, and equivariant features that change with expressions and poses. This segmentation allows independent processing of each feature type, improving efficiency without sacrificing avatar generation quality

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different processing strategies to different feature types based on their specific characteristics. Invariant features are processed to preserve identity information, while equivariant features are processed to capture dynamic expressions and poses. This localized quality approach ensures each feature type is represented with appropriate precision

Inventive Principle:
Principle #3Local quality

4Ease of operation

If existing techniques are used for avatar generation, then processing can be performed, but correlation in the input information is not utilized

Engineering Contradiction:
Improveprocessing simplicityVSAvoidinput information correlation
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent merges multiple input modalities (video, audio, and text information) into a unified feature representation. By combining these different types of input data and their correlations, the system creates a more comprehensive and accurate avatar representation than would be possible using any single input type alone

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS12014454B2Method, electronic device, and computer program product for generating avatar
Publication Date: 2024.06.18 DELL PROD LP
  • US12014454B2 patent drawing
  • US12014454B2 patent drawing
  • US12014454B2 patent drawing

AI summary

Embodiments of the present disclosure relate to a method, an electronic device, and a computer program product for generating an avatar. The method includes generating an indication of correlation among image information, audio information, and text information of a video. The method may further include generating, based on the indication of the correlation, a first feature set and a second feature set representing features of a target object in the video, wherein the first feature set represents invariant features of the target object in the video, and the second feature set represents equivariant features of the target object in the video. The method may further include generating the avatar based on the first feature set and the second feature set. With this method, the generated avatar can be made more accurate and vivid with a better effect, while also reducing data annotation cost, improving operation efficiency, and enhancing user experience.