Digital Avatar Generation via K-NN Graph Semantic Contexts

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies for generating digital avatars lack the capability to create hyper-realistic, AI-driven entities that can interact with users in a natural and personalized manner, offering limited sensory and cognitive capabilities compared to real humans.

Innovation Solution

A machine-learning system that processes user data to generate digital humans with hyper-realistic appearances and behaviors, utilizing deep learning models and modality relationship models to create lifelike interactions through speech, vision, and other sensory modalities, allowing for seamless communication between the digital and physical worlds.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If graphics-based re-animation methods are used to generate digital avatars, then the visual appearance can be created, but the interactions lack naturalness and personalization

Engineering Contradiction:
Improveinteraction naturalnessVSAvoidpersonalization capability
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The patent replaces traditional graphics-based re-animation mechanical systems with AI-driven systems that use deep learning models, modality relationship models, and knowledge graphs to generate hyper-realistic digital human interactions, enabling natural and personalized communication

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system changes the parameters of digital avatar generation by incorporating multiple modalities (audio, video, text) and using machine learning models to dynamically adjust behavioral parameters, semantic contexts, and interaction patterns based on user data and real-time inputs

Inventive Principle:
Principle #35Parameter changes

2Reliability

If traditional digital avatar systems are used, then basic graphical representation is achieved, but sensory and cognitive capabilities are limited

Engineering Contradiction:
Improvesensory capabilityVSAvoidsystem architecture
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the digital human system into multiple specialized components including deep learning models for specific modalities, modality relationship models for integrating different sensory inputs, and knowledge graphs for cognitive reasoning, allowing each component to be optimized independently while working together to achieve comprehensive sensory and cognitive capabilities

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system creates a universal AI-driven platform that handles multiple sensory modalities (vision, audio, text) and cognitive functions (semantic understanding, decision-making, behavior generation) through integrated machine learning models that can process and respond to various types of inputs across different contexts

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Ease of operation

If AI-driven deep learning models are implemented to create hyper-realistic digital humans, then interaction quality improves, but computational complexity and data processing requirements increase

Engineering Contradiction:
Improveinteraction qualityVSAvoidcomputational system
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by pre-processing user data, training deep learning models offline, and pre-computing modality relationships and knowledge graphs before real-time interaction, allowing the digital human to respond naturally during actual use without requiring complex real-time computation for all processing

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces intermediary components such as modality relationship models that mediate between different sensory inputs and the decision-making system, and knowledge graphs that serve as intermediaries between raw data and behavioral generation, simplifying the overall computational architecture while maintaining high interaction quality

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11544886B2Generating digital avatar
Publication Date: 2023.01.03 SAMSUNG ELECTRONICS CO LTD
  • US11544886B2 patent drawing
  • US11544886B2 patent drawing
  • US11544886B2 patent drawing

AI summary

In one embodiment, a method includes, by one or more computing systems: receiving one or more non-video inputs, where the one or more non-video inputs include at least one of a text input, an audio input, or an expression input, accessing a K-NN graph including several sets of nodes, where each set of nodes corresponds to a particular semantic context out of several semantic contexts, determining one or more actions to be performed by a digital avatar based on the one or more identified semantic contexts, generating, in real-time in response to receiving the one or more non-video inputs and based on the determined one or more actions, a video output of the digital avatar including one or more human characteristics corresponding to the one or more identified semantic contexts, and sending, to a client device, instructions to present the video output of the digital avatar.