HMD Facial Expression Synthesis via Mouth Image Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Head-Mounted Displays (HMDs) obstruct the player's face during live game broadcasts, making it difficult for viewers to see the player's expressions, which reduces the immersive experience and enjoyment.

Innovation Solution

An information processing system that acquires images of a user wearing an HMD, estimates the player's expression from their mouth, and produces a facial image to be synthesized with the live game image, allowing the expression to be visible to viewers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Illumination intensity

If a user wears an HMD to play games and broadcast to viewers, then the user's immersion in the video world is increased, but the user's facial expressions are hidden from viewers

Engineering Contradiction:
Improveimmersion in video worldVSAvoidfacial expressions
Core Design Contradiction:
Illumination intensityVSLoss of information

Solution Approach 1:

The patent creates a virtual avatar that copies and represents the user's facial expressions. The system captures the user's mouth movements through a camera, estimates the facial expression, and applies it to a virtual avatar image that is then broadcast to viewers. This allows the user's expressions to be visible without removing the HMD, thus preserving both immersion and expression visibility.

Inventive Principle:
Principle #26Copying

2Adaptability or versatility

If the HMD covers both eyes and nose to provide immersive VR experience, then the sense of immersion is increased, but the photographed image of the player hides the great part of the expression

Engineering Contradiction:
Improvesense of immersionVSAvoidfacial expression visibility
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent introduces a virtual avatar as an intermediary between the user's actual face and the viewers. Instead of showing the user's real face (which is obscured by the HMD), the system uses the avatar to mediate and convey the user's expressions. The avatar serves as a substitute that preserves the user's identity and expressions while maintaining the immersive HMD experience.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Device complexity

If only the mouth is used for expression estimation, then the system can work with limited visible area, but the expression estimation must be accurate from partial information

Engineering Contradiction:
Improvesystem simplicityVSAvoidexpression estimation accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent employs machine learning algorithms that enable the system to automatically learn and recognize facial expressions from mouth images alone. The pre-trained model self-adapts to extract meaningful expression information from the limited visible area, performing the complex task of accurate expression recognition without requiring additional sensors or more visible facial areas.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS10896322B2Information processing device, information processing system, facial image output method, and program
Publication Date: 2021.01.19 SONY INTERACTIVE ENTERTAINMENT LLC
  • US10896322B2 patent drawing
  • US10896322B2 patent drawing
  • US10896322B2 patent drawing

AI summary

There is provided an information processing device including an image acquiring portion configured to acquire a photographed image obtained by photographing a user wearing a head mounted display, an expression estimating portion configured to estimate an expression of the user from an image of a mouth of the user included in the photographed image, an facial image producing portion configured to produce a facial image responding to the estimated expression of the user, and an output portion configured to output an image including the facial image.