Video Communication Image Converter for Speaker Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video communication systems face challenges in identifying speakers and conveying delicate facial expressions and environmental information, leading to increased processing burdens and potential misinterpretation in professional settings, particularly when animation characters or avatars are used.

Innovation Solution

A video communication system that converts image data into picture-like images using image feature quantities, reducing the number of colors and processing requirements, allowing for better identification and communication without the need for face recognition and feature point extraction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If animation characters or avatars are displayed on the videophone, then the speaker can be represented in a creative manner, but the speaker identification becomes difficult and professional communication is adversely affected

Engineering Contradiction:
Improvecreative representationVSAvoidspeaker identification
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent applies parameter changes by transforming image data through color space conversion (RGB to HSV and subsequent processing) to create picture-like images that maintain recognizability while reducing realism. This parameter transformation allows the image to be neither fully realistic nor completely abstract, resolving the contradiction between creative representation and speaker identification.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If face recognition and feature point extraction are performed to generate virtual characters, then the speaker can be represented, but the processing load increases and high-performance CPU is required

Engineering Contradiction:
Improvevirtual character generationVSAvoidprocessing load
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent extracts only the essential processing steps needed for picture-like image generation, specifically color space conversion and image processing, while omitting complex face recognition and feature point extraction algorithms. This selective extraction reduces processing load while maintaining the ability to generate representative images.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent uses simpler, more efficient image processing algorithms that can be implemented with standard CPU capabilities rather than requiring high-performance processors. The processing approach uses readily available computational resources and standard image processing techniques, making the system more accessible and reducing hardware requirements.

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

3Ease of operation

If animation characters or avatars are displayed, then the communication can be made more engaging, but delicate facial expressions and environmental information are difficult to convey

Engineering Contradiction:
Improvecommunication engagementVSAvoidfacial expression detail
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent applies parameter changes through color space transformation and image processing that preserve essential visual information while creating a picture-like appearance. This allows delicate facial expressions and environmental details to remain visible and conveyable, unlike fully animated characters that would lose such information.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS8164613B2Video communication system, terminal, and image converter
Publication Date: 2012.04.24 NEC CORP
  • US8164613B2 patent drawing
  • US8164613B2 patent drawing
  • US8164613B2 patent drawing

AI summary

Image input device (57) of a mobile phone captures an image of the face of the speaker and stores the captured image data in image memory (53). Communication image generator (52) reads the image data stored in image memory (53) and converts the image data into illustration image data representing an illustration-like image of the speaker. Communication image generator (52) stores the illustration image data in image memory (53). Central controller (51) reads the illustration image data from image memory (53), and sends the illustration image data via wireless device (54) and antenna (59). A mobile phone of the party who the speaker is talking to receives the illustration image data, and displays an illustration-like image of the speaker based on the image data.