Automatic Person Identification in Media Using Speaker and Image Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current communication devices lack efficient methods for automatically identifying and tagging individuals in pictures and videos, making it difficult to organize and share media content effectively.

Innovation Solution

The method involves using speaker recognition and image recognition to identify individuals in pictures and videos, storing associated information, and tagging them for easy retrieval and sharing, with features like automatic email sending and organized storage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual tagging of individuals in pictures and videos is used, then user control and accuracy are improved, but time consumption and operational complexity increase

Engineering Contradiction:
Improveidentification accuracyVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system automatically identifies and tags individuals in pictures and videos without requiring manual user input. The processor autonomously performs face recognition, extracts identifying information, and applies tags to media files, enabling the system to serve itself rather than requiring continuous user intervention for each tagging operation.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system pre-processes media files by automatically detecting faces and generating identifying information during or immediately after capture. By performing identification and tagging operations in advance rather than on-demand, the system reduces the time users would otherwise spend manually organizing and labeling their photo and video collections.

Inventive Principle:
Principle #10Preliminary action

2Productivity

If automatic identification using speaker recognition and image recognition is implemented, then productivity and ease of operation are improved, but device complexity increases

Engineering Contradiction:
Improvemedia organization efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The communication device integrates multiple recognition functions (speaker recognition and image recognition) into a single unified system. The processor can switch between or combine these different recognition methods depending on the situation, allowing one device to perform both voice-based and visual-based identification tasks across various media types including pictures, videos, and audio recordings.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system combines speaker recognition technology with image recognition technology in a unified identification framework. By merging these two separate recognition systems, the device can cross-validate identification results and improve overall accuracy while maintaining a single integrated processing architecture rather than requiring separate standalone systems.

Inventive Principle:
Principle #5Merging (Combining)

3Loss of information

If comprehensive tagging information is stored, then information completeness and retrieval accuracy are improved, but storage requirements and processing overhead increase

Engineering Contradiction:
Improveinformation completenessVSAvoidstorage requirements
Core Design Contradiction:
Loss of informationVSQuantity of substance

Solution Approach 1:

The system extracts only the essential identifying information (such as names, key attributes, and recognition confidence scores) from the media files and stores this metadata separately from the actual picture and video data. By extracting and storing only the critical tagging information rather than duplicating or analyzing the entire media content repeatedly, the system maintains information completeness while minimizing storage overhead.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS8144939B2Automatic identifying
Publication Date: 2012.03.27 SONY GROUP CORP
  • US8144939B2 patent drawing
  • US8144939B2 patent drawing
  • US8144939B2 patent drawing

AI summary

A method includes storing a picture or video that includes a first person. The method also includes automatically identifying the first person using speaker recognition or image recognition and tagging the picture or video with information indicating that the first person is in the picture or video.