Speech-Based Content Tagging for Automatic Person Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing electronic devices require manual input of tags for content, such as photos and videos, which is cumbersome and time-consuming, especially for adding person tags to videos.

Innovation Solution

An electronic device that stores person tag information using speech data, automatically associating speech data with user information and storing it as tag information during content generation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual input procedure is used to add person tags to content, then tag information can be stored in content, but the operation process becomes complicated and time-consuming

Engineering Contradiction:
Improvetag information storageVSAvoidtagging process
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent replaces the manual mechanical input process with an automatic speech recognition system. The electronic device collects speech data from the content, performs speech recognition to identify persons, and automatically generates and stores person tags without requiring manual input operations.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service tagging by automatically analyzing the content's speech data and generating person tags independently. The electronic device performs speech recognition, matches speech patterns with registered user information, and stores the generated tags back into the content without external manual intervention.

Inventive Principle:
Principle #25Self-service

2Reliability

If manual input procedure is used to add person tags to content, then tag information can be stored in content, but time consumption increases

Engineering Contradiction:
Improvetag information storageVSAvoidtagging time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements preliminary action by pre-registering user speech patterns and personal information in the system before actual tagging is needed. When content requires tagging, the system directly compares speech data against the pre-established database, significantly reducing the time required for person identification and tag generation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The manual tagging process is replaced with automated speech recognition and pattern matching systems that process and analyze content continuously, eliminating the time-consuming manual steps of watching content, identifying persons, and entering tag information.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Ease of operation

If automatic speech-based tagging is implemented, then tagging process is simplified and time is reduced, but device complexity increases

Engineering Contradiction:
Improvetagging processVSAvoidsystem structure
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent applies universality by designing a multi-functional system that integrates speech data collection, speech recognition, pattern matching, and tag generation within a single electronic device framework. The same speech recognition engine serves multiple purposes including content analysis, person identification, and tag creation, reducing overall system complexity despite the advanced capabilities.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system introduces a speech pattern database as an intermediary layer between the speech recognition engine and the tag generation process. This database stores pre-processed user speech patterns and serves as a reference for matching, simplifying the complexity by organizing data in a structured intermediate format that facilitates efficient comparison and identification.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12620398B2Electronic device stores tag information of content
Publication Date: 2026.05.05 SAMSUNG ELECTRONICS CO LTD
  • US12620398B2 patent drawing
  • US12620398B2 patent drawing
  • US12620398B2 patent drawing

AI summary

An electronic device according to an embodiment comprises a memory, a display, and a processor operatively connected to the memory and the display, wherein the processor may be configured to: collect speech data; match the collected speech data with user information related to the collected speech data and store, in the memory, association information between the collected speech data and the user information; when generating content, detect speech data of the content input that is input during generation of the content; and when there is user information matching with the detected speech data in the memory, store the user information matching with the detected speech data of the content as tag information of the content.