Speech-Based Content Tagging for Automatic Person Identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing electronic devices require manual input of tags for content, such as photos and videos, which is cumbersome and time-consuming, especially for adding person tags to videos.
Innovation Solution
An electronic device that stores person tag information using speech data, automatically associating speech data with user information and storing it as tag information during content generation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual input procedure is used to add person tags to content, then tag information can be stored in content, but the operation process becomes complicated and time-consuming
Solution Approach 1:
The patent replaces the manual mechanical input process with an automatic speech recognition system. The electronic device collects speech data from the content, performs speech recognition to identify persons, and automatically generates and stores person tags without requiring manual input operations.
Solution Approach 2:
The system enables self-service tagging by automatically analyzing the content's speech data and generating person tags independently. The electronic device performs speech recognition, matches speech patterns with registered user information, and stores the generated tags back into the content without external manual intervention.
2Reliability
If manual input procedure is used to add person tags to content, then tag information can be stored in content, but time consumption increases
Solution Approach 1:
The patent implements preliminary action by pre-registering user speech patterns and personal information in the system before actual tagging is needed. When content requires tagging, the system directly compares speech data against the pre-established database, significantly reducing the time required for person identification and tag generation.
Solution Approach 2:
The manual tagging process is replaced with automated speech recognition and pattern matching systems that process and analyze content continuously, eliminating the time-consuming manual steps of watching content, identifying persons, and entering tag information.
3Ease of operation
If automatic speech-based tagging is implemented, then tagging process is simplified and time is reduced, but device complexity increases
Solution Approach 1:
The patent applies universality by designing a multi-functional system that integrates speech data collection, speech recognition, pattern matching, and tag generation within a single electronic device framework. The same speech recognition engine serves multiple purposes including content analysis, person identification, and tag creation, reducing overall system complexity despite the advanced capabilities.
Solution Approach 2:
The system introduces a speech pattern database as an intermediary layer between the speech recognition engine and the tag generation process. This database stores pre-processed user speech patterns and serves as a reference for matching, simplifying the complexity by organizing data in a structured intermediate format that facilitates efficient comparison and identification.
Data Source
AI summary
An electronic device according to an embodiment comprises a memory, a display, and a processor operatively connected to the memory and the display, wherein the processor may be configured to: collect speech data; match the collected speech data with user information related to the collected speech data and store, in the memory, association information between the collected speech data and the user information; when generating content, detect speech data of the content input that is input during generation of the content; and when there is user information matching with the detected speech data in the memory, store the user information matching with the detected speech data of the content as tag information of the content.


