Voice Tag Image Management System

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The inefficiency of searching for specific photos among a large collection in electronic devices, such as smartphones and tablets, due to the lack of effective tagging and categorization methods, despite advancements in camera technology that increase photo capture frequency.

Innovation Solution

An electronic device equipped with a voice input module to obtain voice data, analyze it for metadata information, and register voice tags for specific images, which can then be assigned to similar images based on determined metadata conditions, facilitating efficient search and organization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional photo management methods (folder arrangement, date sorting) are used, then images can be organized by basic criteria, but searching for specific photos among large collections becomes inefficient

Engineering Contradiction:
Improvephoto search efficiencyVSAvoidtime to find desired photo
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical search methods (browsing folders, scrolling through thumbnails) with voice-based acoustic input. Users can speak natural language queries like 'show me photos of my cat' instead of manually navigating through thousands of images, dramatically reducing search time and improving efficiency.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces voice tags as an intermediary layer between photos and search queries. These voice tags serve as mediators that capture semantic meaning from user speech and link it to photo metadata, enabling more intuitive and accurate photo retrieval compared to conventional keyword tags.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If voice tags are applied to all similar images, then search accuracy improves, but processing time and computational resources increase

Engineering Contradiction:
Improvephoto identification accuracyVSAvoidtime to process and register voice tags
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies voice tags selectively rather than uniformly to all images. The system identifies specific photos that match the voice query criteria and applies tags only to those relevant images, rather than processing the entire photo library. This localized approach maintains high search accuracy while reducing unnecessary processing time.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent implements a balanced approach where voice tags are applied to a subset of images that are most relevant to the query. Rather than tagging every possible matching image (excessive action) or being too restrictive (insufficient action), the system identifies and tags the optimal subset of photos that best satisfy the voice search criteria, achieving good results with reasonable processing time.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentEP3010219B1Method and apparatus for managing images using a voice tag
Publication Date: 2020.12.02 SAMSUNG ELECTRONICS CO LTD
  • EP3010219B1 patent drawingFigure 1
  • EP3010219B1 patent drawingFigure 2
  • EP3010219B1 patent drawingFigure 3

AI summary

An electronic device is provided. The electronic device includes a voice input module which receives a voice from an outside to generate voice data, a memory which stores one or more images or videos, and a processor which is electrically connected to the voice input module and the memory. The memory includes instructions, when executed by the processor, causing the electronic device to link at least one of the voice data, the first metadata information based on the voice data, or second metadata information generated from the voice data and/or the first metadata information with the second image or video.