Automatic Speech Recognition for Contextual Image Tagging

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Manual image and video tagging is time-consuming and inefficient, especially when dealing with large collections or unique tags, as existing automated solutions require identical tags across all media.

Innovation Solution

The method involves processing audio streams through automatic speech recognition and object recognition algorithms to automatically generate and associate keywords with images or videos, allowing for contextual tagging with minimal user input.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual tagging is used, then tagging accuracy and context relevance are improved, but time consumption and labor effort increase significantly

Engineering Contradiction:
Improvetagging accuracyVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system enables automatic tagging by having the audio content itself provide the tags through speech recognition. The audio stream processes itself to generate relevant keywords and tags, eliminating the need for manual human intervention while maintaining high accuracy in tagging.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual tagging process with an automated speech recognition system. Instead of human operators manually assigning tags, the system uses automatic speech recognition to extract keywords from audio streams and apply them as tags to video content, dramatically reducing time consumption.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If existing automated tagging solutions are used, then time consumption is reduced, but tagging versatility and uniqueness are limited due to identical tags across all media

Engineering Contradiction:
Improvetagging efficiencyVSAvoidtag uniqueness
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The system applies local quality by generating unique tags specific to each video segment based on its individual audio content. Instead of applying uniform tags across all media, the speech recognition extracts context-specific keywords that are locally relevant to each particular video portion, enabling both efficiency and uniqueness.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent segments the tagging process by processing audio streams in relation to specific video segments rather than applying blanket tags to entire collections. This segmentation allows each video segment to receive customized tags based on its unique audio content, maintaining versatility while achieving automated efficiency.

Inventive Principle:
Principle #1Segmentation

3Productivity

If automated speech recognition is implemented, then tagging speed increases, but system complexity and processing requirements increase

Engineering Contradiction:
Improvetagging speedVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system achieves universality by using the speech recognition engine for multiple purposes: transcribing speech, extracting keywords, generating tags, and identifying contextually relevant terms. This multi-functionality reduces the need for separate specialized systems, managing complexity while maintaining high tagging speed.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11715302B2Automatic tagging of images using speech recognition
Publication Date: 2023.08.01 STREEM INC
  • US11715302B2 patent drawing
  • US11715302B2 patent drawing
  • US11715302B2 patent drawing

AI summary

Methods for automatically tagging one or more images and/or video clips using a audio stream are disclosed. The audio stream may be processed using an automatic speech recognition algorithm, to extract possible keywords. The image(s) and/or video clip(s) may then be tagged with the possible keywords. In some embodiments, the image(s) and/or video clip(s) may be tagged automatically. In other embodiments, a user may be presented with a list of possible keywords extracted from the audio stream, from which the user may then select to manually tag the image(s) and/or video clip(s).