Deep Tagging Non-Speech Sounds in Recorded Audio

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current content searching technologies are inadequate in locating specific portions of recorded audio and video content, as they lack the ability to effectively tag and index non-speech sounds, making it difficult for users to navigate directly to specific points within large files.

Innovation Solution

A method and system for deep tagging recorded audio that detects non-speech sounds and associates descriptive terms with their time of occurrence, allowing for the creation of searchable tags stored as metadata, enabling users to search and play back specific segments of content based on these tags.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional content searching is used, then users can search for files, but users cannot locate specific portions within large audio/video files

Engineering Contradiction:
Improvelocation precisionVSAvoidtagging system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the audio/video file into multiple portions and assigns individual tags to specific segments rather than tagging the entire file. This allows precise location identification within large files by dividing them into manageable, searchable units with descriptive tags associated with specific time codes or segment identifiers.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces tags as intermediary elements that bridge the gap between user search queries and specific file portions. These tags serve as mediators that connect descriptive keywords to precise locations within the audio/video content, enabling users to navigate directly to relevant segments without manually searching through entire files.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If deep tagging of non-speech sounds is implemented, then content accessibility improves, but processing complexity increases

Engineering Contradiction:
Improvecontent accessibilityVSAvoidsound detection system complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent extracts and identifies specific non-speech sound events from the audio track by detecting acoustic patterns that differ from speech. It separates these non-speech sounds (such as laughter, applause, or environmental noises) from the speech content and assigns descriptive tags to them, allowing independent indexing and searching of these distinct audio elements.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system automatically detects, classifies, and tags non-speech sounds without requiring manual annotation. The sound detection algorithm autonomously identifies acoustic patterns, determines their types, and generates appropriate tags, enabling the system to self-serve the tagging function and improve content accessibility without proportionally increasing operational complexity.

Inventive Principle:
Principle #25Self-service

3Speed

If searchable tags are stored as metadata, then navigation speed improves, but storage requirements increase

Engineering Contradiction:
Improvenavigation speedVSAvoidmetadata storage volume
Core Design Contradiction:
SpeedVSQuantity of substance

Solution Approach 1:

The patent stores tags as metadata associated with specific local portions of the audio/video file rather than creating separate databases or centralized indexes. Each tag is embedded locally with its corresponding segment, allowing rapid direct access to specific portions without requiring extensive external storage or complex query processing, thus improving navigation speed while minimizing additional storage requirements.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS9972340B2Deep tagging background noises
Publication Date: 2018.05.15 KYNDRYL INC
  • US9972340B2 patent drawing
  • US9972340B2 patent drawing
  • US9972340B2 patent drawing

AI summary

In a computer system for navigating to a location in recorded content, a computer receives a descriptive term or phrase associated with a searchable tag. The searchable tag corresponds to a point-in-time at which a non-speech sound occurred during the recording of recorded content of a communication between a plurality of participants. The recorded content includes speech from one or more of the plurality of participants, the descriptive term includes an automatically generated phonetic translation of the non-speech sound, and the non-speech sound was transmitted to the plurality of participants during the recording. The computer navigates to a location in the recorded content corresponding to the point-in-time at which the non-speech sound occurred.