Deep Tagging Non-Speech Sounds in Recorded Audio
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current content searching technologies are inadequate in locating specific portions of recorded audio and video content, as they lack the ability to effectively tag and index non-speech sounds, making it difficult for users to navigate directly to specific points within large files.
Innovation Solution
A method and system for deep tagging recorded audio that detects non-speech sounds and associates descriptive terms with their time of occurrence, allowing for the creation of searchable tags stored as metadata, enabling users to search and play back specific segments of content based on these tags.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional content searching is used, then users can search for files, but users cannot locate specific portions within large audio/video files
Solution Approach 1:
The patent segments the audio/video file into multiple portions and assigns individual tags to specific segments rather than tagging the entire file. This allows precise location identification within large files by dividing them into manageable, searchable units with descriptive tags associated with specific time codes or segment identifiers.
Solution Approach 2:
The patent introduces tags as intermediary elements that bridge the gap between user search queries and specific file portions. These tags serve as mediators that connect descriptive keywords to precise locations within the audio/video content, enabling users to navigate directly to relevant segments without manually searching through entire files.
2Ease of operation
If deep tagging of non-speech sounds is implemented, then content accessibility improves, but processing complexity increases
Solution Approach 1:
The patent extracts and identifies specific non-speech sound events from the audio track by detecting acoustic patterns that differ from speech. It separates these non-speech sounds (such as laughter, applause, or environmental noises) from the speech content and assigns descriptive tags to them, allowing independent indexing and searching of these distinct audio elements.
Solution Approach 2:
The system automatically detects, classifies, and tags non-speech sounds without requiring manual annotation. The sound detection algorithm autonomously identifies acoustic patterns, determines their types, and generates appropriate tags, enabling the system to self-serve the tagging function and improve content accessibility without proportionally increasing operational complexity.
3Speed
If searchable tags are stored as metadata, then navigation speed improves, but storage requirements increase
Solution Approach 1:
The patent stores tags as metadata associated with specific local portions of the audio/video file rather than creating separate databases or centralized indexes. Each tag is embedded locally with its corresponding segment, allowing rapid direct access to specific portions without requiring extensive external storage or complex query processing, thus improving navigation speed while minimizing additional storage requirements.
Data Source
AI summary
In a computer system for navigating to a location in recorded content, a computer receives a descriptive term or phrase associated with a searchable tag. The searchable tag corresponds to a point-in-time at which a non-speech sound occurred during the recording of recorded content of a communication between a plurality of participants. The recorded content includes speech from one or more of the plurality of participants, the descriptive term includes an automatically generated phonetic translation of the non-speech sound, and the non-speech sound was transmitted to the plurality of participants during the recording. The computer navigates to a location in the recorded content corresponding to the point-in-time at which the non-speech sound occurred.


