Indexing Group Communication Audio Streams
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems face challenges in accurately recognizing and searching for specific content in digital recordings of group communications, particularly when multiple speakers are involved, due to overlapping speech and reduced recording fidelity, making it difficult to identify and replay desired sections.
Innovation Solution
A method and system that index and search recognized content in media streams by processing separate audio streams from multiple speakers, using customized speech recognition models, and synchronizing them to create a combined stream with accurate location indexing, allowing for real-time search and playback of specific keywords or speakers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple speakers speak simultaneously in a group communication recording, then the communication is more natural and interactive, but the speech recognition accuracy deteriorates due to overlapping and blended speech
Solution Approach 1:
The patent segments the mixed audio signal from multiple speakers into separate speaker-specific audio streams using source separation technology. Each segmented stream is then processed independently by speech recognition engines, allowing accurate recognition even when speakers overlap in the original recording.
2Device complexity
If a single microphone is used to record group communications, then the recording device is simpler, but the speech quality and recognition accuracy worsen due to reduced fidelity and limited listening position
Solution Approach 1:
The patent introduces source separation technology as an intermediary processing step between the single microphone recording and speech recognition. This intermediary separates the mixed audio into distinct speaker streams, effectively compensating for the limitations of the simple single-microphone recording setup.
3Ease of manufacture
If users replay different sections of recorded communication to find specific topics, then no additional search technology is needed, but the time required to locate desired content increases significantly
Solution Approach 1:
The patent performs preliminary speech recognition and transcribes the audio recording into text with time stamps before the user needs to search. This preliminary conversion to searchable text format allows users to quickly find specific topics without manually replaying audio sections, significantly reducing content location time.
4Stability of the object's composition
If speech from multiple speakers is blended together in a recording, then the recording captures the natural flow of conversation, but the ability to distinguish and search for content from individual speakers deteriorates
Solution Approach 1:
The patent segments the blended speech into separate speaker streams while preserving the temporal structure of the conversation. Each speaker's contributions are identified and separated, allowing both maintenance of natural conversation flow and enabling speaker-specific content distinction and search.
Data Source
AI summary
In one embodiment, indexing content in streamed data includes receiving streams of audio data encoding a recording of a live ongoing group communication, where each stream of audio data encodes a different one of multiple voices. Each of the streams of audio data is provided to a recognizer to cause separate recognition of words in each of the streams. The recognized words are indexed to corresponding locations in each of the streams, and the streams are combined into a combined stream of audio data by synchronizing at least one common location in the streams. Embodiments allow accurate recognition of speech in group communications in which multiple speakers have simultaneously spoken, and accurate search of content encoded and processed from such speech.


