Voice Analysis System Excluding Non-Speech Audio
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing communication session recording technologies include substantial non-substantive content like music on hold and IVR queries, which complicate accurate speech recognition and phonetic analysis, and consume valuable recording resources without providing useful information.
Innovation Solution
A voice analysis system that identifies and excludes non-speech components, such as music and IVR queries, from communication session recordings, allowing for more efficient processing and resource allocation to speech analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If entire communication sessions including music on hold and IVR queries are recorded, then complete communication data is captured, but processing accuracy and resource efficiency deteriorate
Solution Approach 1:
The patent extracts and removes non-speech components (music on hold, IVR queries, announcements) from the recorded communication session, retaining only the relevant speech portions for analysis. This extraction process directly improves speech recognition accuracy by eliminating distracting non-speech audio while still preserving complete communication data through metadata tracking of removed segments.
Solution Approach 2:
The patent segments the communication session into distinct speech and non-speech portions, allowing selective processing of only the speech segments for phonetic analysis while maintaining awareness of the complete session through metadata. This segmentation enables improved measurement precision without losing information about the full communication flow.
2Loss of information
If entire communication sessions including music on hold and IVR queries are recorded, then complete communication data is captured, but processing time and resource consumption increase
Solution Approach 1:
By extracting and removing non-speech components (music, IVR queries, announcements) from the recording, the system reduces the amount of data requiring processing while maintaining complete communication data through metadata. This directly reduces processing time and resource consumption without losing information about the full session.
Solution Approach 2:
The system performs preliminary identification and marking of non-speech portions during or after recording, so that subsequent processing can focus only on relevant speech segments. This preliminary action reduces the processing burden and time required for analysis while preserving complete communication data through metadata tracking.
3Productivity
If non-speech components are excluded from processing, then processing efficiency improves, but complete audio data is lost
Solution Approach 1:
The patent segments the audio stream into speech and non-speech portions, processing only the speech segments for analysis while maintaining metadata that tracks the complete audio data. This segmentation enables improved processing efficiency without losing information, as the metadata preserves awareness of the full audio session.
Solution Approach 2:
The system introduces metadata as an intermediary that tracks and preserves information about non-speech portions (music, IVR queries, announcements) separately from the audio processing. This metadata acts as a mediator that allows efficient processing of only speech data while maintaining complete information about the full communication session.
Data Source
AI summary
Systems and methods for analyzing communication sessions are provided. A representative method includes: recording the communication session; identifying those portions of the communication session not containing speech of at least one of an agent and a customer; and performing processing on the recording of the communication session based, at least in part, on whether the portions contain speech of at least one of the agent and the customer.


