On-Hold Call Audio Monitoring for Automatic Agent Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Users have to manually monitor voice communication sessions on hold, which increases power consumption and requires frequent device interactions, leading to battery drain and inconvenience.
Innovation Solution
An on hold client monitors the audio stream to automatically detect when a session is no longer on hold, using machine learning and audio analysis to determine the end of the hold status, and provides user interface notifications without requiring user intervention.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a caller manually monitors the on hold call by increasing volume, activating speakerphone, and repeatedly activating the screen, then the caller can determine when a service representative becomes active, but power consumption increases and battery drains expedited
Solution Approach 1:
The system performs self-service by automatically monitoring the audio stream to detect when a service representative becomes available, eliminating the need for the caller to manually increase volume, activate speakerphone, or repeatedly wake the screen. The automated voice detection and speaker diarization technologies enable the system to independently determine representative availability and notify the caller, thereby resolving the contradiction between reliable detection and energy consumption.
Solution Approach 2:
The patent replaces manual mechanical operations (volume adjustment, speakerphone activation, screen tapping) with automated acoustic field analysis. The system uses audio stream processing, voice activity detection, and speaker diarization to automatically monitor for representative availability, substituting physical user actions with electronic signal processing that consumes minimal power while maintaining reliable detection.
2Reliability
If a caller manually monitors the on hold call by making multiple inputs (volume adjustment, speakerphone activation, screen activation), then the caller can detect service representative availability, but the number of user inputs increases significantly
Solution Approach 1:
The system performs self-service by automatically monitoring the audio stream to detect when a service representative becomes available, eliminating the need for the caller to manually increase volume, activate speakerphone, or repeatedly wake the screen. The automated voice detection and speaker diarization technologies enable the system to independently determine representative availability and notify the caller, thereby resolving the contradiction between reliable detection and energy consumption.
Solution Approach 2:
The patent introduces an intermediary automated monitoring system that acts as a mediator between the caller and the service representative. This intermediary continuously analyzes the audio stream, detects voice patterns indicating representative availability, and notifies the caller through the existing interface, thereby eliminating the need for multiple manual user inputs while maintaining reliable detection capability.
3Loss of information
If music is played for a user while they are waiting on hold with interruptions by recorded voices, then the user receives information about the organization, but the user cannot easily distinguish between recorded voices and live service representatives
Solution Approach 1:
The patent replaces manual mechanical operations (volume adjustment, speakerphone activation, screen tapping) with automated acoustic field analysis. The system uses audio stream processing, voice activity detection, and speaker diarization to automatically monitor for representative availability, substituting physical user actions with electronic signal processing that consumes minimal power while maintaining reliable detection.
Solution Approach 2:
The patent introduces an intermediary automated monitoring system that acts as a mediator between the caller and the service representative. This intermediary continuously analyzes the audio stream, detects voice patterns indicating representative availability, and notifies the caller through the existing interface, thereby eliminating the need for multiple manual user inputs while maintaining reliable detection capability.
Data Source
AI summary
Automated monitoring of a voice communication session, when the session is in an on hold status, to determine when the session is no longer in the on hold status. When it is determined that the session is no longer in the on hold status, user interface output is rendered that indicates that the on hold status of the session has ceased. In some implementations, an audio stream of the session can be monitored to determine, based on processing of the audio stream, a candidate end of the on hold status. In response, a response solicitation signal is injected into an outgoing portion of the audio. The audio stream can be further monitored for a response (if any) to the response solicitation signal. The response (if any) can be processed to determine whether the end of the on hold status is an actual end of the on hold status.


