Automated Voice Call Response with Supplemental Word Insertion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Automated voice calls often lack the natural interaction of human conversations, leading to user dissatisfaction due to noticeable delays and unnatural silences caused by processing steps, which can result in lost business opportunities.
Innovation Solution
Incorporation of supplemental words, such as filler words and short sentences, into automated voice calls to mask processing delays and enhance the natural flow of conversations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If automated voice calls use processing steps to generate responses, then productivity is improved, but the naturalness of conversation deteriorates due to noticeable delays and unnatural silences
Solution Approach 1:
The patent introduces supplemental words as an intermediary element between the automated processing system and the user. These filler words (um, uh, hmm, pauses) act as a mediator that masks the mechanical delays of processing while maintaining the illusion of natural human conversation flow, thus resolving the contradiction between automated efficiency and natural interaction
2Ease of operation
If filler words are inserted to mask processing delays, then the perceived naturalness is improved, but device complexity increases due to additional word selection and insertion mechanisms
Solution Approach 1:
The system pre-selects and prepares supplemental words before they are needed during the conversation. The word selection mechanism is established in advance with predefined filler words and insertion rules, allowing the system to quickly insert appropriate words during processing delays without requiring complex real-time decision-making, thus managing complexity while improving naturalness
Data Source
AI summary
A method receives audio data from a call. Services are performed to process the audio data to automatically generate a response, wherein the services include converting the audio data to input text, inputting the input text into a model to automatically generate a text response, and converting the text response to an audio response. Supplemental words are selected based on the input text. The method determines a type of service based on services performed to generate the audio response and determines a position in the response to insert the supplemental words based on the type of service. The supplemental words are provided for insertion in the call at the position to supplement the audio response.


