Voice-Text Modality Conversion for Customer Service Agents

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing customer service systems lack seamless integration between voice and text communications, leading to inefficiencies when switching between modalities and raising concerns about data security when granting AI agents access to production databases.

Innovation Solution

The system combines Speech to Text, a language model, and Text to Speech service agents to enable voice calls to be responded to via text, allowing for seamless transitions between chat-based and voice-based services while maintaining a voice-based connection with users.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If AI agents are given access to production databases to respond to both voice and text communications, then the ability to address customer service needs is improved, but security risks increase

Engineering Contradiction:
Improveability to respond to voice and text communicationsVSAvoidsecurity risks
Core Design Contradiction:
Adaptability or versatilityVSObject-affected harmful factors

Solution Approach 1:

The patent creates a copy of the production database specifically for AI agent operations. This copy allows the AI to query and respond to customer service needs without direct access to the live production database, thereby maintaining security while enabling versatile responses to both voice and text communications.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system introduces an intermediary layer between the AI agent and the production database. This intermediary manages data access and communication, allowing the AI to function with database access while maintaining security protocols. The intermediary translates AI requests into safe database operations.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If human agents monitor both voice and text communications, then they can respond to both modalities, but efficiency decreases compared to taking voice calls directly

Engineering Contradiction:
Improveability to respond via both voice and textVSAvoidagent efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The system enables self-service by allowing the AI agent to autonomously handle both voice and text communications without requiring human agents to manually monitor both modalities. The AI independently processes incoming communications, converts between modalities as needed, and responds appropriately, freeing human agents from inefficient multitasking.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical process of human agents manually monitoring and switching between voice and text channels with an automated AI system. The AI uses speech-to-text conversion and text-to-speech synthesis to bridge modalities automatically, eliminating the inefficiency of human multitasking while maintaining adaptability to both communication types.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Productivity

If chat-based services are used for customer support, then agents can multitask more effectively, but seamless switching between chat and voice services is not achieved

Engineering Contradiction:
Improveagent multitasking capabilityVSAvoidseamless switching between modalities
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The system achieves universality by enabling a single AI agent to handle multiple communication modalities (voice and text) through a unified interface. The AI can receive voice calls, convert them to text for processing, generate text responses, and convert responses back to voice, allowing seamless switching without requiring separate systems for each modality.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system dynamically changes the communication parameter (modality) based on user needs. Speech-to-text conversion allows voice input to be processed as text, while text-to-speech conversion allows text responses to be delivered as voice. This parameter transformation enables seamless switching between modalities while maintaining the efficiency of chat-based processing.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250126207A1Framework for modality conversion between phone and chat conversations
Publication Date: 2025.04.17 ZHU RICHARD
  • US20250126207A1 patent drawing
  • US20250126207A1 patent drawing
  • US20250126207A1 patent drawing

AI summary

Systems and methods combining Speech to Text, a language model, and Text to Speech service agents to accept voice phone calls and respond to those voice phone calls via text, thereby enabling the multitasking benefits of chat-based while maintaining a voice-based connection with a user.