Pre-recorded Audio Navigator for Scalable Simulated Conversations

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for simulating conversations, such as voice-alike talent and speech synthesis, face challenges in scalability and cost, while voice soundboards require a large number of pre-recorded phrases, making them impractical for natural and convincing interactions, especially in large-scale deployments.

Innovation Solution

A system utilizing a pre-recorded audio navigator with a computing device, processor, and memory that selects and outputs recorded speech phrases based on context and user profile information, combined with animation cues, to provide a natural and personalized conversation experience without the need for extensive voice talent or cumbersome voice synthesis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If voice-alike talent is used to provide simulated conversations, then the quality and naturalness of voice performance is improved, but the cost and difficulty of scaling up for large-scale deployments increases significantly

Engineering Contradiction:
Improvevoice performance qualityVSAvoiddeployment complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent uses pre-recorded audio phrases as copies of authentic voice performances. Instead of requiring live voice-alike talent for each interaction, the system captures authentic voice performances once and reuses them through audio playback, eliminating the need for repeated talent recruitment and training while maintaining voice quality consistency across large-scale deployments

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs voice recording and phrase preparation in advance during an offline phase. Audio phrases are pre-recorded, transcribed, and organized into contexts before deployment. This preliminary action eliminates the need for real-time voice generation during operations, simplifying the deployment process and enabling scalable implementation without requiring voice talent on-site

Inventive Principle:
Principle #10Preliminary action

2Ease of manufacture

If a voice soundboard with pre-recorded phrases is used, then the cost is reduced, but the number of phrases required becomes very large, rendering the approach impractical

Engineering Contradiction:
Improvecost effectivenessVSAvoidnumber of pre-recorded phrases
Core Design Contradiction:
Ease of manufactureVSQuantity of substance

Solution Approach 1:

The patent segments the large set of required audio phrases into contextual categories (greetings, farewells, responses, etc.). By organizing phrases according to conversation context rather than storing all possible phrases, the system reduces the total number of recordings needed while maintaining the ability to handle diverse conversational scenarios through contextual matching

Inventive Principle:
Principle #1Segmentation

3Adaptability or versatility

If speech synthesis technology is used to provide unlimited arbitrary phrases, then the versatility is improved, but the naturalness and convincing quality of the voice performance remains inadequate

Engineering Contradiction:
Improvephrase versatilityVSAvoidvoice performance naturalness
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The system uses pre-recorded authentic voice performances as templates instead of synthesized speech. By capturing and replaying actual human voice performances, the system maintains naturalness and convincing quality while achieving versatility through contextual organization of the recorded phrases, avoiding the robotic quality of speech synthesis

Inventive Principle:
Principle #26Copying

Data Source

PatentUS8874444B2Simulated conversation by pre-recorded audio navigator
Publication Date: 2014.10.28 DISNEY ENTERPRISES INC
  • US8874444B2 patent drawing
  • US8874444B2 patent drawing
  • US8874444B2 patent drawing

AI summary

A method is provided for a simulated conversation by a pre-recorded audio navigator, with particular application to informational and entertainment settings. A monitor may utilize a navigation interface to select pre-recorded responses in the voice of a character represented by a performer. The pre-recorded responses may then be queued and sent to a speaker proximate to the performer. By careful organization of an audio database including audio buckets and script-based navigation with shifts for tailoring to specific guest user profiles and environmental contexts, a convincing and dynamic simulated conversation may be carried out while providing the monitor with a user-friendly navigation interface. Thus, highly specialized training is not necessary and flexible scaling to large-scale deployments is readily supported.