Call-Based Audio Sockets for Voice Server Latency Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current telecommunications voice servers face processing delays due to dynamic audio socket allocation and frequent communication of socket information between the T&M subsystem and speech engines, especially under high loads, as they use turn-based speech engine allocation rather than call-based allocation, leading to bottlenecks in componentized architectures.

Innovation Solution

Establishing call-based audio sockets within a media converting component of the voice server, where audio sockets remain available for the duration of a call, allowing continuous communication between the media converter and speech engines without the need for frequent socket re-establishment, thereby reducing latency and bottlenecks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If turn-based speech engine allocation is used, then speech engine resource utilization is maximized, but processing delays and bottlenecks increase due to frequent socket re-establishment and information conveyance

Engineering Contradiction:
Improvespeech engine resource utilizationVSAvoidprocessing delay
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-establishing audio sockets and conveying socket information to the T&M subsystem before speech processing begins. This allows the communication path to be ready in advance, eliminating the need for frequent socket re-establishment during turn-based processing, thereby reducing processing delays while maintaining high speech engine utilization

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces a socket information cache as an intermediary component that stores and manages audio socket information. This intermediary allows the T&M subsystem to quickly retrieve socket information without direct real-time communication with speech engines, reducing bottlenecks and processing delays in the componentized architecture

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If dynamic audio socket allocation is performed for each turn, then speech engine flexibility is maintained, but communication bottlenecks increase due to frequent socket information conveyance between T&M subsystem and speech engines

Engineering Contradiction:
Improvespeech engine flexibilityVSAvoidcommunication overhead
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent merges the socket management function into the T&M subsystem by having it establish and cache socket information centrally. This consolidation eliminates the need for repeated socket establishment communications between the T&M subsystem and multiple speech engines, reducing communication overhead while maintaining the flexibility to allocate different speech engines to different turns

Inventive Principle:
Principle #5Merging (Combining)

3Loss of time

If call-based speech engine allocation is used, then processing delays are reduced through stable communication paths, but speech engine resource utilization decreases due to 1-1 mapping

Engineering Contradiction:
Improveprocessing delayVSAvoidspeech engine resource utilization
Core Design Contradiction:
Loss of timeVSProductivity

Solution Approach 1:

The patent segments the allocation strategy by applying call-based socket establishment for the communication path stability while maintaining turn-based speech engine allocation for resource optimization. The audio sockets are established at the call level to provide stable communication paths, while speech engines can still be dynamically allocated to different turns, achieving both reduced processing delays and high resource utilization

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS7424432B2Establishing call-based audio sockets within a componentized voice server
Publication Date: 2008.09.09 MICROSOFT TECHNOLOGY LICENSING LLC
  • US7424432B2 patent drawing
  • US7424432B2 patent drawing
  • US7424432B2 patent drawing

AI summary

A method of interfacing a telephone application server and a speech engine can include the step of establishing one or more audio sockets in a media converting component of the telephone application server. The audio socket can remain available for approximately a duration of a call. A work unit that requires processing by a speech engine can be detected for the call. An identifier for the audio socket and a data for the work unit can be conveyed to a selected speech engine. Work unit results from the selected speech engine can be received by the media converting component via the previously established audio socket.