Multimedia Telephony System with Visual GUI and Text Input
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Voice over Internet Protocol (VoIP) systems are limited by their audio-only user interface, which restricts user interaction and data entry capabilities, making it difficult to efficiently navigate and select options, especially when dealing with large datasets like voicemails, due to the limitations of DTMF tones and voice recognition technology.
Innovation Solution
A multimedia interactive telephony system that generates dynamic content, allowing communication devices to render graphical user interfaces (GUIs) and accept input from devices like mice and keyboards, enabling visual presentation and interaction beyond audio, and allowing services to interact with users even when not engaged in an active call.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If audio-only interface with DTMF and voice recognition is used, then service interaction is maintained, but user interaction efficiency and data entry capability are limited
Solution Approach 1:
The patent combines audio interface elements with visual GUI elements and text input capabilities into a unified telephony system. Users can interact through multiple modalities (audio, visual, text) simultaneously, merging the strengths of each interface type to overcome the limitations of audio-only interaction while maintaining service compatibility.
Solution Approach 2:
The telephony system is designed to support multiple input and output methods universally. The same service can be accessed through audio prompts, visual displays, text messages, and standard telephony interfaces, making the system adaptable to different user preferences and capabilities without requiring separate service versions.
2Ease of operation
If DTMF keypad input is used, then service navigation is enabled, but control is limited to twelve keys making complex data entry difficult
Solution Approach 1:
The patent adds visual and text-based input dimensions to the traditional audio/DTMF interface. Users can see and read options visually, type text directly, and receive audio feedback simultaneously. This dimensional expansion allows complex data entry and navigation without increasing the physical complexity of the interface hardware.
Solution Approach 2:
The system introduces visual displays and text interfaces as intermediary layers between the user and the service. These intermediaries translate complex service options into visually presentable formats and capture user input more efficiently than DTMF alone, reducing the complexity burden on the user while maintaining service functionality.
3Ease of operation
If sequential audio presentation of list items is used, then user interaction is maintained, but time consumption increases when dealing with large datasets
Solution Approach 1:
The patent presents list items and service options in visual formats that can be perceived and processed simultaneously rather than sequentially. Users can see multiple options at once, read them quickly, and make selections without hearing each item read aloud in sequence, dramatically reducing navigation time for large datasets while maintaining interactive capability.
Solution Approach 2:
The system creates visual copies of audio information and text representations of service options. Instead of hearing each list item read aloud, users see visual representations of the same information, allowing parallel processing of multiple items simultaneously and eliminating the time-consuming sequential audio presentation.
4Adaptability or versatility
If voice recognition is used for input, then interaction capability is expanded, but accuracy is reduced for unpronounceable texts and specialized terminology
Solution Approach 1:
The patent merges voice recognition with text-based input methods and visual interfaces. Users can use voice for natural commands, type text directly for precise input, and see visual options for selection. This combination allows the system to leverage the flexibility of voice while compensating for its inaccuracies with text and visual input methods.
Solution Approach 2:
The system introduces visual displays and text interfaces as intermediary layers that bridge the gap between voice input and service processing. These intermediaries provide precise text-based input options for unpronounceable or specialized terms while maintaining voice recognition for natural interactions, with the visual/text layer correcting any voice recognition errors.
Data Source
AI summary
In a multimedia interactive telephony system, a voice service server generates dynamic content intended for consumption by a communication device. The dynamic content is sent to a gateway where it is transformed from to an intermediate content format appropriate for rendering at the communication device. The user may interact with the transformed dynamic content rendered on the communication device, causing the arguments to be sent to the voice server, thus allowing user interactivity with the voice service. The voice services server may also generate dynamic content for simultaneous consumption by multiple communication devices, each of which may independently render an intermediate content format appropriate to it. The voice services server may also generate the dynamic content for the communication device while the communication device is not currently engaged in an active call.


