Voice Control Hub for Legacy Enterprise Applications
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing consumer voice-based systems, such as Google Assistant and Amazon Alexa, are not suitable for enterprise environments due to lack of public APIs, security risks, inflexibility, and inability to integrate with legacy applications, leading to limitations in custom voice command capabilities and increased latency.
Innovation Solution
A computing system that uses a speech-to-text API to analyze user utterances and generate intents and entities, allowing for voice control of legacy applications within an enterprise environment, with a voice control hub that integrates cloud APIs and machine learning models for natural language understanding and customizable voice commands.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If consumer voice systems (Google Assistant, Amazon Alexa) are used for enterprise control, then voice-based operation is enabled, but security risks and lack of integration with legacy systems occur
Solution Approach 1:
The patent introduces a voice control hub as an intermediary component that sits between the consumer voice system and the enterprise legacy application. This hub receives voice commands from the consumer system, processes them through a command interpreter, and translates actions into legacy application-specific commands via a custom API. This intermediary layer resolves the contradiction by enabling voice operation while maintaining security and integration reliability through controlled, authorized communication channels.
Solution Approach 2:
The system is segmented into distinct functional modules: a voice recognition module for capturing speech, a command interpreter for processing intent, and a legacy application interface for executing actions. This segmentation allows each component to be optimized independently, improving overall system reliability while maintaining ease of voice-based operation.
2Ease of operation
If consumer voice systems with template-based responses are used, then basic voice control is achieved, but customization for specific business use cases is lost
Solution Approach 1:
The command interpreter is designed as a dynamic system that can adapt to different business contexts and legacy applications. Rather than using fixed templates, the system processes natural language commands and dynamically translates them into appropriate actions based on the current application state and user intent. This enables full customization for specific business use cases while maintaining ease of voice control.
Solution Approach 2:
The voice control hub provides universal functionality that can interface with multiple different legacy applications through a standardized command interpreter and API framework. This multi-functional design allows the same voice control infrastructure to serve various business purposes and applications, enhancing adaptability while preserving ease of operation.
3Device complexity
If all access occurs in a captive cloud environment, then centralized control is achieved, but latency increases and on-premises processing is lost
Solution Approach 1:
The system segments processing functions between cloud-based and on-premises components. The voice recognition and command interpretation can occur in the cloud, while the actual application actions and data processing happen locally on the enterprise premises through the legacy application interface. This segmentation reduces latency by keeping critical processing close to the data while maintaining the benefits of centralized voice control infrastructure.
Data Source
AI summary
A computing system for enabling a user to control a legacy application of an enterprise using voice commands includes a processor and a memory storing instructions that, when executed by the one or more processors, cause the computing system to receive a user utterance; generate an output by analyzing the utterance using a speech-to-text application programming interface; and perform an action with respect to an element of the legacy application. A computer-implemented method includes receiving a user utterance; generating an output by analyzing the utterance using a speech-to-text application programming interface; and performing an action with respect to an element of the legacy application. A non-transitory computer readable medium includes program instructions that when executed, cause a computer to receive a user utterance; generate an output by analyzing the utterance using a speech-to-text application programming interface; and perform an action with respect to an element of the legacy application.


