Display Voice Control with UI Text Fallback for Third-Party Apps
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Smart voice systems struggle to recognize and execute voice commands for third-party applications and menu options due to lack of semantic configuration, leading to incomplete or incorrect actions in complex scenarios.
Innovation Solution
A display apparatus equipped with a detector to acquire voice information, a controller to extract keywords, and a method to traverse a configuration library for matching actions, or recognize text and layout information to generate control instructions, enabling execution of user commands even for unconfigured applications.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the smart voice system uses a configuration library with pre-defined semantic functions, then recognition accuracy for built-in applications is improved, but adaptability to third-party applications and menu options deteriorates
Solution Approach 1:
The patent introduces an intermediary component that captures text information and layout information from the display interface. This intermediary layer bridges the gap between the voice input system and the application control system, enabling the voice system to adapt to third-party applications without requiring pre-configuration of semantic functions for each application.
Solution Approach 2:
The system performs preliminary actions by pre-obtaining text information and layout information of the display interface before voice recognition processing. This allows the system to prepare the necessary contextual information in advance, enabling faster and more accurate matching of voice commands to interface elements without requiring real-time configuration.
2Measurement precision
If the system configures semantic functions for all possible applications and menu options, then voice command recognition accuracy is improved, but device complexity and configuration workload increase
Solution Approach 1:
The system implements self-service by automatically obtaining text information and layout information from the display interface without requiring manual configuration. The voice recognition system autonomously adapts to the current interface state by retrieving necessary information dynamically, eliminating the need for developers to pre-configure semantic functions for each application or menu option.
Solution Approach 2:
The system transitions from a static configuration approach to a dynamic one where text information and layout information are obtained in real-time based on the current display state. This dynamic adaptation allows the voice recognition system to adjust to different applications and interface configurations without requiring complex pre-setup procedures.
3Adaptability or versatility
If the system retrieves text information and layout information from the display interface, then adaptability to complex interfaces is improved, but processing time and computational resources increase
Solution Approach 1:
The system performs preliminary retrieval of text information and layout information from the display interface before voice command processing. By obtaining this contextual information in advance, the system reduces the processing time required during actual voice command execution, as the necessary interface data is already available for rapid matching and recognition.
Data Source
AI summary
Some embodiments of the present application disclose a display apparatus and a voice control method for the display apparatus. The display apparatus comprises a display, a detector and a controller. The display is configured to present a user interface, and the detector is configured to acquire user voice information; and the controller is configured to cause the display apparatus to perform: acquiring voice information inputted from a user; in response to the voice information, extracting at least one keyword from the voice information; traversing action items in a configuration library; in response to determining that no action item in the configuration library matches the at least one keyword, obtaining text information of the user interface on the display to in order to determine an control instruction according to the text information.


