Voice Assistant Control Across Application Interfaces
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech assistant technologies require frequent waking up and cannot maintain continuous control over devices when switching between applications, limiting user interaction and convenience.
Innovation Solution
A method and apparatus for speech assistant control that displays a target interface and continuously receives speech data, determining if a control instruction is included in the data, allowing for seamless interaction across different application interfaces without repeated waking up, using speech recognition and user detection techniques.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If the speech assistant exits when switching to other applications, then the system resource occupation is reduced, but the user cannot realize control through the speech assistant when another application is activated
Solution Approach 1:
The speech assistant is designed to provide universal control capabilities across multiple applications. When the target interface belongs to a different application, the speech assistant remains active and continues to receive and process speech data, enabling users to issue control instructions for the running application without exiting the speech assistant interface.
Solution Approach 2:
The speech assistant maintains continuous operation across application switches. The system determines whether the target interface belongs to a different application, and if so, keeps the speech assistant in an active state to continuously receive speech data, ensuring uninterrupted control functionality rather than exiting and requiring re-waking.
2Speed
If the speech assistant wakes up frequently, then the control responsiveness is improved, but the user experience deteriorates due to repeated wake-up operations
Solution Approach 1:
The speech assistant maintains continuous speech data reception across application switches without requiring repeated wake-up operations. The system preserves the active state of the speech assistant when transitioning between applications, allowing users to continuously issue control instructions without interrupting the interaction flow.
Solution Approach 2:
The system proactively determines whether the target interface belongs to a different application before the speech assistant would normally exit. By anticipating the need for continued control capability, the system prepares to maintain the active state, preventing the need for repeated wake-up operations and improving user convenience.
3Ease of operation
If the speech assistant remains active across application switches, then the user experience is improved, but the energy consumption increases
Solution Approach 1:
The speech assistant's operational state is dynamically adjusted based on the application context. The system determines whether the target interface belongs to a different application and adaptively decides whether to maintain the active state or exit, optimizing energy consumption while preserving interaction continuity when needed.
Solution Approach 2:
The system changes the operational parameters of the speech assistant based on application switching events. When a different application is detected, the system modifies the speech assistant's state from exit-to-wake-up cycle to continuous operation mode, and vice versa, thereby optimizing energy usage according to actual user needs.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Provided are a method and apparatus for speech assistant control and a computer-readable storage medium. The method includes: after a speech assistant is waken up, displaying a target interface corresponding to a control instruction corresponding to received speech data; when the target interface is different from an interface of the speech assistant, displaying a speech reception identifier in the target interface and controlling to continuously receive speech data; determining, based on second speech data received when the target interface is displayed, whether a target control instruction to be executed is included in the second speech data; and displaying an interface corresponding to the target control instruction when the target control instruction is included in the second speech data.