Event-Triggered Speech Recognition for App Actions
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing electronic devices with speech recognition functions require additional user inputs, such as wake-up utterances or button presses, to activate intelligence applications when events occur.
Innovation Solution
An electronic device equipped with a communication circuit, display, microphone, and processor that can execute an artificial intelligent application in response to designated events, sense user utterances, and perform actions based on server orders without additional user input.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If additional user input (wake-up utterance or button press) is required to activate intelligence application, then reliability of application execution is improved, but ease of operation deteriorates
Solution Approach 1:
The system performs preliminary analysis of event data before user input is required. When an event occurs, the intelligence application pre-analyzes the event data and prepares potential action orders, so that when the user provides simple input (like a wake-up utterance), the system can quickly execute the pre-prepared actions without requiring complex real-time processing
Solution Approach 2:
The system enables itself to be activated by events automatically. When a designated event occurs, the system automatically triggers the intelligence application and initiates the speech recognition process without requiring manual activation, making the system self-activating based on event conditions
2Ease of operation
If speech recognition function is activated automatically upon event occurrence, then ease of operation is improved, but device complexity increases
Solution Approach 1:
The system segments the speech recognition process into distinct phases: event detection phase, pre-analysis phase, speech sensing phase, and execution phase. Each phase is handled by specific modules (event manager, intelligence application, speech recognition module), allowing the complex functionality to be distributed and managed in manageable segments
Solution Approach 2:
The system dynamically adjusts its operation mode based on event conditions. The speech recognition function is activated dynamically when events occur rather than remaining continuously active or requiring static manual activation. The system transitions between different operational states (idle, event-active, speech-recognition mode) based on real-time conditions
3Measurement precision
If event data is transmitted to external server for processing, then measurement precision of user intent is improved, but loss of time increases
Solution Approach 1:
The system performs preliminary processing of event data locally before transmitting to the server. The intelligence application pre-analyzes the event data and extracts relevant features, so that when data is sent to the external server, the server only needs to perform final intent recognition rather than complete analysis from scratch, reducing overall processing time while maintaining precision
Solution Approach 2:
The intelligence application acts as an intermediary between the event detection system and the external speech recognition server. It preprocesses event data, formats it appropriately, and manages the communication with the external server, optimizing the data exchange process and reducing transmission and processing delays
Data Source
AI summary
A method includes receiving a designated event related to a second application while an execution screen of a first application is displayed on a display. The method also includes executing an artificial intelligent application in response to the designated event. The method further includes transmitting data related to the designated event to an external server, based on the executed artificial intelligent application. Additionally, the method includes sensing a user utterance related to the designated event for a designated period of time. The method also includes transmitting the user utterance to the external server. The method further includes receiving an action order for performing a function related to the user utterance from the external server. The method also includes executing the second application at least based on the received action order. The method further includes outputting a result of performing the function by using the second application.


