Voice Assistant Interrupt Handling via Visual Audio Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing interactive systems that use artificial intelligence for voice recognition and response are limited in that they cannot process additional user queries or conversations on different topics until a response to the initial query is provided, leading to a lack of instant and active response to user needs during conversations.
Innovation Solution
An electronic apparatus equipped with a camera and microphone that analyzes images to determine if a user has input additional voice while responding to the initial voice, allowing it to stop responding to the initial voice and provide a response to the additional voice, enabling more instantaneous and active interaction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of time
If the electronic apparatus provides a response to the initial user voice before detecting additional voice, then the response can be provided completely, but the user must wait until the response is completed to make additional queries
Solution Approach 1:
The electronic apparatus performs preliminary detection of additional user voice during the response provision process using image analysis of the user's mouth movement, before the initial response is fully completed. This allows the system to prepare for and switch to processing the additional query in advance, reducing the waiting time for users to make additional queries.
2Measurement precision
If the electronic apparatus uses only audio signal processing for voice recognition, then the system remains simple, but it cannot distinguish between additional user voice and other sounds during response provision
Solution Approach 1:
The electronic apparatus merges audio signal processing with visual image processing to detect user voice. The processor analyzes both the audio signal from the microphone and the image signal from the camera, combining these two data streams to accurately distinguish additional user voice from other sounds during response provision, thereby improving detection accuracy without significantly increasing system complexity.
Solution Approach 2:
The electronic apparatus introduces image analysis as an intermediary method to verify and confirm the presence of additional user voice. By analyzing the user's mouth movement in the image signal, the system can more accurately determine whether an additional query is being made, serving as a reliable mediator between the audio input and the response processing system.
Data Source
AI summary
An electronic apparatus and a controlling method thereof are provided. The electronic apparatus includes a microphone, a camera, a memory configured to store at least one command, and at least one processor configured to, based on a first user voice being input from a user, provide a response to the first user voice, based on an audio signal including a voice being input while the response to the first user voice is provided, analyze an image captured by the camera and determine whether there is a second user voice uttered by the user in the audio signal, and based on determining that there is the second user voice uttered by the user in the audio signal, stop providing the response to the first user voice and obtain and provide a response to the second user voice.


