Audio-Visual User Orientation for Electronic Devices in Noisy Settings
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing electronic devices face challenges in accurately and quickly detecting the location of a sound source, especially in noisy environments or when the sound signal is weak, which affects their ability to face the user effectively.
Innovation Solution
The electronic device employs a combination of microphones for Direction of Arrival (DOA) estimation and a camera for image scanning, using a processor to determine the direction of the sound source and then moving to orient the camera towards it, with adjustable speeds based on reliability and signal quality, to accurately detect and face the user.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the electronic device uses only microphone-based DOA estimation to detect sound source direction, then the device can operate in dark environments, but the detection accuracy deteriorates in noisy environments or when sound signals are weak
Solution Approach 1:
The patent combines microphone-based DOA estimation with camera-based visual detection to create a hybrid detection system. The processor integrates audio direction information with visual image scanning results to determine the final user direction, thereby improving detection accuracy in noisy environments where either single modality would fail
2Measurement precision
If the electronic device performs comprehensive image scanning to verify sound source location, then the detection accuracy improves, but the response time increases
Solution Approach 1:
The system performs preliminary DOA estimation using microphones before initiating comprehensive image scanning. This preliminary audio-based direction estimation allows the camera to focus its scanning in a predetermined direction, reducing the overall scanning time while maintaining accuracy through subsequent visual verification
Solution Approach 2:
The patent dynamically adjusts the image scanning strategy based on the reliability of DOA estimation. When sound signals are strong and clear, the scanning range is reduced. When sound signals are weak or noisy, the system expands the scanning range to ensure accurate user detection, thereby optimizing the balance between speed and accuracy
3Speed
If the electronic device moves quickly to face the detected direction, then the response speed improves, but the accuracy of facing the correct user direction may deteriorate
Solution Approach 1:
The system uses feedback from the image scanning results to verify and correct the initial DOA-based direction estimation. The processor compares the detected user position from image analysis with the predicted position from audio DOA, and adjusts the final facing direction accordingly, ensuring both speed and accuracy
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach enables the electronic device to accurately and quickly detect the user's location, improving its ability to face the user even in challenging acoustic conditions, by leveraging both sound and visual cues for precise targeting.
Implementation Method 1
a camera 150 and a plurality of microphones 140. The processor 120 receives a user utterance, and detects a first direction based on at least part of the received user utterance
Data Source
Figure 1
Figure 2
Figure 3
AI summary
An electronic device and method are disclosed. The device includes a housing, at least one camera, a plurality of microphones configured to detect a direction of a sound source, at least one driver operable to rotate and/or move at least part of the housing, a wireless communication circuit, a processor operatively connected to the camera, the microphones, the driver, and the wireless communication circuit, and a memory. The processor implements the method, including: receiving a user utterance, detect a first direction from which the user utterance originated, control the driver to rotate and/or move towards the first direction, a first image scan for the first direction and analyze the image for a presence of a user, when the user is not detected, rotate and/or move the at least part of the housing in a second direction, and perform a second image scan.