Voice-Tracking Electronic Housing Using DOA and Image Confirmation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing electronic devices face challenges in accurately and quickly recognizing and facing a user due to limitations in detecting sound source location and image recognition, often resulting in inefficient movement towards the user.
Innovation Solution
An electronic device equipped with cameras and microphones that use Direction of Arrival (DOA) estimation to determine the sound source direction, followed by image scanning to confirm user presence, and adjust movement accordingly, employing multiple speeds and directions to ensure accurate user detection and alignment.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If the electronic device uses only sound source detection to determine user direction, then the response speed is fast, but the accuracy of user detection is insufficient
Solution Approach 1:
The patent combines sound source detection and image recognition into a unified user detection system. The processor integrates results from both microphones (sound source detection) and camera (image recognition) to determine user direction, achieving both fast response and high accuracy by merging multiple detection modalities.
Solution Approach 2:
The processor acts as an intermediary that receives detection results from both sound source detection and image recognition systems, processes and fuses this information, and then controls the driving part to move toward the user. This intermediary processing enables accurate and efficient user localization.
2Measurement precision
If the electronic device uses only image recognition to detect user direction, then the accuracy is high, but the response time is slow
Solution Approach 1:
The system performs sound source detection first as a preliminary action to quickly identify the approximate user direction. This initial detection provides a head start, allowing the image recognition to focus on a smaller search area, thereby reducing overall response time while maintaining high accuracy.
Solution Approach 2:
The user detection process is segmented into two phases: first, sound source detection provides quick directional information; second, image recognition confirms and refines the user location. This segmentation allows each subsystem to operate optimally, with sound providing fast initial guidance and image providing accurate confirmation.
3Speed
If the electronic device moves quickly toward the detected sound source, then the response speed is fast, but the accuracy of facing the user may be reduced due to detection errors
Solution Approach 1:
The system uses feedback from image recognition to verify and correct the direction obtained from sound source detection. If the image recognition confirms the user is in the detected direction, the device moves quickly; if not, it adjusts the movement direction based on the corrected information, ensuring both speed and accuracy.
Solution Approach 2:
The movement control is dynamic, adjusting speed and direction based on the confidence level and consistency of detection results. When both sound and image detections align, the device moves quickly with high confidence; when they differ, the system dynamically adjusts by prioritizing the more reliable detection or performing additional verification.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
The device can accurately and quickly detect and face the user by combining sound source localization with image recognition, enhancing user interaction efficiency.
Implementation Method 1
detect a first direction from which the user utterance originated, based on at least part of the user utterance
Data Source
AI summary
An electronic device and method are disclosed. The device includes a housing, at least one camera, a plurality of microphones configured to detect a direction of a sound source, at least one driver operable to rotate and/or move at least part of the housing, a wireless communication circuit, a processor operatively connected to the camera, the microphones, the driver, and the wireless communication circuit, and a memory. The processor implements the method, including: receiving a user utterance, detect a first direction from which the user utterance originated, control the driver to rotate and/or move towards the first direction, a first image scan for the first direction and analyze the image for a presence of a user, when the user is not detected, rotate and/or move the at least part of the housing in a second direction, and perform a second image scan.


