Beamforming Head Control for Mouth-Tracked Voice Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing communication systems with voice recognition robots have complex apparatus configurations and struggle to accurately recognize user voices, necessitating a simpler configuration that can effectively identify and focus on the user's mouth for accurate voice recognition.
Innovation Solution
A communication system with a head part that can be displaced, equipped with a camera for image capture and a microphone capable of beam-forming in a specific direction, which identifies the user's mouth position and adjusts its position to include it within the beam-forming region, thereby omitting the need for a holding unit and ensuring accurate voice recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a holding unit is used to hold the microphone, then the microphone can be positioned to approach the user's mouth, but the apparatus configuration becomes complicated
Solution Approach 1:
The patent integrates the microphone and camera into a single head part assembly, eliminating the need for a separate holding unit. The microphone and camera are positioned relative to each other within the head part, and both are controlled to face the user's mouth simultaneously through coordinated positioning of the head part, simplifying the overall apparatus configuration while maintaining voice recognition accuracy.
Solution Approach 2:
The head part serves multiple functions: it houses both the camera for capturing user images and the microphone for voice recognition, and acts as a unified positioning unit that can orient both components toward the user's mouth. This multi-functional design eliminates the need for separate holding mechanisms and reduces structural complexity.
2Measurement precision
If the microphone is held by a holding unit, then it can be positioned close to the user's mouth for accurate recognition, but the system requires additional components that complicate the configuration
Solution Approach 1:
The camera and microphone are merged into a single head part assembly with a fixed spatial relationship between them. The control unit coordinates both the camera positioning for mouth detection and the microphone positioning for voice capture, allowing accurate mouth position identification and voice recognition without requiring a separate holding unit for the microphone.
3Measurement precision
If the head part is positioned to include the user's mouth in the beam-forming region, then voice recognition accuracy improves, but the system must accurately identify mouth position from images
Solution Approach 1:
The camera acts as an intermediary device that captures images of the user's face, and the control unit processes these images to identify the mouth position. This identified mouth position information is then used to position the microphone's beam-forming region accurately toward the user's mouth, enabling precise voice recognition through a two-step process of visual detection followed by acoustic targeting.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This configuration simplifies the apparatus while ensuring accurate voice recognition and provides an impression of attentive listening by maintaining the line of sight focused on the user's face, enhancing user interaction.
Implementation Method 1
a microphone provided in the head part and configured to be able to form a beam-forming in a specific direction
Data Source
AI summary
A communication system according to the present disclosure includes a camera configured to be able to photograph a user who is a communication partner and a microphone configured to be able to form a beam-forming in a specific direction. The control unit identifies a position of the mouth of a user using an image of the user taken by the camera and controls a position of a head part so that the identified position of the mouth of the user is included in a region of the beam-forming.


