Smart Glasses Speaker Selection via Gaze Tracking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for interpreting conversations among multiple people speaking different languages are inefficient due to lack of clear speaker selection, excessive resource consumption, and accuracy degradation from crosstalk in short-range communication networks.
Innovation Solution
An apparatus and method using smart glasses with a camera, gaze-tracking camera, and interpretation target processor to capture and track interlocutors, generate a virtual conversation space, and select a target person for interpretation based on user gaze, allowing for efficient and accurate connection of speakers in a conversation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If N-to-N connection method is used for multiple people, then all languages can be interpreted into all languages, but excessive costs are required for the system
Solution Approach 1:
The patent divides the N-to-N interpretation problem into N separate 1-to-N interpretation tasks. Each participant is assigned a dedicated interpretation terminal that interprets only the speech of other participants into their native language, rather than requiring all participants to be connected to all other participants simultaneously. This segmentation reduces the total number of language pairs that need to be supported.
2Ease of operation
If intensity measurement of short-range network is used for connecting multiple users, then connections can be established, but accuracy in interlocutor connection is degraded due to crosstalk
Solution Approach 1:
The patent introduces a camera-based visual identification system as an intermediary between the short-range communication and the interpretation process. The camera captures images of participants, and the system identifies interlocutors based on visual recognition rather than relying solely on signal intensity measurements. This intermediary mechanism eliminates crosstalk issues by providing unambiguous visual identification of the intended interlocutor.
3Ease of operation
If no clarified method of changing speakers is provided, then people can speak freely, but people may speak at the same time
Solution Approach 1:
The patent implements a feedback mechanism where the system continuously monitors who is speaking and provides visual feedback to participants through the smart glasses display. When a participant finishes speaking, the system highlights the next speaker in the sequence, providing clear visual cues about whose turn it is to speak. This feedback loop helps participants understand speaking turns without requiring explicit verbal announcements.
Data Source
AI summary
Provided are an apparatus and method for selecting a speaker by using smart glasses. The apparatus includes a camera configured to capture a front angle video of a user and track guest interpretation interlocutors in the captured video, smart glasses configured to display a virtual space map image including the guest interpretation interlocutors tracked through the camera, a gaze-tracking camera configured to select a target person for interpretation by tracking a gaze of the user so that a guest interpretation interlocutor displayed in the video may be selected, and an interpretation target processor configured to provide an interpretation service in connection with the target person selected through the gaze-tracking camera.


