Zero UI Automatic Interpretation Server Using Speech Energy Comparison
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional automatic interpretation technologies require users to perform unnecessary operations such as touching screens or clicking buttons, and can malfunction when determining who is speaking, disrupting natural conversation and causing incorrect language interpretation.
Innovation Solution
A zero user interface (UI)-based automatic interpretation system and method using a server that communicates with terminal devices equipped with microphones, speakers, and communication functions, which determines the main speech signal by comparing speech energies and performs noise cancellation to provide seamless and accurate interpretation without user intervention.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If a user interface (screen, buttons) is provided for automatic interpretation, then the device can display and control interpretation functions, but the user must perform unnecessary operations (touching screen, clicking buttons) that interrupt natural conversation
Solution Approach 1:
The patent removes the user interface (screen, buttons) from the automatic interpretation device, extracting only the essential speech processing functions. The device now operates purely through speech input and output without requiring visual displays or manual controls, thereby eliminating unnecessary operations while maintaining core interpretation functionality.
Solution Approach 2:
The automatic interpretation device autonomously detects speech signals, determines active speakers, and performs interpretation without requiring user initiation or control. The system self-activates when speech is detected and automatically manages the interpretation process, eliminating the need for user operations entirely.
2Reliability
If the device requires repeated operations to bring the terminal to the mouth, then speech input can be captured, but the conversation flow is interrupted and natural interaction is hindered
Solution Approach 1:
The patent replaces the mechanical requirement of bringing the device to the mouth with an acoustic field-based solution. The device captures speech signals through its microphone array in the surrounding acoustic field, eliminating the need for physical proximity operations while maintaining speech input accuracy through directional audio processing.
3Measurement precision
If the device continuously monitors speech signals from multiple users, then it can identify the active speaker, but it may malfunction and incorrectly interpret the wrong speech signal
Solution Approach 1:
The system continuously monitors speech signals from multiple users and provides feedback by comparing speech energies in real-time. When a user's speech energy exceeds a threshold or surpasses others, the system identifies them as the active speaker and directs interpretation resources accordingly, ensuring accurate speaker identification while maintaining system reliability through continuous verification.
Data Source
AI summary
Provided is a zero user interface (UI)-based automatic interpretation method including receiving a plurality of speech signals uttered by a plurality of users from a plurality of terminal devices, acquiring a plurality of speech energies from the plurality of received speech signals, determining main speech signal uttered in a current utterance turn among the plurality of speech signals by comparing the plurality of acquired speech energies, and transmitting an automatic interpretation result acquired by performing automatic interpretation on the determined main speech signal to the plurality of terminal devices.


