Speech Terminal Trigger Word Detection for Conversation Privacy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech terminals often transmit irrelevant speech operations to drivers during phone conversations, disrupting the conversation between the driver and others on the phone.
Innovation Solution
A speech terminal system that includes an output controller and a speech recognizer, which delays sound data transmission and inhibits output based on speech recognition results, using trigger words to manage command signals and silence compression to prevent irrelevant speech from being transmitted.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If speech data is transmitted in real-time during phone conversations, then communication responsiveness is improved, but irrelevant speech commands disrupt the conversation
Solution Approach 1:
The system performs preliminary speech recognition analysis on incoming speech data before transmitting it to the other party. By detecting trigger words and command structures in advance, the system can identify speech commands that would otherwise be transmitted as normal conversation, thereby preventing irrelevant commands from disrupting the phone conversation while maintaining real-time transmission for normal speech
Solution Approach 2:
The patent introduces an intermediary speech recognition analysis layer between the microphone input and the transmission output. This intermediary component analyzes the speech data structure, identifies trigger words, and determines whether the speech constitutes a command or normal conversation before allowing transmission, thus filtering out harmful irrelevant commands while preserving useful conversation
2Measurement precision
If speech recognition is performed on all sound data, then command accuracy is improved, but processing time increases
Solution Approach 1:
Instead of performing full speech recognition on all sound data, the system applies partial action by first scanning for trigger words using a lightweight detection mechanism. Only when trigger words are detected does the system perform the more intensive full speech recognition analysis. This partial approach maintains high command recognition accuracy while significantly reducing overall processing time and computational resources required
Data Source
AI summary
A speech command generation system includes multiple speech terminals that communicate with each other via a network. Each terminal, which includes a sound pickup device and a speaker. At least one of the terminals converts local picked up sound data to text data, while delaying outputting of the sound data to a remotely communicating terminal, and determines whether the text data includes a trigger word. When the text data includes the trigger word, the outputting of the sound data to the remotely communicating terminal is inhibited.


