Speech Terminal Trigger Word Detection for Conversation Privacy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech terminals often transmit irrelevant speech operations to drivers during phone conversations, disrupting the conversation between the driver and others on the phone.

Innovation Solution

A speech terminal system that includes an output controller and a speech recognizer, which delays sound data transmission and inhibits output based on speech recognition results, using trigger words to manage command signals and silence compression to prevent irrelevant speech from being transmitted.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If speech data is transmitted in real-time during phone conversations, then communication responsiveness is improved, but irrelevant speech commands disrupt the conversation

Engineering Contradiction:
Improvespeech transmission speedVSAvoidirrelevant speech interference
Core Design Contradiction:
SpeedVSObject-generated harmful factors

Solution Approach 1:

The system performs preliminary speech recognition analysis on incoming speech data before transmitting it to the other party. By detecting trigger words and command structures in advance, the system can identify speech commands that would otherwise be transmitted as normal conversation, thereby preventing irrelevant commands from disrupting the phone conversation while maintaining real-time transmission for normal speech

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary speech recognition analysis layer between the microphone input and the transmission output. This intermediary component analyzes the speech data structure, identifies trigger words, and determines whether the speech constitutes a command or normal conversation before allowing transmission, thus filtering out harmful irrelevant commands while preserving useful conversation

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If speech recognition is performed on all sound data, then command accuracy is improved, but processing time increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidprocessing delay
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

Instead of performing full speech recognition on all sound data, the system applies partial action by first scanning for trigger words using a lightweight detection mechanism. Only when trigger words are detected does the system perform the more intensive full speech recognition analysis. This partial approach maintains high command recognition accuracy while significantly reducing overall processing time and computational resources required

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11302318B2Speech terminal, speech command generation system, and control method for a speech command generation system
Publication Date: 2022.04.12 YAMAHA CORP
  • US11302318B2 patent drawing
  • US11302318B2 patent drawing
  • US11302318B2 patent drawing

AI summary

A speech command generation system includes multiple speech terminals that communicate with each other via a network. Each terminal, which includes a sound pickup device and a speaker. At least one of the terminals converts local picked up sound data to text data, while delaying outputting of the sound data to a remotely communicating terminal, and determines whether the text data includes a trigger word. When the text data includes the trigger word, the outputting of the sound data to the remotely communicating terminal is inhibited.