Voice Recognition End-of-Speech Detection via Dynamic Silence Duration

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice recognition systems struggle to accurately determine the end of speech input without pre-setting a category for search items, leading to inappropriate determination of input completion, especially when user features such as voice characteristics are not considered.

Innovation Solution

A voice recognition device and method that extracts features from input voice data to set a duration for determining the end of speech, allowing for flexible determination based on categories like addresses, facility names, telephone numbers, recognition errors, user age, and speech speed, ensuring accurate input completion detection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If a fixed duration is set for determining end of speech based on pre-set categories, then the determination process is simple, but the flexibility and accuracy deteriorate when user features or voice characteristics are not considered

Engineering Contradiction:
Improvedetermination process complexityVSAvoidflexibility to end-of-speech determination
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent makes the duration setting dynamic by adjusting it based on extracted voice features (pitch, volume, timbre) and user characteristics (age, speech habits) rather than using a fixed predetermined value. The processor dynamically determines appropriate duration based on real-time voice analysis, allowing the system to adapt to different users and speaking conditions while maintaining operational simplicity through automated feature extraction.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the parameter of duration from a fixed predetermined value to a variable determined by multiple factors including voice features (pitch, volume, timbre), user age, and speech habits. This parameter transformation allows the system to optimize end-of-speech determination accuracy for different users and contexts without increasing operational complexity, as the changes are automatically calculated based on extracted features.

Inventive Principle:
Principle #35Parameter changes

2Ease of operation

If voice recognition determines end of speech without considering user features, then the system operation is simple, but the accuracy of input completion detection deteriorates

Engineering Contradiction:
Improvesystem operation simplicityVSAvoidaccuracy of input completion detection
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The system performs self-service by automatically extracting voice features (pitch, volume, timbre) and analyzing user characteristics (age, speech habits) to determine optimal duration settings without requiring manual configuration. The processor autonomously adjusts the end-of-speech determination parameters based on real-time voice analysis, maintaining ease of operation while significantly improving detection accuracy through adaptive feature-based determination.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent implements feedback mechanisms where the system continuously monitors voice features and user responses to refine duration settings. By analyzing pitch variations, volume changes, and speech patterns, the system receives feedback on actual speech characteristics and adjusts the end-of-speech determination threshold accordingly, improving accuracy while maintaining simple operation through automated closed-loop control.

Inventive Principle:
Principle #23Feedback

3Productivity

If a predetermined category is set for search items, then the voice recognition processing is efficient, but the adaptability to different user needs and speech patterns deteriorates

Engineering Contradiction:
Improvevoice recognition processing efficiencyVSAvoidadaptability to different user needs
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal voice recognition system that handles multiple user needs and speech patterns through a single adaptive framework. By extracting general voice features (pitch, volume, timbre) and analyzing various user characteristics (age, speech habits, input speed), the system provides multi-functional adaptation across different users and contexts without requiring separate processing paths, maintaining efficiency while enhancing versatility.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system dynamically adapts to different user needs by adjusting duration settings based on real-time analysis of voice features and user characteristics rather than relying on fixed category-based processing. This dynamic approach allows the system to efficiently handle diverse search items and user preferences through automated parameter adjustment, maintaining processing efficiency while significantly improving adaptability to individual user patterns.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11195535B2Voice recognition device, voice recognition method, and voice recognition program
Publication Date: 2021.12.07 TOYOTA JIDOSHA KK
  • US11195535B2 patent drawing
  • US11195535B2 patent drawing
  • US11195535B2 patent drawing

AI summary

A voice recognition device includes a memory and a processor including hardware. The processor is configured to extract a feature of input voice data and set a duration of a silent state after transition of the voice data to the silent state. The duration is used for determining that an input of the voice data is completed.