Speech Processing Method for Adaptive Wake-Up Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing wake-up word models fail to accurately detect wake-up words due to variations in speech rates and speaking habits, leading to instances where the wake-up word is not spoken within the training window, resulting in device failure.

Innovation Solution

A speech processing method that obtains initial speech information exceeding a preset duration threshold, identifies similar speech segments, deletes similar frames to create secondary speech information within the threshold, and analyzes this information to determine user intent, allowing for flexible speech input without requiring users to adjust their speaking speed.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If a fixed training window length is used for wake-up word detection, then the model structure is simple and training is efficient, but users with slow speech rates or long pauses cannot complete their wake-up words within the window, causing detection failure

Engineering Contradiction:
Improvewake-up word detection efficiencyVSAvoidadaptability to different speech rates
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent applies dynamics by making the analysis window length adaptive rather than fixed. The system dynamically adjusts the analysis window length based on the actual speech duration, allowing the window to expand or contract to accommodate different speech rates and pauses while maintaining efficient detection

Inventive Principle:
Principle #15Dynamics

2Adaptability or versatility

If the analysis time window is extended to accommodate slow speakers, then adaptability to different speech rates improves, but the complexity of the analysis system increases and processing efficiency decreases

Engineering Contradiction:
Improveadaptability to different speech ratesVSAvoidspeech analysis system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies parameter changes by adjusting the analysis window length parameter based on speech characteristics. Instead of creating a complex system with multiple fixed windows, the solution changes the single parameter of window length to adapt to different speech durations, maintaining system simplicity while improving adaptability

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If the analysis time window is extended to accommodate slow speakers, then adaptability to different speech rates improves, but processing time and computational resources increase

Engineering Contradiction:
Improveadaptability to different speech ratesVSAvoidspeech processing time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system dynamically determines the analysis window length based on actual speech duration, processing only the necessary time span. This dynamic approach avoids the fixed overhead of extended windows while ensuring slow speakers have sufficient time to complete their wake-up words, optimizing both adaptability and processing efficiency

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12260853B2Speech processing method and apparatus
Publication Date: 2025.03.25 LENOVO (BEIJING) LTD
  • US12260853B2 patent drawing
  • US12260853B2 patent drawing
  • US12260853B2 patent drawing

AI summary

A speech processing method includes obtaining first speech information from a user, determining one or more similar speech segments in the first speech information and deleting one or more similar frames each of the one or more similar speech segments to obtain second speech information, and analyzing the second speech information to determine a user intent corresponding to the first speech information. A duration of the first speech information exceeds a preset analysis duration threshold, and a duration of the second speech information does not exceed the preset analysis duration threshold.