Smart Speaker Wake-Up Selection via Timestamp Comparison

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

When multiple smart speakers coexist, they often respond simultaneously to a wake-up word, leading to chaotic speech interactions and a poor user experience due to multiple devices entering a listening state simultaneously.

Innovation Solution

A method where smart speakers in a wireless network receive and recognize a wake-up word, record timestamps or speech intensities, and compare these values to determine which speaker to wake up first, ensuring only the closest speaker to the user enters a listening state, thereby avoiding simultaneous wake-ups.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If multiple smart speakers respond simultaneously to a wake-up word, then all speakers can be activated, but this causes chaotic speech interaction and noisy environment

Engineering Contradiction:
Improvespeech interaction qualityVSAvoidnoise and chaos
Core Design Contradiction:
ReliabilityVSObject-generated harmful factors

Solution Approach 1:

The patent applies local quality by making each smart speaker evaluate its own local conditions (speech intensity, timestamp) to determine whether to wake up. Instead of all speakers responding uniformly, each speaker independently assesses its proximity to the user based on local acoustic measurements and makes a decentralized decision, thereby eliminating chaotic simultaneous responses while maintaining reliable speech interaction.

Inventive Principle:
Principle #3Local quality

2Adaptability or versatility

If all smart speakers enter listening state simultaneously, then user coverage is maximized, but speech interaction efficiency deteriorates

Engineering Contradiction:
Improveuser coverageVSAvoidspeech interaction efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent implements preliminary action by having smart speakers pre-evaluate speech intensity and record timestamps before actually entering the listening state. This preliminary assessment allows speakers to predict which one should respond based on proximity metrics, enabling efficient single-speaker selection while still maintaining the capability for multiple speakers to be activated if needed for comprehensive user coverage.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If smart speakers use speech intensity comparison to select wake-up candidate, then the closest speaker is selected, but this requires additional recognition processing

Engineering Contradiction:
Improvespeaker proximity detectionVSAvoidrecognition processing
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies self-service by having each smart speaker independently perform speech intensity recognition and timestamp recording on its own, without requiring complex centralized coordination or additional external processing systems. Each speaker uses its built-in microphone and processor to measure local speech intensity, compare it with received intensity data from other speakers, and autonomously determine whether to wake up, thereby achieving precise proximity detection without excessive device complexity.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11437036B2Smart speaker wake-up method and device, smart speaker and storage medium
Publication Date: 2022.09.06 BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
  • US11437036B2 patent drawing
  • US11437036B2 patent drawing
  • US11437036B2 patent drawing

AI summary

The present disclosure discloses a smart speaker wake-up method, a smart speaker wake-up device, a smart speaker and a storage medium, relates to the technical field of speech recognition. The method of the present disclosure is applied to a wireless network including two or more smart speakers, and a specific implementation thereof is: receiving, speech information including a wake-up word; performing a recognition processing to the speech information to obtain identification information corresponding to the wake-up word; and waking up one smart speaker in the wireless network to enter listening state according to the identification information. The present disclosure may be applied to a scenario where multiple smart speakers coexist, so as to quickly select one smart speaker that is most likely to be wakened, avoiding a chaotic speech interaction caused by multiple smart speakers being wakened simultaneously, improving efficiency and quality of speech interaction and achieving better user experience.