Voice-Controlled Wake-Up Using Local Self-Defined Word Features

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice-controlled devices rely on remote systems with powerful computing capabilities for speaker identification, which are vulnerable to network interruptions and require large amounts of speech data for training, making them susceptible to unauthorized activation.

Innovation Solution

A method and processing circuit that utilizes a feature-list-based database to detect self-defined words within the device, allowing for local activation and deactivation without network reliance, using feature collection and screening operations to determine valid speakers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a remote system with powerful computing capabilities is used for speaker identification, then identification accuracy is improved, but network dependency increases and reliability deteriorates

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoidsystem reliability
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent extracts the essential speaker identification functionality from the remote cloud system and implements it locally in the voice-controlled device. By extracting only the necessary feature extraction and comparison algorithms rather than the entire AI processing system, the device achieves reliable local operation while maintaining identification accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent transitions from a centralized cloud-based identification system to a distributed edge-based system. This dimensional shift from cloud-only processing to hybrid cloud-edge architecture enables the device to operate independently of network connectivity while maintaining identification capabilities through local feature databases.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If a remote system with powerful computing capabilities is used for speaker identification, then identification accuracy is improved, but device complexity increases

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the speaker identification system into distinct functional modules: local feature extraction module, feature database storage module, comparison module, and cloud synchronization module. This segmentation allows each component to be optimized independently and simplifies the overall system architecture by distributing functionality across multiple manageable units.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a local feature database as an intermediary between the microphone input and the identification logic. This intermediary layer pre-stores speaker features and handles the actual comparison, acting as a buffer that simplifies the interaction between input audio and identification algorithms, thereby reducing overall system complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If network connection is required for speaker identification, then identification accuracy is improved, but loss of time increases due to network interruptions

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoidactivation response time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary actions by pre-extracting and storing speaker voice features in the local database during device setup or initial usage. This preliminary preparation of identification data enables immediate local comparison without requiring real-time network access, eliminating delays caused by network interruptions while maintaining accurate identification.

Inventive Principle:
Principle #10Preliminary action

4Measurement precision

If cloud-based speaker identification is used, then identification accuracy is improved, but security deteriorates due to unauthorized activation

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoidunauthorized activation
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent applies preliminary anti-action by implementing local feature comparison that actively prevents unauthorized activation before it can occur. The system pre-establishes secure local authentication based on voice features, creating a first line of defense that blocks unauthorized access attempts without requiring cloud verification, thereby preemptively countering the security threat.

Inventive Principle:
Principle #9Preliminary anti-action

Data Source

PatentUS20250218441A1Method and processing circuit for performing wake-up control on voice-controlled device with aid of detecting voice feature of self-defined word
Publication Date: 2025.07.03 REALTEK SEMICON CORP
  • US20250218441A1 patent drawing
  • US20250218441A1 patent drawing
  • US20250218441A1 patent drawing

AI summary

A method for performing wake-up control on a voice-controlled device with aid of detecting voice feature of self-defined word and an associated processing circuit are provided. The method may include: performing feature collection on audio data of at least one audio clip to generate at least one feature list of the at least one audio clip, in order to establish a feature-list-based database in the voice-controlled device; performing the feature collection on audio data of another audio clip to generate another feature list of the other audio clip; and performing at least one screening operation on at least one feature in the other feature list according to the feature-list-based database to determine whether the other audio clip is invalid, in order to selectively ignore the other audio clip or execute at least one subsequent operation, where the at least one subsequent operation includes waking up the voice-controlled device.