Real-Time Audio Collection and Wake-Up Word Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice recognition systems experience interruptions and inaccuracies due to the need for manual wake-up processes, leading to incomplete data processing and poor user experience, as they require users to wake up devices before inputting audio, which can result in missed audio data and incorrect processing.

Innovation Solution

A method and apparatus for real-time audio collection and processing, where audio recognition is performed on collected audio to identify wake-up words, determining their accuracy, and initiating data processing only when the wake-up word is recognized accurately, with inaccurate wake-up words sent to a server for further identification, ensuring complete and accurate data processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual wake-up process is required before audio input, then device can be activated, but audio data may be missed and processing becomes incomplete

Engineering Contradiction:
Improvecompleteness of audio data processingVSAvoiduser operation complexity
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system performs preliminary audio collection before the user completes the wake-up process. The audio collection unit starts collecting audio data immediately, and the processing unit processes this pre-collected audio data after wake-up completion, ensuring no audio data is missed regardless of wake-up timing.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent separates the audio collection function from the wake-up activation process. The audio collection unit operates independently to collect audio data, while the processing unit handles the actual processing after wake-up, extracting the audio collection step from the wake-up sequence to eliminate data loss.

Inventive Principle:
Principle #2Taking out (Extraction)

2Productivity

If wake-up word recognition is performed, then data processing can be initiated, but recognition inaccuracies occur

Engineering Contradiction:
Improvedata processing efficiencyVSAvoidwake-up word recognition accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system implements feedback mechanisms where the processing unit continuously monitors recognition results and adjusts processing based on accuracy assessments. When recognition confidence is sufficient, processing proceeds immediately; when uncertain, the system requests re-input or uses alternative processing paths to ensure accuracy.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent employs dynamic processing strategies where the system adapts its recognition threshold and processing approach based on contextual factors, audio quality, and confidence levels. This allows the system to balance speed and accuracy by adjusting recognition strictness dynamically rather than using a fixed threshold.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS10803861B2Method and apparatus for identifying information
Publication Date: 2020.10.13 BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
  • US10803861B2 patent drawing
  • US10803861B2 patent drawing
  • US10803861B2 patent drawing

AI summary

Embodiments of the present disclosure disclose a method and apparatus for identifying information. One embodiment of the method includes: collecting to-be-processed audio in real-time; performing voice recognition on the to-be-processed audio; performing data-processing on the to-be-processed audio, when the audio is recognized as a wake-up word, the wake-up word is used for instructing performing data-processing on the to-be-processed audio. The embodiment can identify keywords from the to-be-processed audio obtained in real-time and then perform data-processing on the to-be-processed audio, which improves completeness in obtaining the to-be-processed audio and accuracy in performing data-processing on the to-be-processed audio.