Voice Assistant Wake-Up Keyword Recognition Using Segmented Intelligence
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current user terminals face difficulties in recognizing the meaning of user-set wake-up keywords and providing personalized services due to hardware limitations, despite advancements in speech recognition technologies.
Innovation Solution
An integrated intelligence system is developed, incorporating a microphone, speaker, processor, and memory to store NLU and response models, allowing for the selection and processing of voice inputs to generate personalized voice responses based on user-set wake-up keywords.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If speech recognition service is executed with wake-up keyword recognition, then the system can activate voice processing functions, but the user terminal cannot recognize the meaning of the wake-up keyword due to hardware limitations
Solution Approach 1:
The patent divides the intelligence system into two segments: a simple wake-up keyword recognition module running on the user terminal with limited hardware, and a complex natural language understanding module running on a server. This segmentation allows the terminal to perform basic keyword detection without requiring powerful hardware, while the server handles the computationally intensive meaning recognition and personalized service generation.
Solution Approach 2:
The patent introduces a server as an intermediary between the user terminal and the personalized service system. The server receives the wake-up keyword from the terminal, processes it using advanced NLU models, and generates personalized responses. This intermediary enables the terminal to leverage server-side computing power without requiring local hardware upgrades.
2Adaptability or versatility
If multiple response models are stored and selected based on wake-up keywords, then personalized voice responses can be provided, but the system complexity increases
Solution Approach 1:
The patent segments the response model storage and selection functionality between the terminal and server. The terminal stores only simple wake-up keyword models for initial detection, while the server stores multiple complex response models for different personalization scenarios. This division allows personalized responses without requiring the terminal to handle multiple complex models simultaneously.
Solution Approach 2:
The patent changes the parameter of model storage location from local terminal memory to remote server memory. This parameter change enables the system to access multiple response models without increasing terminal hardware complexity, as the models are stored and managed on the server side and transmitted only when needed.
3Loss of information
If the system processes natural language to grasp user intent, then personalized services can be provided, but hardware limitations prevent the terminal from recognizing the meaning of wake-up keywords
Solution Approach 1:
The patent segments the natural language processing task into two parts: simple wake-up keyword detection performed locally on the energy-constrained terminal, and complex intent recognition and meaning analysis performed on the server with充足的 computing resources. This segmentation ensures that energy-intensive NLP operations are not performed on the terminal, preserving battery life while still achieving accurate user intent recognition.
Solution Approach 2:
The server acts as an intermediary that receives raw wake-up keywords from the terminal and performs the energy-intensive natural language understanding. This intermediary approach allows the terminal to minimize energy consumption by offloading complex processing tasks to the server, while still benefiting from accurate intent recognition and personalized service generation.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A method for processing a voice input and a system therefor are provided. The system includes a microphone, a speaker, a processor, and a memory. In a first operation, the processor receives a first voice input including a first wake-up keyword, selects a first response model based on the first voice input, receives a second voice input after the first voice input, processes the second voice input, using an NLU module, and generates a first response based on the processed second voice input. In a second operation, the processor receives a third voice input including a second wake-up keyword different from the first wake-up keyword, selects a second response model based on the third voice input, receives a fourth voice input after the third voice input, processes the fourth voice input, using the NLU module, and generates a second response based on the processed fourth voice input.