Terminal Voice Recognition Using Local Command Word Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice recognition technologies require users to upload personal information to cloud servers, compromising security and increasing network traffic, leading to delayed recognition experiences due to network congestion.
Innovation Solution
Implementing a terminal-based voice recognition method that splits command words using a two-command-word-slot or multi-command-word-slot recognition grammar, allowing for more voice input content recognition with the same number of command words, reducing network reliance and improving user experience.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If cloud server is used for voice recognition, then recognition capability is improved, but network traffic consumption increases and recognition delay increases
Solution Approach 1:
The patent segments the voice recognition system into terminal-based processing components that can operate independently of cloud servers. The terminal device performs local voice feature extraction, candidate text generation, and recognition operations, dividing the previously centralized cloud-based function into distributed terminal capabilities, thereby reducing network dependency and recognition delay
Solution Approach 2:
The terminal device is equipped with self-service voice recognition capabilities through local processing of voice information. The terminal extracts voice features, generates candidate texts, and performs recognition operations autonomously without requiring cloud server intervention, enabling the system to serve itself and eliminating network-related delays
2Measurement precision
If cloud server is used for voice recognition, then recognition capability is improved, but user information security deteriorates
Solution Approach 1:
The patent extracts the voice recognition processing function from the cloud server environment and relocates it to the terminal device. By taking out the recognition operations from the networked cloud environment, the system eliminates the security vulnerability of transmitting and storing user voice information on external servers, keeping all processing local to the user's own device
Solution Approach 2:
The terminal device acts as an intermediary between the user and the voice recognition process, handling all voice information processing locally without requiring cloud server mediation. This intermediary role at the terminal level protects user information by preventing its transmission to external servers while still enabling sophisticated recognition capabilities
3Adaptability or versatility
If more command words are used, then voice input content recognition is improved, but device complexity increases
Solution Approach 1:
The patent introduces a hierarchical dimension to the recognition grammar structure with multiple levels (first level, second level, etc.). Instead of expanding the number of command words linearly, the system organizes them into hierarchical levels where higher-level command words can encompass multiple lower-level options, achieving greater recognition versatility without proportionally increasing overall system complexity
Solution Approach 2:
The patent creates universal command word structures that can serve multiple functions across different contexts. A single command word at a higher hierarchical level can represent multiple specific meanings or actions, allowing the system to recognize diverse voice input content with a relatively small set of multi-functional command words, thereby reducing the need for numerous specialized command words
Data Source
AI summary
An information recognition method and apparatus are provided. The method includes receiving, by a terminal, voice information, extracting a voice feature from the voice information, performing matching calculation on the voice feature and a phoneme string corresponding to each candidate text in multiple candidate texts to obtain a recognition result, where the recognition result includes at least one command word and a label corresponding to the at least one command word, and recognizing, according to the label corresponding to the at least one command word, an operation instruction corresponding to the voice information. A terminal recognizes text information, which is corresponding to voice information input by a user, as an operation instruction.


