Voice Tag Matching Engine for Low-Load Command Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice recognition systems face challenges in reducing computational load, accommodating grammatical structure variations, and verifying user authority, leading to inaccuracies and barriers in language comprehension and usage rights verification.
Innovation Solution
A barrier-free intelligent voice system that performs phonetic and morphological analysis on voice audio to identify independent semantic units, compares these units against user-defined voice tags, and executes corresponding commands, while incorporating an authority verification mechanism to ensure secure and authorized device control.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If grammatical analysis and semantic interpretation are performed on continuous voice commands, then speech understanding capability is improved, but computational load increases significantly
Solution Approach 1:
The patent segments continuous voice commands into independent semantic units through phonetic and morphological analysis. Instead of analyzing the entire sentence structure grammatically, the system divides the voice input into discrete units (words or phrases) that can be independently matched against predefined voice tags. This segmentation approach maintains speech understanding capability while significantly reducing computational complexity by avoiding full grammatical parsing and semantic interpretation of continuous commands.
2Measurement precision
If strict grammatical rules are enforced for voice command recognition, then recognition accuracy is improved, but ease of operation deteriorates due to user limitations
Solution Approach 1:
The patent inverts the traditional approach by not requiring users to follow strict grammatical rules. Instead of analyzing whether user speech conforms to grammatical structures, the system pre-defines multiple voice tags (including terms, names, titles, codes, commands, programs, and messages) and directly matches recognized semantic units against these tags. This inversion maintains recognition accuracy by comparing against known valid inputs while greatly improving ease of operation, as users can speak naturally without worrying about grammatical correctness.
3Device complexity
If voiceprint recognition technology is not used, then system complexity is reduced, but reliability deteriorates due to inability to verify user authority
Solution Approach 1:
The patent integrates multiple functions into the voice recognition system, including not only command recognition but also user authority verification. The system can identify and match voice patterns against stored voiceprints of authorized users, thereby verifying user authority without requiring a separate dedicated voiceprint recognition module. This multi-functionality approach maintains reliability by ensuring only authorized users can execute certain commands while avoiding the complexity of completely separate verification systems.
Data Source
AI summary
A barrier-free intelligent voice system and a method for controlling thereof, wherein multiple words are recognized from a voice audio to create multiple independent semantic units. Meanwhile, the system can continuously determine whether they are one of multiple voice tags created by the user. Thereafter, a target object, a program command, and a remark corresponding to the voice tag can be determined based on the successfully compared voice tag combination. Accordingly, a corresponding program can be started or a remote device can be triggered to operate. The present disclosure can be regarded as an AI intelligent voice processing engine. By allowing users to define different types of voice tag combinations, it can eliminate the grammatical and semantic analysis of natural language processing, eliminate speech translation differences and errors between different languages, effectively reduce the amount of calculations, increase the processing speed of the system, minimize system judgment errors.


