Voice Text Classification via Lexical Template Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current intelligent voice interaction services face challenges in efficiently classifying voice-recognized texts at the early stage due to the need for extensive manual annotation and high initial text corpora collection, resulting in low annotation efficiency.
Innovation Solution
An artificial intelligence-based method and apparatus that acquire and analyze voice queries using a lexical analyzer to match templates in a classifier, allowing for rapid text classification, with optional manual generalization and updating of templates to improve classification accuracy and reduce manual annotation requirements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual annotation is used for text classification, then classification accuracy can be improved, but annotation efficiency deteriorates due to considerable annotating manpower requirements
Solution Approach 1:
The system performs self-service by automatically classifying voice-recognized texts through lexical analysis and template matching, reducing dependency on manual annotation. The classifier autonomously processes texts by comparing lexical structures against predefined templates, enabling the system to serve itself without extensive human intervention while maintaining classification accuracy.
Solution Approach 2:
The patent applies preliminary action by pre-establishing lexical templates for different text categories before classification occurs. These templates contain predefined lexical structures and patterns that enable rapid matching and classification of incoming texts, eliminating the need for time-consuming manual annotation while preserving accuracy through预先 prepared classification criteria.
2Quantity of substance
If extensive text corpora collection is performed, then training data availability is improved, but the cold start problem worsens due to delayed classification capability
Solution Approach 1:
The system performs preliminary action by pre-building a classifier with lexical templates before actual text classification begins. This allows the system to immediately classify texts upon deployment without waiting for extensive corpora collection and manual annotation, effectively solving the cold start problem while still utilizing available text data for template creation.
Solution Approach 2:
The patent replaces the mechanical process of manual text annotation with an automated lexical analysis system. Instead of relying on human annotators to process texts sequentially, the system uses computational lexical structure analysis and template matching to automatically classify texts, dramatically reducing the time required for classification capability deployment.
3Productivity
If automated classification is implemented, then annotation efficiency is improved, but the ability to handle imperfect text corpora deteriorates due to classification challenges
Solution Approach 1:
The patent applies parameter changes by transforming the classification approach from content-based semantic analysis to structural lexical pattern matching. By changing the classification parameter from meaning interpretation to lexical structure comparison against templates, the system achieves high efficiency while maintaining reliability even with imperfect texts, as lexical structures remain detectable regardless of text quality.
Data Source
AI summary
Embodiments of the present disclosure disclose an artificial intelligence based method and apparatus for classifying a voice-recognized text. A specific embodiment of the method includes: acquiring a current interactive text of a voice query from a user; analyzing the current interactive text using a lexical analyzer to obtain a current lexical structure; determining whether the current lexical structure matches a template of a category in a classifier; and classifying, if the current lexical structure matches the template of the category in the classifier, the current interactive text corresponding to the current lexical structure into the category belonging to the matched template. The embodiment can fast classify texts, effectively reduce the magnitude of manually annotated texts, and improve the annotation efficiency in intelligent voice interaction services.


