Voice Text Classification via Lexical Template Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current intelligent voice interaction services face challenges in efficiently classifying voice-recognized texts at the early stage due to the need for extensive manual annotation and high initial text corpora collection, resulting in low annotation efficiency.

Innovation Solution

An artificial intelligence-based method and apparatus that acquire and analyze voice queries using a lexical analyzer to match templates in a classifier, allowing for rapid text classification, with optional manual generalization and updating of templates to improve classification accuracy and reduce manual annotation requirements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual annotation is used for text classification, then classification accuracy can be improved, but annotation efficiency deteriorates due to considerable annotating manpower requirements

Engineering Contradiction:
Improveclassification accuracyVSAvoidannotation efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system performs self-service by automatically classifying voice-recognized texts through lexical analysis and template matching, reducing dependency on manual annotation. The classifier autonomously processes texts by comparing lexical structures against predefined templates, enabling the system to serve itself without extensive human intervention while maintaining classification accuracy.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent applies preliminary action by pre-establishing lexical templates for different text categories before classification occurs. These templates contain predefined lexical structures and patterns that enable rapid matching and classification of incoming texts, eliminating the need for time-consuming manual annotation while preserving accuracy through预先 prepared classification criteria.

Inventive Principle:
Principle #10Preliminary action

2Quantity of substance

If extensive text corpora collection is performed, then training data availability is improved, but the cold start problem worsens due to delayed classification capability

Engineering Contradiction:
Improvetext corpora volumeVSAvoidcold start delay
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The system performs preliminary action by pre-building a classifier with lexical templates before actual text classification begins. This allows the system to immediately classify texts upon deployment without waiting for extensive corpora collection and manual annotation, effectively solving the cold start problem while still utilizing available text data for template creation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces the mechanical process of manual text annotation with an automated lexical analysis system. Instead of relying on human annotators to process texts sequentially, the system uses computational lexical structure analysis and template matching to automatically classify texts, dramatically reducing the time required for classification capability deployment.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Productivity

If automated classification is implemented, then annotation efficiency is improved, but the ability to handle imperfect text corpora deteriorates due to classification challenges

Engineering Contradiction:
Improveannotation efficiencyVSAvoidclassification reliability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent applies parameter changes by transforming the classification approach from content-based semantic analysis to structural lexical pattern matching. By changing the classification parameter from meaning interpretation to lexical structure comparison against templates, the system achieves high efficiency while maintaining reliability even with imperfect texts, as lexical structures remain detectable regardless of text quality.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10762901B2Artificial intelligence based method and apparatus for classifying voice-recognized text
Publication Date: 2020.09.01 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US10762901B2 patent drawing
  • US10762901B2 patent drawing
  • US10762901B2 patent drawing

AI summary

Embodiments of the present disclosure disclose an artificial intelligence based method and apparatus for classifying a voice-recognized text. A specific embodiment of the method includes: acquiring a current interactive text of a voice query from a user; analyzing the current interactive text using a lexical analyzer to obtain a current lexical structure; determining whether the current lexical structure matches a template of a category in a classifier; and classifying, if the current lexical structure matches the template of the category in the classifier, the current interactive text corresponding to the current lexical structure into the category belonging to the matched template. The embodiment can fast classify texts, effectively reduce the magnitude of manually annotated texts, and improve the annotation efficiency in intelligent voice interaction services.