AI Text Mining Method for Automated New Term Discovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current text mining methods for new term discovery in natural language processing rely heavily on manual setting of feature thresholds, which is labor-intensive and inefficient, especially in the rapidly evolving internet era where new terms emerge frequently.
Innovation Solution
An artificial intelligence-based text mining method that uses machine learning algorithms to select new terms from domain candidate terms by obtaining domain candidate term features, calculating term quality scores, and determining new terms based on these scores, thereby eliminating the need for manual threshold setting.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If manual setting of feature thresholds is used for new term discovery, then the process can be controlled and adjusted, but the manpower cost is high and the efficiency is low
Solution Approach 1:
The system automatically discovers new terms by computing statistical features and applying machine learning classification without requiring manual threshold setting. The classifier self-adjusts to identify new terms based on learned patterns from training data, eliminating the need for continuous manual intervention while maintaining high efficiency
Solution Approach 2:
The patent replaces the mechanical manual threshold-setting process with an automated machine learning system. The mechanical operation of manually adjusting and setting thresholds is substituted by an electronic/computational system that automatically computes statistical features and classifies new terms using trained models
2Adaptability or versatility
If manual setting of feature thresholds is used for new term discovery, then the process can be customized, but the process is labor-intensive and time-consuming
Solution Approach 1:
The system performs preliminary action by pre-training machine learning classifiers with labeled training data before actual new term discovery. This preliminary training phase enables the system to automatically adapt to different domains and requirements without manual threshold setting during the actual discovery process, saving significant time
Solution Approach 2:
The patent changes the approach from manually setting fixed threshold parameters to dynamically computing statistical features (such as degree of solidification and degree of freedom) and using machine learning classifiers that adapt parameters automatically based on input data characteristics and domain requirements
3Measurement precision
If statistical methods with manual thresholds are used, then the selection criteria are clear, but the system cannot adapt to rapidly emerging new terms
Solution Approach 1:
The system transitions from static manual thresholds to dynamic machine learning classification. The classifier continuously adapts to newly emerging terms by learning from training data and adjusting its decision boundaries automatically, enabling the system to remain accurate while adapting to rapidly changing language usage and new term creation
Data Source
AI summary
This application discloses a text mining method based on artificial intelligence performed by a computer device. This application includes: obtaining domain candidate term features corresponding to domain candidate terms; obtaining term quality scores corresponding to the domain candidate terms according to the domain candidate term features; determining a new term from the domain candidate terms according to the term quality scores corresponding to the domain candidate terms; obtaining an associated text according to the new term; and determining a domain seed term as a domain new term in response to determining according to the associated text that the domain seed term satisfies a domain new term mining condition. By this application, new terms can be automatically selected from domain candidate terms based on a machine learning algorithm, thereby reducing manpower costs and well adapting to the rapid emergence of special new terms in the Internet era.


