AI Text Mining Method for Automated New Term Discovery

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current text mining methods for new term discovery in natural language processing rely heavily on manual setting of feature thresholds, which is labor-intensive and inefficient, especially in the rapidly evolving internet era where new terms emerge frequently.

Innovation Solution

An artificial intelligence-based text mining method that uses machine learning algorithms to select new terms from domain candidate terms by obtaining domain candidate term features, calculating term quality scores, and determining new terms based on these scores, thereby eliminating the need for manual threshold setting.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If manual setting of feature thresholds is used for new term discovery, then the process can be controlled and adjusted, but the manpower cost is high and the efficiency is low

Engineering Contradiction:
Improvemanual controlVSAvoidnew term discovery efficiency
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The system automatically discovers new terms by computing statistical features and applying machine learning classification without requiring manual threshold setting. The classifier self-adjusts to identify new terms based on learned patterns from training data, eliminating the need for continuous manual intervention while maintaining high efficiency

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual threshold-setting process with an automated machine learning system. The mechanical operation of manually adjusting and setting thresholds is substituted by an electronic/computational system that automatically computes statistical features and classifies new terms using trained models

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If manual setting of feature thresholds is used for new term discovery, then the process can be customized, but the process is labor-intensive and time-consuming

Engineering Contradiction:
Improveprocess customizationVSAvoidtime for threshold setting
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system performs preliminary action by pre-training machine learning classifiers with labeled training data before actual new term discovery. This preliminary training phase enables the system to automatically adapt to different domains and requirements without manual threshold setting during the actual discovery process, saving significant time

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes the approach from manually setting fixed threshold parameters to dynamically computing statistical features (such as degree of solidification and degree of freedom) and using machine learning classifiers that adapt parameters automatically based on input data characteristics and domain requirements

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If statistical methods with manual thresholds are used, then the selection criteria are clear, but the system cannot adapt to rapidly emerging new terms

Engineering Contradiction:
Improveterm selection accuracyVSAvoidadaptability to new terms
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system transitions from static manual thresholds to dynamic machine learning classification. The classifier continuously adapts to newly emerging terms by learning from training data and adjusting its decision boundaries automatically, enabling the system to remain accurate while adapting to rapidly changing language usage and new term creation

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20230111582A1Text mining method based on artificial intelligence, related apparatus and device
Publication Date: 2023.04.13 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • US20230111582A1 patent drawing
  • US20230111582A1 patent drawing
  • US20230111582A1 patent drawing

AI summary

This application discloses a text mining method based on artificial intelligence performed by a computer device. This application includes: obtaining domain candidate term features corresponding to domain candidate terms; obtaining term quality scores corresponding to the domain candidate terms according to the domain candidate term features; determining a new term from the domain candidate terms according to the term quality scores corresponding to the domain candidate terms; obtaining an associated text according to the new term; and determining a domain seed term as a domain new term in response to determining according to the associated text that the domain seed term satisfies a domain new term mining condition. By this application, new terms can be automatically selected from domain candidate terms based on a machine learning algorithm, thereby reducing manpower costs and well adapting to the rapid emergence of special new terms in the Internet era.