Sentiment Analysis Model Training with Automated Tag Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current sentiment analysis models require extensive manual annotation of large datasets, consuming significant resources and time, which is inefficient and costly.

Innovation Solution

A method and apparatus that utilize a pre-established sentiment analysis model trained using untagged and tagged sample data, where tag information is generated using multiple tag generation models to extend the sample data, reducing the need for manual annotation and improving accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual annotation is performed on large amounts of data to establish a sentiment analysis model, then the training effect and model accuracy are improved, but the consumption of manpower and material resources increases significantly

Engineering Contradiction:
Improvemodel accuracyVSAvoidmanpower and material resources
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent applies preliminary action by pre-establishing a tag generation model before training the sentiment analysis model. This tag generation model automatically generates tag information for untagged sample data, preparing the data in advance without requiring manual annotation during the main training process. This resolves the contradiction by performing the labor-intensive tagging work automatically before the actual model training, reducing resource consumption while maintaining model accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements self-service through the tag generation model that automatically generates tag information for untagged sample data using pre-trained word vectors and classification algorithms. The system serves itself by automatically creating training data without external manual intervention, thereby reducing manpower resources while ensuring sufficient training data for accurate model performance.

Inventive Principle:
Principle #25Self-service

2Reliability

If extensive manual annotation is performed to distribute sentiment words and sentence patterns widely, then the training effect is improved, but the time consumption increases

Engineering Contradiction:
Improvetraining effectVSAvoidannotation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs preliminary action by pre-training word vectors and establishing the tag generation model before the sentiment analysis model training. This preparation phase automatically creates a comprehensive distribution of sentiment words and sentence patterns through algorithmic processing, eliminating the need for time-consuming manual annotation while ensuring wide coverage for effective training.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces the mechanical manual annotation system with an automated computational system. The tag generation model uses pre-trained word vectors, text segmentation, and classification algorithms to automatically generate tag information, substituting human manual work with automated mechanical processes that are faster and more scalable, thereby reducing time loss while maintaining training effectiveness.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If multiple tag generation models are used to generate tag information, then the accuracy of tag generation is improved, but the device complexity increases

Engineering Contradiction:
Improvetag generation accuracyVSAvoidmodel structure complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies merging by combining multiple tag generation models (first, second, and third tag generation models) into a unified sentiment analysis training system. Each model focuses on different aspects (all words, sentiment words, non-sentiment words), and their outputs are integrated to generate comprehensive tag information. This resolves the contradiction by merging multiple specialized models into a coordinated system that improves accuracy while managing complexity through functional specialization.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent implements segmentation by dividing the tag generation task into three separate models, each handling different aspects: the first model processes all words, the second focuses on sentiment words, and the third handles non-sentiment words. This segmentation allows each model to specialize and achieve higher accuracy in its specific domain while the overall system manages complexity through modular architecture.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11062089B2Method and apparatus for generating information
Publication Date: 2021.07.13 BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
  • US11062089B2 patent drawing
  • US11062089B2 patent drawing
  • US11062089B2 patent drawing

AI summary

A method and an apparatus for generating information are provide according to embodiments of the disclosure. A specific embodiment of the method comprises: acquiring to-be-analyzed information according to a target keyword; and inputting the to-be-analyzed information into a pre-established sentiment analysis model to generate sentiment orientation information of the to-be-analyzed information. The sentiment analysis model is obtained through following training: acquiring untagged sample data and tagged sample data; generating tag information corresponding to the untagged sample data using a pre-established tag generation model, and using the untagged sample data and the generated tag information as extended sample data, the tag generation model being used to represent a corresponding relationship between the untagged sample data and the tag information; and obtaining the sentiment analysis model by training using the tagged sample data and the extended sample data.