Autonomous Sentiment Analysis Training Corpus Builder

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Manual selection and labeling of training samples for sentiment analysis computer models is time-consuming and labor-intensive, leading to lower confidence in sentiment determination, especially when dealing with large volumes of data.

Innovation Solution

An autonomous system that extracts semantic and sentiment elements from textual data, generalizes these elements, generates informative ranking scores, selects informative samples, and constructs a training corpus without human intervention, using components like named entity recognizers, feature term detectors, opinion word detectors, and semantic generalizers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual selection and labeling of training samples is used, then high quality training data can be obtained, but the process becomes time-consuming and labor-intensive

Engineering Contradiction:
Improvequality of training dataVSAvoidtime for sample selection and labeling
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs self-service by automatically selecting and labeling training samples without human intervention. The sentiment analysis model autonomously processes textual data, extracts sentiment elements, generates labels, and ranks samples for inclusion in the training corpus, eliminating the need for manual labor while maintaining data quality

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The manual mechanical process of sample selection and labeling is replaced with an automated computational system. The patent substitutes human operators with a computer-based sentiment analysis model that uses natural language processing, semantic analysis, and machine learning algorithms to perform tasks previously requiring human judgment and effort

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If manual selection of training samples is used, then expert support can ensure domain-specific accuracy, but the process requires significant human resources

Engineering Contradiction:
Improvedomain-specific accuracyVSAvoidhuman resource requirements
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system achieves domain-specific reliability through self-service mechanisms where the automated model continuously learns from and adapts to domain-specific patterns in the data. The sentiment analysis system autonomously identifies domain-relevant features and adjusts its labeling criteria without requiring external expert intervention for each sampling decision

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent creates a universal system that can handle multiple domains and types of textual data. The sentiment analysis model is designed to be domain-agnostic initially but can adapt to specific domains through automated learning, making it versatile for product reviews, social media, customer service, and other text-based sentiment analysis tasks without requiring domain-specific manual configuration

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Productivity

If random sampling of training data is used, then processing speed increases, but confidence in sentiment determination decreases

Engineering Contradiction:
Improvedata processing speedVSAvoidconfidence in sentiment determination
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system changes the parameter of sample selection from random to informed/ranked selection based on sentiment analysis metrics. By calculating sentiment scores, confidence levels, and diversity metrics for each sample, the system transforms the selection process into a parameter-driven optimization problem where samples are chosen based on their informational value rather than random chance

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces an intermediary ranking mechanism that mediates between the large volume of available data and the limited training corpus size. The sentiment analysis model acts as an intermediary layer that evaluates all candidate samples, assigns rankings based on multiple criteria (sentiment clarity, diversity, representativeness), and selects the top-ranked samples, thereby maintaining both speed and reliability

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS10824812B2Method and apparatus for informative training repository building in sentiment analysis model learning and customization
Publication Date: 2020.11.03 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10824812B2 patent drawing
  • US10824812B2 patent drawing
  • US10824812B2 patent drawing

AI summary

The methods, systems, and computer program products described herein provide ways to generate an informative training corpus of samples for use in machine training a high-quality sentiment analysis computer model. In some aspects, a method is disclosed including receiving a plurality of training samples, extracting semantic and sentiment elements of one or more of the training samples, generalizing the semantic and sentiment elements of the one or more of the training samples, generating an informative ranking score for one or more of the training samples based on the generalized semantic and sentiment elements, selecting informative training samples from the plurality of training samples based at least in part on the generated informative ranking scores, and adding the selected informative training samples to an informative training corpus.