Stance Classification Using Text Analytics and Rule-Based Classifiers

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current text analytics methods lack the ability to efficiently classify the stance of individuals on debatable topics from large volumes of digital textual information, particularly in identifying the positions of persons mentioned in encyclopedic entries and indices, which limits the generation of comprehensive databases and visualizations of stances.

Innovation Solution

A computerized text analysis method that automatically searches digital resources using debatable topics and personal derivations, applies rule-based classifiers and machine learning algorithms to determine stances, and generates stance graphs to visualize the positions of persons on debatable topics, utilizing resources like Wikipedia and lexical databases to identify hyperlinks and associations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual analysis of digital textual information is used, then accuracy of stance classification is maintained, but processing time and labor requirements increase significantly

Engineering Contradiction:
Improveaccuracy of stance classificationVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual human analysis with automated computer-based text analytics systems. The system uses natural language processing algorithms, machine learning models, and computational methods to automatically classify stances of individuals on debatable topics from large volumes of digital text, eliminating the need for manual processing while maintaining classification accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system employs self-learning machine learning models that automatically improve their classification accuracy through training on annotated data. The automated pipeline includes self-contained steps for text retrieval, processing, feature extraction, and stance classification that operate independently without requiring continuous human intervention.

Inventive Principle:
Principle #25Self-service

2Productivity

If automated text analytics are applied to large volumes of digital text, then processing speed increases, but complexity of the analysis system increases

Engineering Contradiction:
Improveprocessing speedVSAvoidcomplexity of analysis system
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent divides the complex text analytics pipeline into distinct modular segments: text retrieval module, text processing module, feature extraction module, classification module, and visualization module. Each segment performs a specific function and can be independently optimized or replaced, reducing overall system complexity while maintaining high processing speed.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system uses a universal text processing framework that handles multiple tasks through a single integrated platform. The same infrastructure supports text retrieval from various sources, multiple types of text processing operations, different classification models, and various visualization outputs, reducing the need for separate specialized systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Quantity of substance

If comprehensive search of digital resources is performed, then completeness of stance identification improves, but time required for data collection increases

Engineering Contradiction:
Improvecompleteness of stance identificationVSAvoidtime for data collection
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The system performs preliminary actions by pre-indexing and pre-processing digital textual resources before actual analysis begins. Text data is pre-processed, structured, and organized in databases, allowing rapid retrieval and analysis during the actual stance identification task without requiring time-consuming on-the-fly processing.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces intermediary data structures and indexing mechanisms that mediate between the raw digital text resources and the final stance classification. These intermediaries include structured data formats, knowledge graphs, and indexed databases that enable efficient querying and comprehensive coverage without direct processing of all raw text.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Reliability

If rule-based classification is applied, then interpretability of results improves, but flexibility in handling diverse textual patterns decreases

Engineering Contradiction:
Improveinterpretability of resultsVSAvoidflexibility in handling textual patterns
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent merges rule-based classification systems with machine learning-based classification systems into a hybrid approach. The rule-based component provides interpretability and handles clear-cut cases, while the machine learning component provides flexibility and adapts to diverse textual patterns. Both components work together to process text data, with the system automatically selecting the appropriate method based on the characteristics of the input text.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS11341188B2Expert stance classification using computerized text analytics
Publication Date: 2022.05.24 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11341188B2 patent drawing
  • US11341188B2 patent drawing
  • US11341188B2 patent drawing

AI summary

A computerized text analysis method that comprises: searching a resource of information with a search query comprising at least one of: (a) the specific debatable topic, and (b) a personal derivation of the specific debatable topic, to obtain a list of indices whose index subject contains the personal derivation and/or the specific debatable topic; determining, by applying a rule-based classifier, whether the index subject of each of the indices is (i) in favor of the debatable topic or (ii) against the debatable topic; detecting, in each of the indices, hyperlinks to encyclopedic entries whose entry subjects are person names; and determining that: if the index subject of each of the one or more indices is in favor of the specific debatable topic, then the persons are in favor of the specific debatable topic, and vice versa.