Unsupervised Fringe Belief Detection via Word Co-occurrence

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for detecting fringe beliefs in text data are limited by their reliance on language-specific techniques, supervised machine learning, and the need for training data, which can fail to identify emerging beliefs that have not been fact-checked or are not mainstream, and often require significant user configuration and labeled data.

Innovation Solution

A computer-implemented method using unsupervised machine learning techniques to identify fringe beliefs in text data by analyzing word co-occurrences and statistical properties, without relying on training data or specific vocabulary lists, allowing for language-independent and flexible identification of anomalous clusters in text datasets.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If language-specific techniques and supervised machine learning are used to detect fringe beliefs, then detection accuracy for known beliefs is improved, but the system fails to identify emerging beliefs and requires significant user configuration and labeled data

Engineering Contradiction:
Improvedetection accuracyVSAvoidability to identify emerging beliefs
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system performs self-training by automatically selecting and labeling training examples from the input data itself, eliminating the need for external labeled datasets. The algorithm iteratively identifies candidate fringe beliefs, uses them to train the classifier, and refines its detection capabilities autonomously, enabling it to adapt to emerging beliefs without user configuration

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system transitions from supervised learning with fixed parameters to an adaptive approach where training parameters (labeled data, vocabulary lists) are dynamically generated from the data itself. This allows the system to detect emerging beliefs by changing its learning parameters based on observed patterns rather than relying on pre-configured settings

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If supervised machine learning with labeled data is used, then detection precision for mainstream fringe beliefs is improved, but the system cannot detect beliefs that have not been fact-checked or are not on the radar of investigators

Engineering Contradiction:
Improvedetection precisionVSAvoidemergent fringe beliefs
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The system performs preliminary unsupervised clustering analysis to identify potential fringe belief clusters before applying supervised classification. This preliminary action creates candidate training examples from the data itself, allowing the system to detect emerging beliefs that have not yet been labeled by fact-checkers or investigators

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system introduces an intermediary unsupervised learning stage that bridges the gap between raw data and supervised classification. This intermediary process generates training examples and vocabulary lists from the data itself, enabling detection of emerging beliefs without relying on external labeled sources

Inventive Principle:
Principle #24Intermediary (Mediator)

3Difficulty of detecting and measuring

If deep neural networks with training datasets are used for anomaly detection, then detection capability is improved, but the system requires training data that is often unavailable and needs user configuration

Engineering Contradiction:
Improveanomaly detection capabilityVSAvoidconfiguration requirements
Core Design Contradiction:
Difficulty of detecting and measuringVSDevice complexity

Solution Approach 1:

The system eliminates the need for external training datasets by performing self-training using the input data itself. The algorithm automatically generates training examples, selects relevant vocabulary, and configures its parameters without user intervention, reducing both data requirements and configuration complexity while maintaining anomaly detection capability

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

Instead of requiring training data to detect anomalies, the system inverts the approach by using unsupervised clustering on the input data to generate the training examples needed for supervised detection. This inversion eliminates the dependency on external training datasets and user configuration

Inventive Principle:
Principle #13The other way round (Inversion)

4Measurement precision

If sentiment analysis and language-specific features are used, then detection accuracy for English text is improved, but the system cannot be applied to other languages and requires extensive configuration

Engineering Contradiction:
Improvedetection accuracyVSAvoidlanguage independence
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system uses unsupervised topic modeling and clustering techniques that are language-agnostic, allowing the same algorithm to detect fringe beliefs in multiple languages without reconfiguration. The approach identifies patterns based on word co-occurrence and statistical properties rather than language-specific features, providing universal applicability across different languages

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12099538B2Identifying fringe beliefs from text
Publication Date: 2024.09.24 GALISTEO CONSULTING GRP
  • US12099538B2 patent drawing
  • US12099538B2 patent drawing
  • US12099538B2 patent drawing

AI summary

Disclosed is a flexible, scalable method for automatically distinguishing between the ‘mainstream’ and ‘fringe’ in any text dataset, such as unstructured social media data. The disclosed method allows analysts quickly to pinpoint text articulating fringe beliefs or theories, without prior knowledge of the nature of the fringe beliefs, and can be applied to text data in any language, provided the data can be represented electronically. The method works automatically and without any preconceived notions of what is important or what vocabulary is used in the context of certain beliefs. An analyst's attention is then quickly focused either on new conspiracy theories taking hold (potentially allowing decision-makers to act to interdict the ‘next QAnon insurrection’), or on well-founded beliefs that simply are not yet mainstream. Either way, analysts are empowered better to ‘connect the dots’ for a more complete understanding of the information environment.