Unsupervised Multilingual Topic Modeling for Propaganda Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for analyzing multilingual text introduce biases and inefficiencies, as they rely on predetermined topics, human scoring, and language-specific assessments, making it difficult to identify propaganda and sentiment in large datasets without introducing cultural or systemic biases.
Innovation Solution
An unsupervised multilingual topic modeling approach that uses a term-by-document matrix and Pointwise Mutual Information weighting to identify prevalent and distinctive topics across languages, eliminating the need for pre-defined sentiment-bearing words and human assessment, and allowing for flexible adaptation to different formats and languages.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If predetermined topics and pre-evaluated word scoring functions are used, then opinion analysis can be performed, but bias is introduced into the system
Solution Approach 1:
The system performs preliminary unsupervised topic modeling to discover topics that actually exist in the data before analyzing opinion. This preliminary action reveals the true topic structure without preconceptions, allowing subsequent opinion analysis to be grounded in actual data patterns rather than predetermined assumptions.
Solution Approach 2:
The system uses the data itself to define topics and opinion expressions through unsupervised learning, rather than relying on external pre-evaluated word lists. The data serves its own analysis needs by automatically revealing its inherent structure and meaning patterns.
2Productivity
If human scorers are used to assess sentiment attributes, then sentiment analysis can be performed, but unpredictable bias and high costs are introduced
Solution Approach 1:
The system replaces the mechanical process of human scoring with an automated computational approach using unsupervised topic modeling and pattern recognition. This substitution eliminates human bias and inconsistency while maintaining the ability to assess sentiment attributes at scale.
Solution Approach 2:
The system creates a computational model that copies and generalizes from limited annotated examples rather than relying on continuous human judgment. By learning patterns from a small set of labeled data, the system can automatically apply these patterns to analyze sentiment in large volumes of text without additional human involvement.
3Adaptability or versatility
If different groups of humans are used for each language, then language-specific sentiment can be assessed, but cultural biases are introduced
Solution Approach 1:
The system uses a universal unsupervised topic modeling approach that works across multiple languages without requiring language-specific customization. The same computational methodology automatically adapts to different languages, ensuring consistent and unbiased sentiment analysis across linguistic boundaries.
4Productivity
If predetermined topics are used in opinion analysis, then opinion can be measured, but the presence of topics saying something about opinion or writer priorities is not addressed
Solution Approach 1:
The system performs preliminary unsupervised topic modeling to discover what topics actually exist in the data before analyzing opinion. This reveals the true topic structure and allows the system to capture information about which topics are present and their relative importance, rather than forcing data into predetermined categories.
Data Source
AI summary
Disclosed is a scalable method for automatically deriving the topics discussed most prevalently in unstructured, multilingual text, and simultaneously revealing which topics are more biased towards one or another ‘information space’. The concept of the ‘information space’ is derived from Russian strategic doctrine on information warfare; an example of an ‘information space’ would be the portion of social media in which the Russian language is used. The disclosed method leverages this concept, in conjunction with unsupervised multilingual machine learning, to determine, automatically and without any built-in bias or preconceived notions of what is important, which topics are more discussed, for example, in one language than another. An analyst's attention can then be focused on the most important differences between national discourses, and insight more quickly gained into the areas (both topics and geographic regions) in which propaganda of the sort envisaged in Russian strategic doctrine may be taking hold.


