Document Credibility Index via Topic-Sentiment Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for determining document credibility are often subjective and inaccurate, particularly in the context of the Internet and World Wide Web, where text is frequently reproduced for reasons other than semantic content, leading to skewed credibility determinations based on source reliability or frequency of topic expressions.
Innovation Solution
The approach involves determining sentiments corresponding to topics within a document, using topic and sentiment models to calculate a credibility index as a weighted average of topic-sentiment scores, focusing on the document's content rather than external source credibility.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If credibility is determined based on source reliability or frequency of topic expressions, then the process is simplified and can be automated, but the accuracy of credibility determination deteriorates due to subjective biases and text reproduction for non-semantic reasons
Solution Approach 1:
The patent replaces manual, subjective credibility assessment with an automated computational system that uses natural language processing and machine learning algorithms to objectively analyze document sentiments and determine credibility, thereby maintaining automation while improving measurement precision through data-driven analysis rather than human judgment
Solution Approach 2:
The patent introduces sentiment analysis as an intermediary layer between the raw document text and the final credibility determination. This intermediary process extracts and quantifies sentiments corresponding to topics in the document, providing an objective bridge that improves accuracy while maintaining automated operation
2Measurement precision
If credibility is determined by gathering data from a large number of knowledgeable people, then the accuracy of credibility assessment improves, but the effort and time required deteriorates significantly
Solution Approach 1:
The patent enables documents to self-assess their own credibility through automated sentiment analysis of their content. The system extracts sentiments directly from the document text without requiring external human evaluators, thereby achieving accurate credibility determination while eliminating the time-consuming process of gathering data from multiple knowledgeable people
3Ease of operation
If text frequency is used as a proxy for credibility, then the measurement process is simplified, but the accuracy deteriorates because text is reproduced for reasons other than semantic content
Solution Approach 1:
The patent changes the measurement parameter from simple text frequency to sentiment scores corresponding to topics in the document. This transformation maintains ease of automated operation while improving accuracy by capturing the semantic meaning and emotional tone of the text rather than merely counting occurrences, thereby distinguishing between genuine belief and reproduction for other purposes
Data Source
AI summary
A plurality of topics encompassed in a document are determined and, for each such topic, a sentiment for that topic is likewise determined. Thereafter, credibility of the document is determined based on the resulting plurality of sentiments. In one embodiment, credibility of at least one target document is established by first determining, for each of a plurality of portions of the at least one target document, at least one topic encompassed in the portion to provide a plurality of target topics. Likewise, sentiment scores are determined for each portion. Thereafter, for each prior topic of a plurality of prior topics, a topic-sentiment score is determined based on sentiment scores corresponding to those portions of the plurality of portions having a target topic corresponding to the prior topic. A credibility index is determined based on the resulting plurality of topic-sentiment scores.


