Automated Document Credibility Scoring via Semantic Similarity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Manually verifying the accuracy of a large number of documents in a database is time-consuming, and existing techniques for generating document summarization do not effectively address the credibility of content within these documents.

Innovation Solution

A method and system that utilize a processor to extract topics from documents, generate topic combinations, obtain summaries, perform semantic similarity tests, and calculate credibility scores based on similarity percentages to automatically determine the credibility of content in documents stored in a database.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual verification of document accuracy is performed, then credibility assessment can be achieved, but time consumption increases significantly

Engineering Contradiction:
Improvecredibility assessment accuracyVSAvoidverification time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system enables documents to assess their own credibility through automated processing. The processor extracts topics, generates summaries, and calculates credibility scores for each document independently, eliminating the need for manual verification while maintaining assessment accuracy through self-contained automated analysis

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual verification process with an automated computational system. The processor performs topic extraction, summary generation, and credibility scoring through algorithmic operations, substituting human manual assessment with automated mechanical processing that is both faster and scalable

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If existing document summarization techniques are used, then summarization can be generated, but credibility determination is not effectively addressed

Engineering Contradiction:
Improvesummarization efficiencyVSAvoidcredibility determination
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent merges document summarization with credibility determination into a single integrated process. The system simultaneously generates summaries and calculates credibility scores by combining topic extraction, summary generation, and semantic similarity testing, achieving both productivity and reliability in one unified system

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The processor performs multiple functions: extracting topics from documents, generating summaries based on topic combinations, conducting semantic similarity tests, and calculating credibility scores. This multi-functional approach allows the system to both summarize documents efficiently and determine their credibility reliably through a single automated process

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11170034B1System and method for determining credibility of content in a number of documents
Publication Date: 2021.11.09 IDOX AI CORP
  • US11170034B1 patent drawing
  • US11170034B1 patent drawing
  • US11170034B1 patent drawing

AI summary

A method for determining credibility of content in a number of documents includes: obtaining topics from each document; for each document, generating topic combinations, each topic combination being a subset of the topics of the document; for each topic combination, obtaining a summary from the corresponding document; performing a semantic similarity test on each pair of two summaries that are respectively from two documents, so as to obtain a similarity percentage between the two summaries; for a group of the topic combinations that are identical combinations of topic(s), calculating a credibility score for the group based on the similarity percentage(s) calculated for the summaries that correspond to the topic combinations in the group.