Semantic Graph Accuracy via External Corpus Validation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing computational linguistics methods face challenges in accurately measuring the quality of semantic similarity graphs, particularly for unsupervised learning techniques, as they lack effective methods for assessing performance due to the absence of a training set and discernible mechanisms for testing results.
Innovation Solution
A scoring system that leverages exogenous information to quantify the quality of semantic graphs by using an external dataset to evaluate the accuracy of edges in the graph, based on neighboring nodes and their connections, providing a more reliable assessment by avoiding self-consistent misleading evaluations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If human review is used to assess algorithm accuracy, then measurement precision may be improved, but productivity deteriorates due to time and cognitive limitations
Solution Approach 1:
The patent introduces an external corpus as an intermediary reference standard to assess the accuracy of semantic graph edges. Instead of relying solely on human reviewers, the system uses pre-existing external documents that contain ground truth relationships to automatically evaluate whether edges in the generated semantic graph are accurate, thereby maintaining measurement precision while eliminating human time constraints
Solution Approach 2:
The patent replaces the mechanical human review process with an automated computational evaluation system. The system automatically compares edges in the semantic graph against relationships found in the external corpus using computational linguistics methods, substituting human cognitive assessment with machine-based measurement that achieves both high precision and scalability
2Measurement precision
If manual review of each document is performed, then measurement precision improves, but loss of time increases making it difficult to review large collections
Solution Approach 1:
The patent extracts a separate external corpus from the main document collection to serve as an independent reference standard. This external corpus contains pre-established relationships that can be used to evaluate the semantic graph without requiring review of every document in the main collection, thereby reducing time loss while maintaining assessment precision
Solution Approach 2:
The external corpus is prepared in advance with pre-established ground truth relationships before the semantic graph generation process. This preliminary preparation allows the evaluation system to quickly assess graph accuracy by comparing against pre-validated references, eliminating the need for time-consuming document-by-document review during the assessment phase
3Device complexity
If self-consistent evaluation methods are used, then device complexity is reduced, but reliability deteriorates due to misleading results
Solution Approach 1:
The patent introduces an external corpus as an independent intermediary reference that is separate from the semantic graph generation process. This external reference acts as an unbiased ground truth standard, preventing the self-consistency bias where the evaluation system might validate its own outputs. The external corpus provides objective criteria for assessing edge accuracy without being influenced by the internal logic of the semantic graph construction
Solution Approach 2:
Instead of having the semantic graph evaluation system validate itself (self-consistent approach), the patent inverts the evaluation direction by using an external corpus to validate the graph. The external corpus serves as the primary truth source, and the semantic graph is evaluated against it, reversing the traditional self-validation approach and thereby improving reliability
Data Source
AI summary
Provided is a process including: obtaining a semantic similarity graph having nodes corresponding to documents in an analyzed corpus and edges indicating semantic similarity between pairs of the documents; for at least a plurality of nodes in the graph, evaluating accuracy of the edges based on neighboring nodes and an external corpus by performing operations including: identifying the neighboring nodes based on adjacency to the respective node in the graph; selecting documents from an external corpus based on references in the selected documents to entities mentioned in the documents of the neighboring nodes; and determining how semantically similar the respective node is to the selected documents.


