Synonym Identification via Anchor Text Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for determining synonymous names of entities in fact databases are impractical, expensive, and prone to human errors due to the vast number of objects in repositories, making it difficult for users to find relevant answers when searching with different names.
Innovation Solution
A method that identifies source documents for an entity, processes anchor texts from linking documents to generate synonym candidates, and selects synonymous names for storage in the repository, using a system architecture that includes importers, janitors, a build engine, and a service engine to manage and query facts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional methods (consulting people familiar with entities) are used to determine synonymous names, then accuracy of synonym identification may be maintained, but productivity becomes insufficient and cost increases due to the vast number of objects
Solution Approach 1:
The system automatically determines synonymous names by processing anchor texts from linking documents itself, without requiring human experts. The computer-implemented method extracts and analyzes anchor texts to generate synonym candidates, enabling the system to serve itself in identifying synonyms at scale.
Solution Approach 2:
The patent replaces the manual mechanical process of consulting human experts with an automated computational system. The method uses computer algorithms to process anchor texts, generate synonym candidates, and store them in the repository, substituting human cognitive work with automated text processing.
2Reliability
If manual methods are used to determine synonymous names, then quality control may be maintained, but human errors increase and cost increases
Solution Approach 1:
The system performs synonym identification automatically without human intervention, eliminating human errors associated with manual methods. The automated process consistently applies the same algorithmic rules to all objects, ensuring uniform quality control across the entire repository.
3Quantity of substance
If the repository contains vast numbers of objects with multiple synonymous names, then completeness of information improves, but search functionality deteriorates when name mismatches occur
Solution Approach 1:
The system enhances search functionality by storing multiple synonymous names for each object, enabling the repository to respond to various search queries using different names for the same entity. This multi-functional approach allows users to search using any synonymous name and still retrieve the correct object.
Data Source
AI summary
A repository contains objects representing entities. The objects also include facts about the represented entities. The facts are derived from source documents. A synonymous name of an object is determined by identifying a source document from which one or more facts of the entity represented by the object were derived, identifying a plurality of linking documents that link to the source document through hyperlinks, each hyperlink having an anchor text, processing the anchor texts in the plurality of linking documents to generate a collection of synonym candidates for the entity represented by the object, and selecting a synonymous name for the entity represented by the object from the collection of synonym candidates.


