String Disambiguation via Clique Graphs and Arrival Probability
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing reader aggregators are inefficient in analyzing and providing relevant content to users, as they lack effective disambiguation and entity ranking capabilities, leading to time-consuming manual checks and resource drainage on computing devices.
Innovation Solution
A computer-implemented method using a disambiguation database and clique graphs to disambiguate strings in articles and rank entities based on their relevance, by analyzing articles to identify candidate entities and generating probabilities through metrics such as link frequency and arrival probability, ensuring accurate content delivery to users.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If reader aggregators download all new articles from websites, then users receive comprehensive content updates, but computing device resources are drained and users must manually check websites
Solution Approach 1:
The patent extracts and analyzes specific entities and strings from downloaded articles using NLP techniques, rather than processing entire articles. This selective extraction of relevant information (entities, strings, clique graphs) reduces the computational burden while maintaining content delivery effectiveness.
Solution Approach 2:
The system performs preliminary disambiguation and entity ranking analysis on articles before presenting them to users. By pre-processing articles to identify and rank relevant entities, the system reduces the need for manual user checking and optimizes resource usage during actual reading.
2Quantity of substance
If reader aggregators provide all new content from subscribed websites, then users have access to complete information feeds, but analysis capability and content configuration are limited
Solution Approach 1:
The patent segments articles into discrete entities and strings, creating structured representations (clique graphs) that can be independently analyzed and ranked. This segmentation transforms unstructured text into analyzable components, enhancing analysis capability while maintaining information completeness.
Solution Approach 2:
The system adds a new dimension of analysis by creating clique graphs and entity rankings based on disambiguation metrics. This transforms the flat information feed into a multi-dimensional structure that enables sophisticated content configuration and relevance ranking.
3Ease of operation
If manual website checking is performed, then users control content access, but time consumption increases
Solution Approach 1:
The system performs automatic disambiguation, entity identification, and ranking operations without user intervention. Users simply access their personalized feed, and the system self-manages the complex tasks of analyzing articles, identifying entities, and ranking content based on user interests.
Solution Approach 2:
The system uses user interaction data and entity ranking metrics to continuously improve content delivery. By analyzing which entities and articles users engage with, the system refines its disambiguation and ranking algorithms, providing increasingly personalized content without requiring manual user checking.
Data Source
AI summary
Aspects of the present disclosure may involve a computer implemented method of disambiguating a string from an article involving an electronic device including one or more hardware processing units, and accessing a disambiguation database comprising a plurality of string-entity combinations and associated metrics for each of the plurality of string-entity combinations. The associated metrics may include a metric associated with an arrival probability of linking at a web page for a particular entity after a specified number of links from a starting page. The method may involve generating a clique graph for each candidate entity of the article, and generating a probability that a particular candidate entity matches a particular string associated with the particular candidate entity as a function of score attributes generated from the clique graph and the arrival probability of linking at the page for the particular entity after the specified number of links.


