Unified Entity Storage Architecture for Scalable Coreference Resolution
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing database architectures struggle to scale effectively with large volumes of entities and complex relationships in unstructured data, leading to inefficiencies in entity-centric information management.
Innovation Solution
The implementation of a unified entity storage architecture, known as the Knowledge Base, which combines persistent storage and intelligent data caching to enable rapid storage and retrieval of entities, concepts, relationships, and metadata, along with advanced analytical functions and powerful search capabilities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If traditional database architectures are used to store entity information, then data storage capability is provided, but the system cannot scale effectively with large volumes of entities and complex relationships
Solution Approach 1:
The system segments entity information into multiple coreference units, each representing a distinct entity or concept. These coreference units are stored as separate records with their own attributes, allowing the database to scale by adding individual entity records without increasing overall system complexity. The segmentation enables independent management of each entity while maintaining relationships through shared attributes.
Solution Approach 2:
The patent introduces a new dimensional approach by organizing data around coreference units rather than traditional database tables. This dimensional change allows the system to represent complex relationships through attribute comparisons between coreference units, transforming the scaling problem from a database architecture issue to a data organization issue that can be handled more efficiently.
2Adaptability or versatility
If unstructured text data is processed to extract entity attributes, then information from diverse data sources can be integrated, but the quality and consistency of extracted attributes varies
Solution Approach 1:
The system applies local quality by extracting and storing specific attributes relevant to each coreference unit's context. Different attributes are extracted based on the specific unstructured data source and entity type, allowing high-quality extraction tailored to each data source's characteristics while maintaining consistency across the overall system through standardized attribute frameworks.
Solution Approach 2:
The coreference unit acts as an intermediary structure between raw unstructured text data and the final structured entity representation. This intermediary layer processes and standardizes attributes extracted from diverse data sources, filtering out inconsistencies and transforming varied data formats into consistent attribute representations that can be reliably compared and resolved.
3Measurement precision
If attributes from both structured and unstructured data are compared to resolve coreference, then resolution accuracy improves, but processing time and computational complexity increase
Solution Approach 1:
The system applies partial action by selectively comparing only the most relevant and informative attributes between coreference units, rather than comparing all possible attributes. This selective comparison approach maintains high resolution accuracy by focusing on discriminative attributes while significantly reducing processing time by excluding redundant or less informative attribute comparisons.
Solution Approach 2:
The patent changes the parameter of attribute comparison by transforming unstructured text attributes into structured formats that can be efficiently compared with existing structured data attributes. This parameter transformation enables faster comparison operations while maintaining the rich information content needed for accurate coreference resolution, reducing computational complexity through standardized attribute representations.
Data Source
AI summary
In some aspects, the present disclosure relates to coreference resolution. In one embodiment, a method includes obtaining unstructured text data including a plurality of references corresponding to entities, and determining, from the unstructured text data, attributes associated with the entities. The method also includes obtaining structured data including predefined attributes associated with the entities, and comparing attributes associated with a first coreference unit with attributes associated with a second coreference unit. The first coreference unit is a sub-entity representation having the attributes determined from the unstructured text data and the second coreference unit is a sub-entity representation having the predefined attributes. The method further includes determining, based on the comparison, whether the first coreference unit and the second coreference unit both correspond to the same entity.


