Text Entity Extraction via Information Mesh and Self-Service Catalogues
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional text analysis tools rely heavily on custom catalogues for extracting text entities, which are costly and error-prone to generate and maintain, and have a high reliance on these catalogues for quality and relevance.
Innovation Solution
A system that reduces reliance on custom catalogues by using a text entity extractor and an information mesh to identify and score text entities based on a built-in catalog, with a data structure that includes mesh entities, attributes, and relations, allowing for relevance determination and type assignment through a query process.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If custom catalogues are used to extract text entities, then the quality and relevance of extracted entities is improved, but the cost and complexity of generating and maintaining the catalogues increases
Solution Approach 1:
The system automatically generates and maintains custom catalogues by learning from the organization's existing data sources and text analysis results, eliminating the need for manual catalogue creation and updates. The catalogue evolves autonomously based on extracted entities and their contextual relationships.
Solution Approach 2:
The patent introduces an information mesh as an intermediary layer between the text analysis engine and custom catalogues. This mesh automatically learns entity relationships and structures from unstructured data, translating raw text patterns into structured catalogue formats without manual intervention.
2Reliability
If manual custom catalogues are generated, then the relevance of extracted entities to domain-specific terminology is improved, but the time and resources required for catalogue maintenance increase
Solution Approach 1:
The system performs automatic catalogue generation and maintenance through machine learning algorithms that continuously analyze text sources and update the information mesh. This self-service mechanism eliminates manual catalogue management while maintaining high relevance to domain-specific terminology through contextual learning.
Solution Approach 2:
The information mesh operates continuously to learn from new text sources and update entity relationships in real-time. This continuous learning process ensures the catalogue remains current and relevant without requiring periodic manual updates, thereby eliminating time loss associated with maintenance cycles.
3Ease of operation
If generic catalogues are used without custom configuration, then the ease of operation is improved, but the quality and relevance of extracted entities deteriorates
Solution Approach 1:
The system automatically adapts generic catalogues to domain-specific requirements by learning from the organization's data sources. The information mesh autonomously customizes entity extraction patterns without requiring manual configuration, thereby maintaining ease of operation while improving extraction quality through adaptive learning.
Solution Approach 2:
The catalogue structure transitions from static generic definitions to dynamic, context-aware patterns that adapt based on the organization's specific terminology and data characteristics. This dynamic adaptation allows the system to maintain operational simplicity while achieving domain-specific extraction accuracy.
Data Source
AI summary
A system includes a data structure comprising a plurality of mesh entities, the data structure associating each of the plurality of mesh entities with a respective name and a respective one or more attribute values, and associating each of the plurality of mesh entities with one or more relations to one or more other ones of the plurality of mesh entities. Some aspects include reception of a file comprising text, identification of text entities from the text, identification of first mesh entities from the plurality of mesh entities based on the identified text entities, determination, for each of the first mesh entities, of a name and one or more attribute values, and determination of a relevance associated with each identified text entity based on the determined name and one or more attribute values.


