Entity Information Extraction Using Ontology-Based String Sorting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for extracting entity information from digital media are cumbersome, time-consuming, and inefficient, especially when information is structured differently, leading to contextually irrelevant results.
Innovation Solution
A method and system that refine target data using a predefined syntax, generate strings, sort them by length, identify entity types based on an ontology, assign labels, and map them to a predefined signature to extract relevant entity information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If conventional searching techniques are used to extract entity information, then the process is simple, but the extraction becomes cumbersome and time-consuming when information is structured differently
Solution Approach 1:
The patent replaces manual mechanical sifting through webpages with an automated computational system that uses algorithms to refine target data, generate strings, and map them to predefined signatures, thereby eliminating time-consuming manual efforts while maintaining extraction accuracy
Solution Approach 2:
The system performs preliminary actions by pre-defining syntax rules and signatures before extraction begins. The algorithm refines target data according to predefined syntax patterns, allowing rapid automated extraction without manual intervention during the actual search process
2Productivity
If existing searching techniques are used, then the process is straightforward, but results become contextually irrelevant when data structure varies
Solution Approach 1:
The patent changes the parameters of data processing by transforming raw target data through algorithmic refinement according to predefined syntax, then generating multiple string representations. This parameter transformation allows the system to accurately identify entity types and map them to correct signatures regardless of the original data structure, maintaining high accuracy while preserving retrieval speed
Solution Approach 2:
The system segments the extraction process into distinct stages: refining target data, generating strings, sorting by length, identifying entity types, assigning labels, and mapping to signatures. This segmentation allows each stage to optimize for both speed and precision independently, with the overall process maintaining high productivity while improving measurement precision
3Adaptability or versatility
If manual extraction methods are used, then flexibility is maintained, but the process becomes complex and inefficient
Solution Approach 1:
The patent creates a universal extraction system that handles multiple data structures through a single algorithmic framework. The predefined syntax rules and signature mapping work universally across different data formats, providing adaptability to various structures without increasing process complexity, as the same algorithmic steps apply to all cases
Data Source
AI summary
Disclosed is a method and a system for extracting entity information from target data. The method comprises: providing the target data; refining the target data to obtain at least one base entity information having a plurality of base entity units using an algorithm, wherein the algorithm is based on a predefined syntax; generating a plurality of strings for each of the base entity information, wherein the plurality of strings comprises at least one base entity unit among the plurality of base entity units; sorting the plurality of strings in a decreasing order of length of the plurality of strings; identifying an entity type of the plurality of strings, based on an ontology, by processing the plurality of strings sequentially; assigning labels to the plurality of strings based on the entity type; and mapping the labelled plurality of strings to a predefined signature to obtain the entity information.

