Entity Information Extraction Using Ontology-Based String Sorting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for extracting entity information from digital media are cumbersome, time-consuming, and inefficient, especially when information is structured differently, leading to contextually irrelevant results.

Innovation Solution

A method and system that refine target data using a predefined syntax, generate strings, sort them by length, identify entity types based on an ontology, assign labels, and map them to a predefined signature to extract relevant entity information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If conventional searching techniques are used to extract entity information, then the process is simple, but the extraction becomes cumbersome and time-consuming when information is structured differently

Engineering Contradiction:
Improveease of entity information extractionVSAvoidtime required for manual sifting through webpages
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical sifting through webpages with an automated computational system that uses algorithms to refine target data, generate strings, and map them to predefined signatures, thereby eliminating time-consuming manual efforts while maintaining extraction accuracy

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system performs preliminary actions by pre-defining syntax rules and signatures before extraction begins. The algorithm refines target data according to predefined syntax patterns, allowing rapid automated extraction without manual intervention during the actual search process

Inventive Principle:
Principle #10Preliminary action

2Productivity

If existing searching techniques are used, then the process is straightforward, but results become contextually irrelevant when data structure varies

Engineering Contradiction:
Improvespeed of information retrievalVSAvoidaccuracy of entity information extraction
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent changes the parameters of data processing by transforming raw target data through algorithmic refinement according to predefined syntax, then generating multiple string representations. This parameter transformation allows the system to accurately identify entity types and map them to correct signatures regardless of the original data structure, maintaining high accuracy while preserving retrieval speed

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system segments the extraction process into distinct stages: refining target data, generating strings, sorting by length, identifying entity types, assigning labels, and mapping to signatures. This segmentation allows each stage to optimize for both speed and precision independently, with the overall process maintaining high productivity while improving measurement precision

Inventive Principle:
Principle #1Segmentation

3Adaptability or versatility

If manual extraction methods are used, then flexibility is maintained, but the process becomes complex and inefficient

Engineering Contradiction:
Improveadaptability to different data structuresVSAvoidcomplexity of extraction process
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent creates a universal extraction system that handles multiple data structures through a single algorithmic framework. The predefined syntax rules and signature mapping work universally across different data formats, providing adaptability to various structures without increasing process complexity, as the same algorithmic steps apply to all cases

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11270073B2Method and system for extracting entity information from target data
Publication Date: 2022.03.08 INNOPLEXUS AG
  • US11270073B2 patent drawing
  • US11270073B2 patent drawing

AI summary

Disclosed is a method and a system for extracting entity information from target data. The method comprises: providing the target data; refining the target data to obtain at least one base entity information having a plurality of base entity units using an algorithm, wherein the algorithm is based on a predefined syntax; generating a plurality of strings for each of the base entity information, wherein the plurality of strings comprises at least one base entity unit among the plurality of base entity units; sorting the plurality of strings in a decreasing order of length of the plurality of strings; identifying an entity type of the plurality of strings, based on an ontology, by processing the plurality of strings sequentially; assigning labels to the plurality of strings based on the entity type; and mapping the labelled plurality of strings to a predefined signature to obtain the entity information.