NLP Entity Recognition via Database Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for identifying complex business entities from unstructured text data in supplier relationship management systems are manual and inefficient, failing to automatically link extracted data to structured data and map relationships effectively.

Innovation Solution

A system utilizing Natural Language Processing (NLP) techniques to identify and match text segments from documents against predefined entities, employing tagging techniques, semantic analysis, and consolidation to recognize complex entities, with a database storage unit and processor for entity recognition and data integration.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual processes are used to check and associate document identifiers with structured data, then flexibility and adaptability are maintained, but productivity is low and time consumption is high

Engineering Contradiction:
Improveprocessing speedVSAvoidautomation level
Core Design Contradiction:
ProductivityVSExtent of automation

Solution Approach 1:

The system enables self-service by automatically performing entity recognition, text segment identification, and data association without human intervention. The processor autonomously analyzes documents, identifies entities using NLP techniques, matches them against structured data, and populates databases, eliminating the need for manual agent intervention while maintaining high adaptability to different document types and business scenarios

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual process with an automated electronic system using Natural Language Processing techniques. The processor executes algorithms that perform entity recognition, text segmentation, and data matching, substituting human cognitive and manual operations with computational processes that achieve higher productivity and consistency

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Loss of time

If manual entity identification is performed, then measurement precision can be maintained through human judgment, but loss of time occurs due to manual processing

Engineering Contradiction:
Improveprocessing timeVSAvoidentity recognition accuracy
Core Design Contradiction:
Loss of timeVSMeasurement precision

Solution Approach 1:

The system implements continuous automated processing where the processor continuously analyzes documents, identifies entities, and associates them with structured data without interruption. This continuous automated action eliminates the time loss associated with manual processing while maintaining precision through consistent application of NLP algorithms and structured matching rules

Inventive Principle:
Principle #20Continuity of useful action

Solution Approach 2:

The system incorporates feedback mechanisms where the processor evaluates matching results, refines entity recognition based on structured data constraints, and adjusts text segment identification to improve accuracy. This feedback loop ensures high measurement precision in entity recognition while maintaining rapid automated processing speeds

Inventive Principle:
Principle #23Feedback

3Productivity

If automated NLP techniques are implemented, then productivity increases and time is reduced, but device complexity increases due to multiple processing components

Engineering Contradiction:
Improvedata processing throughputVSAvoidsystem architecture complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent implements a universal processor-based system that performs multiple functions including entity recognition, text segmentation, data matching, and database population through a single integrated platform. This multi-functional approach increases productivity across different document types and business scenarios while managing device complexity through standardized NLP techniques and unified architecture that can be applied universally across various processing tasks

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS8229883B2Graph based re-composition of document fragments for name entity recognition under exploitation of enterprise databases
Publication Date: 2012.07.24 SAP SE
  • US8229883B2 patent drawing
  • US8229883B2 patent drawing
  • US8229883B2 patent drawing

AI summary

Methods and systems are described that involve recognizing complex entities from text documents with the help of structured data and Natural Language Processing (NLP) techniques. In one embodiment, the method includes receiving a document as input from a set of documents, wherein the document contains text or unstructured data. The method also includes identifying a plurality of text segments from the document via a set of tagging techniques. Further, the method includes matching the identified plurality of text segments against attributes of a set of predefined entities. Lastly, a best matching predefined entity is selected for each text segment from the plurality of text segments.In one embodiment, the system includes a set of documents, each document containing text or unstructured data. The system also includes a database storage unit that stores a set of predefined entities, wherein each entity contains a set of attributes. Further, the system includes a processor to identify a plurality of text segments from a document via a set of tagging techniques and to match the identified plurality of text segments against the set of attributes.