Keyword Extraction Using Relationship Maps and NLP
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing keyword extraction techniques are insufficient for capturing important information in corporate documents like meeting notes and weekly reports, as they fail to effectively utilize statistical, natural language, and positional data, along with relationships between keywords.
Innovation Solution
A method that combines Natural Language Processing (NLP) information, frequency analysis, and co-occurrence analysis to extract keywords from documents by processing sentences into terms, forming clusters based on co-occurrence, evaluating connectivity, and calculating scores to identify top keywords.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional keyword extraction techniques are used, then the process is simple, but the extraction accuracy and ability to capture theme information is insufficient
Solution Approach 1:
The patent combines multiple extraction techniques including statistical analysis, natural language processing, and relationship map analysis into a unified keyword extraction system. This integration allows the system to leverage the strengths of each individual technique while achieving superior overall accuracy in capturing theme information from documents.
Solution Approach 2:
The invention creates a composite extraction approach that integrates diverse data types (frequency data, NLP data, positional data) and multiple analysis methods. This composite methodology similarly to how composite materials combine different substances to achieve enhanced properties, enables the system to overcome the limitations of any single extraction technique.
2Reliability
If multiple analysis techniques are combined to improve keyword extraction, then the extraction performance improves, but the processing complexity increases
Solution Approach 1:
The patent divides the keyword extraction process into distinct modular stages: statistical analysis stage, natural language processing stage, relationship map construction stage, and integration stage. Each stage processes specific types of data independently and produces intermediate results that are combined in subsequent stages, making the complex overall process more manageable and maintainable.
Solution Approach 2:
The relationship map serves as an intermediary structure that connects and integrates results from different analysis techniques. It stores and represents the relationships between keywords discovered through various methods, providing a unified framework for synthesizing information from statistical, NLP, and contextual analyses before final keyword selection.
3Loss of information
If relationship maps between keywords are utilized, then the thematic information capture improves, but the data processing time increases
Solution Approach 1:
The patent pre-processes the document data by constructing the relationship map and calculating co-occurrence statistics before the actual keyword extraction takes place. This preliminary preparation organizes the data in a structured format that facilitates faster and more accurate keyword identification during the extraction phase, reducing the time required for the critical analysis steps.
Data Source
AI summary
Disclosed herein is a method of extracting keywords from a document based on certain statistical, positional and natural language data, as well as relationship maps between the keywords. Under this method, document data are processed to obtain an NLP result for each sentence of the document, and based on the NLP result, words in the document are filtered and grouped into terms; a frequency analysis as well as a co-occurrence analysis are performed over the terms to output one or more keywords representing the document.


