CNN Keyword Extraction from Unstructured Documents

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for extracting relevant keywords from unstructured documents, such as resumes, are time-intensive and inefficient, especially when dealing with the need to manually classify and update job descriptions with latest skill sets, as they often rely on ineffective supervised learning techniques.

Innovation Solution

A method and device utilizing a Convolution Neural Network (CNN) model and keyword repository to split documents into keyword samples, determine relevancy scores, and classify keywords as relevant or non-relevant, thereby automating the extraction of relevant keywords.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual keyword extraction and classification is performed, then accuracy and relevance of keywords can be ensured, but time consumption and labor intensity increase significantly

Engineering Contradiction:
Improvekeyword relevance accuracyVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system enables self-service by allowing the CNN model to automatically extract and classify keywords from unstructured documents without requiring manual intervention. The model processes documents independently, generating relevant keywords and job descriptions autonomously, thereby eliminating the time-consuming manual extraction process while maintaining accuracy through the trained neural network.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual process of keyword extraction and classification with an automated neural network system. The CNN model substitutes human labor by processing unstructured text data, identifying keywords, and generating job descriptions through computational algorithms, thus transforming a manual mechanical process into an automated digital system.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If supervised learning techniques are used for keyword extraction, then training can be guided by labeled data, but effectiveness on unstructured documents remains limited

Engineering Contradiction:
Improvetraining guidanceVSAvoidextraction effectiveness
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent applies parameter changes by transitioning from traditional supervised learning parameters to deep learning parameters. The CNN model uses multiple layers of parameters (convolutional filters, pooling layers, fully connected layers) to process unstructured document data, enabling the system to automatically learn patterns and extract keywords effectively without relying on pre-labeled training data for every possible scenario.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If unstructured documents are processed manually, then flexibility and adaptability are maintained, but processing speed and efficiency decrease

Engineering Contradiction:
Improvedocument processing flexibilityVSAvoidprocessing speed
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The system enables self-service by allowing the CNN model to automatically extract and classify keywords from unstructured documents without requiring manual intervention. The model processes documents independently, generating relevant keywords and job descriptions autonomously, thereby eliminating the time-consuming manual extraction process while maintaining accuracy through the trained neural network.

Inventive Principle:
Principle #25Self-service

4Device complexity

If conventional keyword extraction methods are used, then simplicity of implementation is maintained, but extraction accuracy and relevance fall short

Engineering Contradiction:
Improveimplementation simplicityVSAvoidkeyword extraction accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent applies parameter changes by transitioning from traditional supervised learning parameters to deep learning parameters. The CNN model uses multiple layers of parameters (convolutional filters, pooling layers, fully connected layers) to process unstructured document data, enabling the system to automatically learn patterns and extract keywords effectively without relying on pre-labeled training data for every possible scenario.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11416532B2Method and device for identifying relevant keywords from documents
Publication Date: 2022.08.16 WIPRO LTD
  • US11416532B2 patent drawing
  • US11416532B2 patent drawing
  • US11416532B2 patent drawing

AI summary

A method of identifying relevant keywords from a document is disclosed. The method includes splitting text of the document into a plurality of keyword samples, such that each of the plurality of keyword samples comprises a predefined number of keywords extracted in a sequence. Further, each pair of adjacent keyword samples in the plurality of samples includes a plurality of common words. The method further includes determining a relevancy score for each of the plurality of keyword samples based on at least one of a trained Convolution Neural Network (CNN) model and a keyword repository. The method further includes classifying keywords from each of the plurality of keyword samples as relevant keywords or non-relevant keywords based on the relevancy score determined for each of the plurality of keyword samples.