Application Document Generation Using Topic Clustering and Entity Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current systems face challenges in processing and analyzing unstructured information in application documents, leading to longer processing and response times, frustration, and potential misdirection in communications, due to the need for advanced programming and increased resource usage.

Innovation Solution

A method and system that cluster historical application documents by topic, select a representative document for each cluster, and engage users in a dialogue to populate new application documents, utilizing the representative document to request necessary information and ensure complete and accurate data entry.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Difficulty of detecting and measuring

If unstructured information is processed using advanced programming and automated systems, then information processing capability is improved, but processing time and resource usage increase

Engineering Contradiction:
Improveinformation processing capabilityVSAvoidprocessing time
Core Design Contradiction:
Difficulty of detecting and measuringVSLoss of time

Solution Approach 1:

The system performs preliminary actions by pre-processing historical application documents to extract entities and relationships before they are needed for processing new documents. This includes pre-clustering documents by topic and pre-identifying structured information patterns, so that when new documents arrive, the processing can be significantly accelerated by reusing these pre-computed structures.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates simplified copies of unstructured information by extracting key entities and relationships into structured formats. Instead of processing the entire unstructured text, the system generates condensed representations (copies) that retain the essential information while being much easier and faster to process by automated systems.

Inventive Principle:
Principle #26Copying

2Difficulty of detecting and measuring

If unstructured information is processed using advanced programming and automated systems, then information processing capability is improved, but resource expenses increase

Engineering Contradiction:
Improveinformation processing capabilityVSAvoidresource expenses
Core Design Contradiction:
Difficulty of detecting and measuringVSLoss of energy

Solution Approach 1:

The system creates simplified copies of unstructured information by extracting key entities and relationships into structured formats. Instead of processing the entire unstructured text, the system generates condensed representations (copies) that retain the essential information while being much easier and faster to process by automated systems.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system discards redundant and irrelevant portions of unstructured information while recovering and retaining only the essential structured data elements. By filtering out unnecessary information and keeping only the critical entities and relationships, the system reduces processing load and resource consumption while maintaining processing capability.

Inventive Principle:
Principle #34Discarding and recovering

3Device complexity

If generic application documents with few structured fields are used, then device complexity is reduced, but information completeness deteriorates

Engineering Contradiction:
Improveapplication document complexityVSAvoidinformation completeness
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

The system implements dynamic adaptation by automatically adjusting the structure of application documents based on the specific topic and content being processed. Rather than using a fixed generic template, the system dynamically generates or selects appropriate structured fields and formats tailored to each document's requirements, thereby maintaining both simplicity and completeness.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes parameters of application documents based on the input content, such as adjusting the number and type of structured fields, the level of detail required, and the specific format used. This allows the same base system to handle diverse document types effectively without requiring complex predefined templates for every possible scenario.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11423219B2Generation and population of new application document utilizing historical application documents
Publication Date: 2022.08.23 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11423219B2 patent drawing
  • US11423219B2 patent drawing
  • US11423219B2 patent drawing

AI summary

One embodiment provides a method, including: obtaining a plurality of previously submitted application documents, wherein each of the previously submitted application documents comprises information provided by a user who initiated a given previously submitted application document; clustering the plurality of previously submitted application documents into clusters of application documents based upon topics of the previously submitted application documents; selecting a representative application document; identifying entities contained within a given representative application document, wherein each of the entities corresponds to information to be entered into a new application document created from the given representative application document; and engaging in a dialogue with a user to create the new application document utilizing a similar representative application document to request information from the user, wherein the similar representative application document comprises a representative application document of a cluster having a topic similar to a topic of the new application document.