Automated Enterprise Document Harvesting with Feedback-Based Retraining

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for generating purpose-specific documents, such as cybersecurity incident reports, are time-consuming and prone to overlooking crucial content, leading to potential financial penalties and legal issues due to incomplete or inaccurate document completion.

Innovation Solution

A system utilizing algorithms to automatically extract and populate data fields in documents from various enterprise sources, with subsequent editing and retraining based on accuracy, ensuring comprehensive and accurate content inclusion.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If algorithms automatically extract information to populate data fields, then document creation speed is improved, but accuracy of information extraction may deteriorate

Engineering Contradiction:
Improvedocument creation speedVSAvoidinformation extraction accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system monitors editing of extracted information and uses this feedback to determine algorithm accuracy. When accuracy is determined to be below thresholds, the algorithm is automatically retrained using the edited information, creating a continuous improvement loop that maintains both speed and accuracy

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The algorithm performs preliminary information extraction and population of data fields before human review. This preliminary action enables rapid document creation while allowing subsequent accuracy verification through editing monitoring and selective retraining

Inventive Principle:
Principle #10Preliminary action

2Loss of information

If algorithms extract information from multiple enterprise sources, then completeness of document content is improved, but risk of overlooking crucial content increases

Engineering Contradiction:
Improvecontent completenessVSAvoidcontent accuracy
Core Design Contradiction:
Loss of informationVSReliability

Solution Approach 1:

The system monitors edits made to extracted information from multiple sources and uses this feedback to identify when crucial content was overlooked or incorrectly extracted. This feedback mechanism triggers selective retraining to improve the algorithm's ability to reliably extract important information from diverse sources

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The algorithm is designed to extract information from multiple types of enterprise sources including emails, social media conversations, direct messages, online discussions, document repositories, and code repositories. This multi-functional capability ensures comprehensive content collection while the feedback mechanism maintains reliability

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If algorithm retraining is performed based on editing feedback, then extraction accuracy is improved, but system complexity increases

Engineering Contradiction:
Improveextraction accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

Editing feedback is collected and stored as training data before retraining is initiated. This preliminary preparation of training data simplifies the retraining process by having ready-to-use examples that directly improve extraction accuracy without requiring complex data collection procedures

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system performs automatic retraining using the collected editing feedback without requiring external intervention. The algorithm self-improves by using its own performance gaps (identified through editing monitoring) as training data, reducing the need for complex external training infrastructure

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12443788B2Automated document harvesting and regenerating by crowdsourcing in enterprise social networks
Publication Date: 2025.10.14 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US12443788B2 patent drawing
  • US12443788B2 patent drawing
  • US12443788B2 patent drawing

AI summary

Various systems and methods are presented regarding automatically generating a document for a particular topic, and further utilizing an algorithm to automatically identify information in a resource that pertains to the topic. The document can be a purpose-specific document, created by a template configured to combine data fields to build the document. Each data field can be subject-matter specific with an associated algorithm. The algorithm can be trained to recognize and extract content that potentially meets the subject matter of the data field. The data field can be populated with the content. Subsequent editing of the content populating the data field can be monitored to determine an effectiveness of the algorithm to identify data of interest in the data repository. In the event of low success, the algorithm can be retrained. Further, the document can be regenerated if an algorithm is retrained, new information added to the repository, etc.