Targeted Partial Corpus Re-enrichment via Surface Form Indexing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Re-enriching an entire corpus to take advantage of natural language processing (NLP) model enhancements is a computationally expensive and time-consuming process, especially when dealing with large datasets like hundreds, thousands, or millions of documents, as it requires processing the entire corpus rather than targeted portions.

Innovation Solution

Implementing a targeted partial re-enrichment method that identifies NLP model enhancements traceable to surface forms within the corpus, allowing for selective re-enrichment of specific passages or documents rather than the entire corpus, along with the option to preview updates before committing them to the database.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If the entire corpus is re-enriched to take advantage of NLP model enhancements, then the completeness and accuracy of annotations are improved, but the computational cost and time required increase significantly

Engineering Contradiction:
Improveannotation accuracyVSAvoidre-enrichment time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent divides the corpus into segments based on surface form occurrences. Instead of re-enriching the entire corpus, only the specific segments (documents and passages) containing the enhanced surface forms are identified and re-enriched. This segmentation approach maintains annotation accuracy for affected areas while avoiding unnecessary processing of unrelated portions.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by treating different parts of the corpus differently. Enhanced surface forms receive targeted re-enrichment with updated NLP models, while other parts of the corpus retain their existing annotations. This localized approach ensures high annotation accuracy where needed while minimizing overall computational expenditure.

Inventive Principle:
Principle #3Local quality

2Manufacturing precision

If the entire corpus is re-enriched to take advantage of NLP model enhancements, then the completeness and accuracy of annotations are improved, but the computational resources required increase significantly

Engineering Contradiction:
Improveannotation accuracyVSAvoidcomputational resources
Core Design Contradiction:
Manufacturing precisionVSUse of energy by moving object

Solution Approach 1:

The patent extracts only the necessary portions of the corpus that contain enhanced surface forms for re-enrichment. By extracting and processing only these specific segments rather than the entire corpus, the system achieves improved annotation accuracy for relevant areas while dramatically reducing computational resource consumption.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies partial action by performing re-enrichment on only the necessary subset of the corpus rather than the complete corpus. This partial processing approach provides sufficient annotation accuracy for enhanced surface forms while avoiding the excessive computational resources that would be required for full corpus re-enrichment.

Inventive Principle:
Principle #16Partial or excessive action

3Productivity

If targeted partial re-enrichment is performed on specific portions of the corpus, then the computational resources and time required are reduced, but the scope of updated annotations is limited

Engineering Contradiction:
Improvere-enrichment efficiencyVSAvoidnumber of updated annotations
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The patent performs preliminary action by first identifying all enhanced surface forms and their locations in the corpus before executing re-enrichment. This preliminary identification step enables the system to efficiently determine the exact scope of required updates, ensuring that the maximum necessary annotations are updated while maintaining high productivity through targeted processing.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11537660B2Targeted partial re-enrichment of a corpus based on NLP model enhancements
Publication Date: 2022.12.27 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11537660B2 patent drawing
  • US11537660B2 patent drawing
  • US11537660B2 patent drawing

AI summary

Techniques for targeted partial re-enrichment include determining that at least one natural language processing (NLP) request is associated with at least one surface form, the NLP request being for a corpus, a database comprising preexisting annotations associated with the corpus. An index query related to the at least one surface form is performed to generate index query results, the index query results including identification of portions of the corpus affected by the NLP request. A scope of the NLP request related to the database is determined based on the index query results, the scope including identification of impacted candidate annotations of the preexisting annotations affected by the NLP request. An NLP service is performed on the corpus according to the scope and the portions, thereby resulting in updates. The updates are committed to the database associated with the corpus.