Funding Information Extraction Using Ensemble NER Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods are inefficient in automatically extracting funding information from scientific articles, as this requires distinguishing funding details from other acknowledgments and is embedded in various locations within text documents, necessitating advanced natural language processing and machine learning techniques.

Innovation Solution

A system and method utilizing a combination of natural language processing and machine learning models, including multiple annotators with different named-entity recognition models and an ensemble mechanism, to identify and extract funding information from text documents, optimizing accuracy and reducing computational time.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional methods are used to extract funding information from scientific articles, then the extraction process can be performed with simple tools, but the detection accuracy is low and false positives are high

Engineering Contradiction:
Improvedetection accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the funding information extraction task into multiple specialized components: paragraph classification (identifying paragraphs containing funding information), entity recognition (identifying specific funding entities within paragraphs), and information extraction (extracting structured funding details). This segmentation allows each component to be optimized independently, improving overall detection accuracy while managing system complexity through modular design.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces intermediary components including trained machine learning classifiers that act as mediators between raw text and funding information extraction. These classifiers are trained on labeled datasets to recognize patterns indicative of funding information, serving as an intermediate processing layer that improves detection accuracy by filtering and preparing data before final extraction.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If advanced natural language processing and machine learning techniques are used, then detection accuracy improves, but computational time increases

Engineering Contradiction:
Improvedetection accuracyVSAvoidcomputational time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary actions by pre-training machine learning classifiers on labeled datasets containing funding information examples. This pre-training establishes foundational knowledge that accelerates subsequent extraction tasks. Additionally, the system performs preliminary paragraph classification before detailed entity recognition, filtering out non-relevant paragraphs early to reduce computational burden on subsequent processing stages.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent divides the processing pipeline into sequential segments: first classifying paragraphs to identify those containing funding information, then performing entity recognition only on relevant paragraphs, and finally extracting structured information. This segmentation reduces overall computational time by avoiding intensive processing on irrelevant text portions while maintaining high detection accuracy through specialized processing at each stage.

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If multiple annotators with different named-entity recognition models are used, then recall and precision are maximized, but device complexity increases

Engineering Contradiction:
Improverecall and precisionVSAvoidnumber of annotators
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges multiple named-entity recognition models into a unified ensemble system that processes funding information extraction collectively. Instead of running separate independent annotators, the system combines their capabilities through an ensemble mechanism that aggregates results from multiple models, maximizing recall and precision while managing complexity through integrated processing architecture.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent implements feedback mechanisms where the ensemble of annotators learns from their collective performance. The system uses training data to teach the annotators what constitutes accurate funding information identification, and continuously refines their performance based on feedback from labeled examples. This feedback loop enables multiple annotators to work协同ively, improving recall and precision while the learned patterns reduce redundant complexity.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS10740560B2Systems and methods for extracting funder information from text
Publication Date: 2020.08.11 ELSEVIER INC
  • US10740560B2 patent drawing
  • US10740560B2 patent drawing
  • US10740560B2 patent drawing

AI summary

Systems and methods of extracting funding information from text are disclosed herein. The method includes receiving a text document, extracting paragraphs from the text document using a natural language processing model or a machine learning model, and classifying, using a machine learning classifier, the paragraphs as having funding information or not having funding information. The method further includes labeling, using a first annotator, potential entities within the paragraphs classified as having funding information, and labeling, using a second annotator, potential entities within the paragraphs classified as having funding information, where the first annotator implements a first named-entity recognition model and the second annotator implements a second named-entity recognition model that is different from the first named-entity recognition model. The method further includes extracting the potential entities from the paragraphs classified as having funding information and determining, using an ensemble mechanism, funding information from the potential entities.