Funding Information Extraction Using Ensemble NER Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods are inefficient in automatically extracting funding information from scientific articles, as this requires distinguishing funding details from other acknowledgments and is embedded in various locations within text documents, necessitating advanced natural language processing and machine learning techniques.
Innovation Solution
A system and method utilizing a combination of natural language processing and machine learning models, including multiple annotators with different named-entity recognition models and an ensemble mechanism, to identify and extract funding information from text documents, optimizing accuracy and reducing computational time.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional methods are used to extract funding information from scientific articles, then the extraction process can be performed with simple tools, but the detection accuracy is low and false positives are high
Solution Approach 1:
The patent segments the funding information extraction task into multiple specialized components: paragraph classification (identifying paragraphs containing funding information), entity recognition (identifying specific funding entities within paragraphs), and information extraction (extracting structured funding details). This segmentation allows each component to be optimized independently, improving overall detection accuracy while managing system complexity through modular design.
Solution Approach 2:
The patent introduces intermediary components including trained machine learning classifiers that act as mediators between raw text and funding information extraction. These classifiers are trained on labeled datasets to recognize patterns indicative of funding information, serving as an intermediate processing layer that improves detection accuracy by filtering and preparing data before final extraction.
2Measurement precision
If advanced natural language processing and machine learning techniques are used, then detection accuracy improves, but computational time increases
Solution Approach 1:
The patent performs preliminary actions by pre-training machine learning classifiers on labeled datasets containing funding information examples. This pre-training establishes foundational knowledge that accelerates subsequent extraction tasks. Additionally, the system performs preliminary paragraph classification before detailed entity recognition, filtering out non-relevant paragraphs early to reduce computational burden on subsequent processing stages.
Solution Approach 2:
The patent divides the processing pipeline into sequential segments: first classifying paragraphs to identify those containing funding information, then performing entity recognition only on relevant paragraphs, and finally extracting structured information. This segmentation reduces overall computational time by avoiding intensive processing on irrelevant text portions while maintaining high detection accuracy through specialized processing at each stage.
3Measurement precision
If multiple annotators with different named-entity recognition models are used, then recall and precision are maximized, but device complexity increases
Solution Approach 1:
The patent merges multiple named-entity recognition models into a unified ensemble system that processes funding information extraction collectively. Instead of running separate independent annotators, the system combines their capabilities through an ensemble mechanism that aggregates results from multiple models, maximizing recall and precision while managing complexity through integrated processing architecture.
Solution Approach 2:
The patent implements feedback mechanisms where the ensemble of annotators learns from their collective performance. The system uses training data to teach the annotators what constitutes accurate funding information identification, and continuously refines their performance based on feedback from labeled examples. This feedback loop enables multiple annotators to work协同ively, improving recall and precision while the learned patterns reduce redundant complexity.
Data Source
AI summary
Systems and methods of extracting funding information from text are disclosed herein. The method includes receiving a text document, extracting paragraphs from the text document using a natural language processing model or a machine learning model, and classifying, using a machine learning classifier, the paragraphs as having funding information or not having funding information. The method further includes labeling, using a first annotator, potential entities within the paragraphs classified as having funding information, and labeling, using a second annotator, potential entities within the paragraphs classified as having funding information, where the first annotator implements a first named-entity recognition model and the second annotator implements a second named-entity recognition model that is different from the first named-entity recognition model. The method further includes extracting the potential entities from the paragraphs classified as having funding information and determining, using an ensemble mechanism, funding information from the potential entities.


