Custom Knowledgebase Construction with NLP Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing knowledgebases in genetics and genomics rely heavily on manual construction by subject matter experts, which is time-consuming and costly, and lack comprehensive automation in extracting and curating assertions and biological sequences for specific biological fields.

Innovation Solution

A semi-automated method involving natural language processing to extract assertions from publications, followed by manual editing by experts to construct custom knowledgebases and sequence datasets, allowing for the automatic extraction and association of biological sequences with assertions, enabling efficient construction and curation of knowledgebases for specific fields like antibiotic resistance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual construction of knowledgebases by subject matter experts is used, then the completeness and accuracy of the knowledgebase is improved, but the time consumption and cost increase significantly

Engineering Contradiction:
Improvecompleteness and accuracy of knowledgebaseVSAvoidtime consumption
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary automated extraction of assertions and biological sequences from publications before manual curation. This preliminary action prepares the data in advance, reducing the time experts need to spend on initial data collection while maintaining the ability to manually refine the results for accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary automated processing layer between raw publications and the final knowledgebase. This intermediary system extracts and structures data using natural language processing and sequence analysis, serving as a bridge that reduces the manual workload while preserving expert oversight for quality control.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If manual construction of knowledgebases by subject matter experts is used, then the quality of the knowledgebase is improved, but the cost increases significantly

Engineering Contradiction:
Improvequality of knowledgebaseVSAvoidcost
Core Design Contradiction:
ReliabilityVSEase of manufacture

Solution Approach 1:

The system performs preliminary automated extraction of assertions and biological sequences from publications before manual curation. This preliminary action prepares the data in advance, reducing the time experts need to spend on initial data collection while maintaining the ability to manually refine the results for accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary automated processing layer between raw publications and the final knowledgebase. This intermediary system extracts and structures data using natural language processing and sequence analysis, serving as a bridge that reduces the manual workload while preserving expert oversight for quality control.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If automated extraction of assertions from publications is used, then the productivity is improved, but the completeness and accuracy may be reduced

Engineering Contradiction:
Improveextraction speedVSAvoidcompleteness and accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system performs preliminary automated extraction of assertions and biological sequences from publications before manual curation. This preliminary action prepares the data in advance, reducing the time experts need to spend on initial data collection while maintaining the ability to manually refine the results for accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements a feedback mechanism where extracted assertions and sequences are reviewed and validated by subject matter experts. The experts can correct errors and provide feedback that improves the automated extraction process, creating a continuous improvement loop that enhances both speed and accuracy over time.

Inventive Principle:
Principle #23Feedback

4Reliability

If existing knowledgebases rely entirely on manual construction, then the accuracy is maintained, but the adaptability to specific biological fields is reduced

Engineering Contradiction:
ImproveaccuracyVSAvoidadaptability to specific biological fields
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system performs preliminary automated extraction of assertions and biological sequences from publications before manual curation. This preliminary action prepares the data in advance, reducing the time experts need to spend on initial data collection while maintaining the ability to manually refine the results for accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary automated processing layer between raw publications and the final knowledgebase. This intermediary system extracts and structures data using natural language processing and sequence analysis, serving as a bridge that reduces the manual workload while preserving expert oversight for quality control.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS9563741B2Constructing custom knowledgebases and sequence datasets with publications
Publication Date: 2017.02.07 BATTELLE MEMORIAL INST
  • US9563741B2 patent drawing
  • US9563741B2 patent drawing
  • US9563741B2 patent drawing

AI summary

Illustrative embodiments of custom knowledgebases and sequence datasets, as well as related methods, are disclosed. In one illustrative embodiment, one or more computer-readable media may comprise a custom knowledgebase and an associated sequence dataset. The custom knowledgebase may comprise a plurality of assertions that have been automatically extracted from a plurality of publications, where each of the plurality of assertions encodes a relationship between a subject and an object. The sequence dataset may comprise a plurality of called biological sequences, where each of the plurality of called biological sequences is associated with one or more of the plurality of assertions of the custom knowledgebase.