Hybrid Machine-User Learning System for Scientific Data Annotation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for analyzing and extracting scientific data from textual formats, particularly in biological assays, are inefficient due to expert-specific jargon and lack of scalable solutions for large-scale data mining, hindering computer software's ability to perform data mining operations effectively.

Innovation Solution

A hybrid machine-user learning system that uses a computer with an algorithm to identify key words and phrases in textual documents, matches them with semantic definitions, and stores accurate semantic definition-key word pairs, enabling efficient annotation and storage of scientific data in a format like RDF triples.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual annotation by experts is used, then accuracy of scientific data annotation is improved, but time consumption and scalability are worsened

Engineering Contradiction:
Improveannotation accuracyVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent introduces an intermediary system consisting of automated NLP algorithms and machine learning models that act as mediators between the raw scientific text and the final annotated data. These intermediaries process the text to generate initial annotations, which are then refined through user verification, thereby maintaining high accuracy while significantly reducing the time required compared to purely manual annotation.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements feedback mechanisms where users verify and correct automated annotations, and these corrections are fed back into the machine learning models to improve future predictions. This iterative feedback loop allows the system to maintain expert-level accuracy while scaling to large datasets, as the automated system learns from user corrections to improve its performance over time.

Inventive Principle:
Principle #23Feedback

2Productivity

If automated NLP algorithms are used, then scalability and productivity are improved, but annotation accuracy is worsened

Engineering Contradiction:
Improveannotation speedVSAvoidannotation accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system implements feedback mechanisms where users verify and correct automated annotations, and these corrections are fed back into the machine learning models to improve future predictions. This iterative feedback loop allows the system to maintain expert-level accuracy while scaling to large datasets, as the automated system learns from user corrections to improve its performance over time.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent introduces an intermediary system consisting of automated NLP algorithms and machine learning models that act as mediators between the raw scientific text and the final annotated data. These intermediaries process the text to generate initial annotations, which are then refined through user verification, thereby maintaining high accuracy while significantly reducing the time required compared to purely manual annotation.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of operation

If expert-specific jargon is used in scientific texts, then communication effectiveness among scientists is improved, but computer software's ability to process and mine data is worsened

Engineering Contradiction:
Improvescientific communication effectivenessVSAvoiddata mining capability
Core Design Contradiction:
Ease of operationVSDifficulty of detecting and measuring

Solution Approach 1:

The patent introduces an intermediary layer of semantic annotation that bridges the gap between expert scientific jargon and computer processing capabilities. The system annotates text with standardized semantic terms and ontologies that translate complex scientific language into machine-understandable representations, enabling automated data mining while preserving the original scientific meaning and communication effectiveness.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system transforms the parameters of scientific text by mapping expert jargon to standardized semantic representations. This parameter transformation allows the same text to serve both human scientific communication and automated data processing purposes, as the semantic annotations create a bridge between the two domains without requiring changes to the original scientific writing style.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS9594743B2Hybrid machine-user learning system and process for identifying, accurately selecting and storing scientific data
Publication Date: 2017.03.14 COLLABORATIVE DRUG DISCOVERY
  • US9594743B2 patent drawing
  • US9594743B2 patent drawing
  • US9594743B2 patent drawing

AI summary

A process for identifying, accurately selecting, and storing scientific data that is present in textual formats. The process includes providing scientific data located in a text document and searching the text document using a computer and selecting a plurality of key words and phrases using an algorithm. The selected key words and phrases are matched with a plurality of semantic definitions and a plurality of semantic definition-key words and phrase pairs are created. The created plurality of semantic definition-key words and phrase pairs are displayed to a user via a computer user interface and the user selects which of the created plurality of semantic definition-key words and phrase pairs are accurate. The process also includes storing the selected and accurate semantic definition-key words and phrase pairs in computer memory.