Supervised Machine Learning System for Automated Hypothesis Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current techniques for identifying quality hypotheses in scientific research are inefficient, often producing irrelevant results, requiring manual intervention, and introducing noise due to the difficulty in defining negative examples, which hampers the generation of accurate and significant research conclusions.

Innovation Solution

A supervised machine learning system (ISML) that learns from known relations to predict novel hypotheses by constructing an interaction network, using features like SimMeSH, JaccardArticleCoOccurence, and SumPub, and applying classifiers like Naïve Bayes, SVM, and logistic regression to rank the likelihood of interactions between information items.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If current techniques for identifying quality hypotheses are used, then manual intervention can be applied, but productivity is low and time consumption is high

Engineering Contradiction:
Improveautomated hypothesis identificationVSAvoidtime for identifying hypotheses
Core Design Contradiction:
Extent of automationVSLoss of time

Solution Approach 1:

The patent replaces manual research processes with a computational system that automatically identifies quality hypotheses. The system uses computer algorithms to process scientific literature, extract entities, and predict relationships, substituting the mechanical manual review process with an automated information processing system that operates at machine speed.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service hypothesis identification by automatically performing literature screening, entity extraction, and relationship prediction without requiring manual intervention. The computational framework independently processes scientific data, generates hypothesis predictions, and ranks them by quality, allowing the system to serve its own hypothesis generation function autonomously.

Inventive Principle:
Principle #25Self-service

2Measurement precision

If current techniques are used to identify quality hypotheses, then some hypotheses can be generated, but measurement precision is low due to noise and irrelevant results

Engineering Contradiction:
Improveprecision of hypothesis identificationVSAvoidnoise and irrelevant results
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent applies local quality by differentiating between various types of information in the literature and applying appropriate processing methods to each. The system identifies and extracts specific entity types (proteins, genes, chemicals) with their local characteristics, and uses type-specific features for relationship prediction, thereby improving precision by treating different information locally rather than uniformly.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system changes parameters by using multiple feature representations (co-occurrence frequency, semantic similarity, structural features) and adjusting the weighting of different evidence types based on their reliability. The hypothesis quality score dynamically adjusts based on the strength and consistency of multiple parameters, filtering out noise by requiring convergence of multiple independent evidence streams.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If supervised machine learning is applied to predict relationships, then productivity increases, but device complexity increases due to multiple classifiers and features

Engineering Contradiction:
Improvehypothesis generation rateVSAvoidsystem architecture complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the hypothesis identification task into distinct modular components: literature processing module, entity extraction module, feature computation module, and hypothesis prediction module. Each module handles a specific aspect of the workflow and can be independently optimized or replaced, managing complexity through functional segmentation while maintaining high overall productivity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system achieves universality by designing a multi-functional platform that can handle different types of scientific relationships (protein-protein interactions, gene-disease associations, drug-target bindings) using the same underlying machine learning framework. The unified architecture processes diverse data types and predicts various relationship types, reducing complexity by avoiding separate specialized systems for each relationship type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11120348B2System and methods for predicting probable relationships between items
Publication Date: 2021.09.14 UNIV OF MASSACHUSETTS
  • US11120348B2 patent drawing
  • US11120348B2 patent drawing
  • US11120348B2 patent drawing

AI summary

The present invention relates generally to identifying relationships between items. Certain embodiments of the present invention are configurable to identify the probability that a certain event will occur by identifying relationships between items. Certain embodiments of the present invention provide an improved supervised machine learning system.