Colon Polyp Diagnosis Using Microbial Feature Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for diagnosing colon polyps using machine learning models face challenges due to noise in unprocessed samples and biases in bacterial metagenome analysis, leading to degraded performance in identifying causative factors of colorectal cancer.

Innovation Solution

A diagnostic apparatus and method that involves analyzing a mixture of a sample and a gut environment-like composition to extract microbial data, selecting microbe-related features using a predetermined algorithm, training a machine learning model with these features, and diagnosing colon polyps based on the output, focusing on specific microbial families and genera such as Oscillospirales, Burkholderiales, and Lactobacillales.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If bacterial metagenome analysis is performed without culturing samples, then the analysis process is simplified and faster, but the accuracy in identifying causative factors of colorectal cancer deteriorates due to large bias between samples

Engineering Contradiction:
Improveanalysis speedVSAvoidaccuracy in identifying causative factors
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent applies preliminary action by performing culturing of bacterial samples before metagenome analysis. This preliminary culturing step enriches the bacterial population and reduces bias between samples, thereby improving the accuracy of causative factor identification while maintaining a relatively efficient analysis process

Inventive Principle:
Principle #10Preliminary action

2Loss of time

If a machine learning model is trained using unprocessed samples of respective subjects as training data, then the data processing time is reduced, but the performance of the machine learning model significantly degrades due to noise in the training data

Engineering Contradiction:
Improvedata processing timeVSAvoidperformance of machine learning model
Core Design Contradiction:
Loss of timeVSReliability

Solution Approach 1:

The patent applies preliminary action by processing and preparing the training data in advance through culturing and standardized extraction procedures. This preliminary processing reduces noise in the training data, improving machine learning model performance while the standardized protocol keeps the overall processing time manageable

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary processing step between sample collection and machine learning training. This intermediary step involves culturing bacteria and extracting microbial data in a standardized manner, which acts as a mediator to reduce noise and improve the quality of training data for the machine learning model

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20230215570A1Method and apparatus for diagnosing colon plyp using machine learning model
Publication Date: 2023.07.06 HEM PHARM INC
  • US20230215570A1 patent drawing
  • US20230215570A1 patent drawing
  • US20230215570A1 patent drawing

AI summary

A method of diagnosing the presence or absence of colon polyps by using a machine learning model, which is performed by a diagnostic apparatus, includes: analyzing a mixture of a sample collected from a subject and a gut environment-like composition; extracting a plurality of microbial data based on an analysis result of the mixture; selecting a microbe-related feature to be used for the machine learning model from the plurality of microbial data based on a predetermined feature selection algorithm; training the machine learning model by using the microbe-related feature to predict the presence or absence of colon polyps for each of the microbial data; and diagnosing the presence or absence of colon polyps based on an output value of the machine learning model by inputting, into the trained machine learning model, the microbial data extracted based on the analysis result of the mixture of the sample collected from the subject and the gut environment-like composition, wherein the microbe-related feature includes the content of at least one kind of microbes selected from families belonging to the order Oscillospirales, the order Burkholderiales, the order Saccharimonadales, the order Lactobacillales, the order Bacteroidales, the order Clostridiales, the order Erysipelotrichales, the order Bacteroidales and the order Lachnospirales.