PU Classification for Nonconforming Data in DNA Sequencing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Advanced sensing devices face challenges in accurately identifying and removing nonconforming data from measurement sets due to high noise levels and unclear object properties, particularly in next-generation DNA sequencers, where thermal and quantum noise interfere with minute measurements, leading to low accuracy and difficulty in applying conventional noise filtering methods.

Innovation Solution

A machine learning-based method using a Positive Unlabeled (PU) classification approach to identify and remove nonconforming data by learning from feature values such as waveform characteristics, including wave height, wavelength, kurtosis, and inertia moments, which enhances the accuracy of classification analysis in noisy measurement environments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional noise filtering methods are used, then the measurement process is simple, but the measurement accuracy is low due to high noise levels

Engineering Contradiction:
Improvemeasurement accuracyVSAvoidcomplexity of noise filtering method
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent transforms the noise filtering problem from signal processing domain to machine learning classification domain by changing the parameters and methods used. Instead of conventional filtering based on signal characteristics, it employs PU classification with multiple feature parameters (waveform shape, amplitude, duration) to identify and remove nonconforming data, thereby improving measurement accuracy in high-noise environments

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces conventional mechanical/electrical noise filtering methods with a machine learning-based classification system. The PU classification algorithm substitutes traditional signal processing mechanisms, using learned patterns from positive and unlabeled data to distinguish conforming from nonconforming measurements, achieving better accuracy without simple filtering

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If signal strength-based noise removal is used, then the method is easy to implement, but it fails when noise signal level is larger than target signal

Engineering Contradiction:
Improvereliability of noise removalVSAvoidease of noise removal method
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent changes the fundamental approach from signal strength comparison to multi-parameter classification. Instead of relying on amplitude differences, it uses PU classification with diverse features including waveform shape, duration, and temporal patterns, enabling reliable noise removal even when noise exceeds target signal strength

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces machine learning classifiers as an intermediary between raw measurement data and final results. The PU classification system acts as a mediator that learns to distinguish conforming from nonconforming data through training on positive examples and unlabeled data, providing reliable filtering without direct signal strength comparison

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If knowledge-based noise filtering is applied, then filtering can be targeted, but it cannot be applied when object properties are unclear

Engineering Contradiction:
Improveapplicability to unknown objectsVSAvoidprecision of noise filtering
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent performs preliminary learning action by training the PU classification system on positive example data and unlabeled data before actual measurement. This preliminary training enables the system to adapt to specific measurement contexts and object properties, making it versatile for unknown objects while maintaining precise, targeted noise filtering through learned patterns

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent enables the measurement system to serve itself by automatically learning and adapting to object properties through PU classification. The system extracts features directly from measurement data, trains classifiers autonomously, and applies learned patterns for noise filtering, eliminating the need for external knowledge input while maintaining high precision and adaptability

Inventive Principle:
Principle #25Self-service

Data Source

PatentEP3623793B1Identification classification analysis method and classification analysis device
Publication Date: 2022.09.07 AIPORE INC
  • EP3623793B1 patent drawingFigure 1A~1B
  • EP3623793B1 patent drawingFigure 2A~2B
  • EP3623793B1 patent drawingFigure 3

AI summary

The present invention provides an identification method by which appropriately identifies nonconforming data from a measurement data set, for example, contributes to improve the reliability of measurement results by the advanced sensing device, a classification analysis method which can perform the classification analysis with high accuracy for the measurement data, an identification device, a classification analysis device, a storage medium for identification and a storage medium for classification analysis. A feature value is obtained in advance which indicates the feature of waveform of pulse signal, and the feature value obtained is set as the learning data for machine learning. The nonconforming data identified with high accuracy by the classifier based on the PU classification technique are removed from the analyzed data, and by using the feature quantity obtained from said analyzed data as a variable, the classification analysis on the analyte is performed by executing the classification analysis program.