cfDNA End-Motif Analysis Using Multidimensional ML Features
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing techniques for analyzing cell-free DNA lack accuracy in determining properties and classifying pathologies, particularly due to the limitations in utilizing end motifs and sequence information.
Innovation Solution
The use of pre-end and post-end motifs, combined with machine learning techniques, to analyze multidimensional data structures for increased accuracy in determining properties and classifying pathologies, along with the use of stem-loop adapters to enhance sequencing efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing techniques for analyzing cell-free DNA are used, then the analysis process is simple, but the accuracy in determining properties and classifying pathologies is insufficient
Solution Approach 1:
The patent segments the DNA analysis into multiple dimensional features including pre-end motifs, post-end motifs, 5'-end motifs, and 3'-end motifs. Each segment is analyzed separately and then combined through machine learning models to achieve comprehensive and accurate pathology classification, resolving the contradiction between simple analysis and high accuracy.
Solution Approach 2:
The patent introduces multidimensional data structures that incorporate multiple types of end motifs and their positional information. By transforming one-dimensional sequence data into multidimensional feature spaces, the system achieves enhanced accuracy in pathology classification while maintaining manageable analysis complexity through structured data organization.
2Productivity
If traditional sequencing methods are used, then the sequencing process is straightforward, but the sequencing efficiency and accuracy are limited
Solution Approach 1:
The patent employs stem-loop adapters that perform preliminary actions during library preparation to enable efficient sequencing. The adapters are designed to facilitate strand-specific sequencing and improve the efficiency of detecting end motifs, thereby enhancing productivity without significantly increasing operational complexity.
3Measurement precision
If only basic end motif analysis is performed, then the analysis is simple and quick, but the accuracy in pathology classification is insufficient
Solution Approach 1:
The patent merges multiple types of end motif analyses (pre-end, post-end, 5'-end, and 3'-end motifs) into a unified machine learning framework. By combining these diverse features and processing them through integrated models, the system achieves high accuracy in pathology classification while optimizing analysis time through parallel processing and efficient algorithms.
Solution Approach 2:
The patent utilizes machine learning models that can dynamically adjust and optimize parameters such as motif lengths, positional weights, and feature combinations. This parameter optimization allows the system to achieve high classification accuracy while minimizing analysis time by identifying the most informative features and reducing computational overhead.
Data Source
AI summary
This disclosure provides techniques for analyzing end motifs, e.g., nucleotides in a reference genome outside the outmost coordinates of an aligned sequenced fragment, as well as machine learning techniques that use multidimensional data structures to achieve increased accuracy in determining a property (e.g., classification of a pathology or fractional concentration of clinically-relevant DNA) of a sample or of the subject from which a sample is obtained. Various end motifs are described and used for determining such properties. Various encodings of cfDNA molecules are also described, e.g., for use with molecule-level and sample-level models. 4-end sequencing techniques are described that reduce dimer artifacts. Cleavage profiles of 3′ ends around CpG sites are also used to detect pathologies.


