CNN Denoising ATAC-Seq Data Reducing NGS Read Requirements
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
ATAC-seq datasets require extensive processing and are costly due to the large number of NGS reads needed for accurate analysis, making it challenging to study isolated cell types, such as human cancer samples, with limited sample availability and high processing times.
Innovation Solution
A denoising process using a convolutional neural network (CNN) architecture that transforms suboptimal ATAC-seq datasets into equivalent quality datasets with up to five times fewer reads, reducing processing time and cost, and effectively denoises datasets from different cell or tissue types without requiring extensive sequence information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a large number of NGS reads are used to obtain accurate ATAC-seq data, then measurement precision is improved, but loss of time and loss of substance increase due to extensive processing requirements
Solution Approach 1:
The patent applies preliminary action by training a deep learning model in advance on high-quality ATAC-seq datasets. The trained model can then rapidly process suboptimal datasets without requiring extensive processing time, as the heavy computational work was performed during the preliminary training phase. This resolves the contradiction by shifting computational burden from the analysis stage to the training stage.
Solution Approach 2:
The patent introduces a deep learning model as an intermediary between suboptimal ATAC-seq data and high-quality results. This intermediary model, trained on high-quality data, acts as a bridge that transforms noisy, fast-to-obtain data into high-quality analytical results without requiring direct processing of large numbers of NGS reads, thus reducing processing time while maintaining accuracy.
2Measurement precision
If a large number of NGS reads are used to obtain accurate ATAC-seq data, then measurement precision is improved, but loss of substance increases due to sample requirements
Solution Approach 1:
The deep learning model is pre-trained on high-quality ATAC-seq datasets generated from sufficient samples. This preliminary training captures the characteristics of high-quality data, enabling the model to later enhance suboptimal datasets that require fewer physical samples. This resolves the contradiction by decoupling sample requirements from the analysis phase.
Solution Approach 2:
The patent uses copying by creating a computational model that replicates the characteristics of high-quality ATAC-seq data. Instead of requiring physical copies of high-quality samples, the model captures and reproduces the essential features of high-quality data, allowing suboptimal datasets with fewer samples to be transformed into high-quality results through computational copying rather than physical replication.
3Loss of time
If suboptimal ATAC-seq datasets with fewer reads are used, then loss of time and loss of substance are reduced, but measurement precision deteriorates
Solution Approach 1:
The deep learning model serves as an intermediary that processes suboptimal datasets and enhances their quality. The model takes noisy, low-precision data as input and outputs enhanced data with improved measurement precision, effectively bridging the gap between fast-to-obtain suboptimal data and high-quality results without requiring extensive processing of raw reads.
Solution Approach 2:
The patent applies parameter changes by transforming the quality parameters of ATAC-seq data through the deep learning model. The model changes key parameters such as signal-to-noise ratio, peak detection accuracy, and overall data quality, converting suboptimal datasets with fewer reads into high-quality datasets that meet analytical standards.
4Loss of substance
If suboptimal ATAC-seq datasets with fewer reads are used, then loss of substance is reduced, but measurement precision deteriorates
Solution Approach 1:
The deep learning model creates a computational copy of high-quality data characteristics. By training on high-quality datasets, the model learns to reproduce essential features and patterns, allowing it to enhance suboptimal datasets that require fewer physical samples while maintaining measurement precision through computational rather than physical means.
Solution Approach 2:
The model transforms the quality parameters of datasets generated from minimal samples. It changes critical parameters such as signal-to-noise ratio, peak calling accuracy, and data reliability, enabling high-quality analysis from suboptimal datasets that require reduced sample input.
Data Source
AI summary
The present invention provides methods, systems, computer program products that use deep learning with neural networks to denoise ATAC-seq datasets. The methods, systems, and programs provide for increased efficiency, accuracy, and speed in identifying genomic sites of chromatin accessibility in a wide range of tissue and cell types.


