Context-Based Genomic Read Compression for Biomarker Signatures
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods lack efficient ways to compress nucleic acid sequence read data for immune-oncology biomarkers, necessitating improved systems for reducing memory requirements and detecting enrichment or loss of signatures in tumor microenvironments for predicting responses to immune-oncology treatments.
Innovation Solution
A method and system that organize targeted genes into categories, compress read data to form reduced-data values, and compare these values to baselines to determine enrichment or loss of functional contexts, using a processor and memory for efficient data storage and analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If nucleic acid sequence read data is stored in full detail, then measurement precision is maintained, but memory requirements increase significantly
Solution Approach 1:
The patent extracts only the essential information from full nucleic acid sequence read data by counting reads mapping to specific genes and organizing them into functional categories. This extraction process retains the biologically relevant signal while discarding redundant sequence detail, thereby reducing memory requirements while preserving detection accuracy for immune-oncology biomarkers
Solution Approach 2:
Instead of storing complete sequence data and then analyzing it, the patent inverts the approach by directly counting and categorizing reads during the sequencing process. This inversion transforms the data structure from detailed sequences to aggregated read counts organized by functional gene categories, achieving compression while maintaining analytical utility
2Measurement precision
If comprehensive gene expression data is analyzed, then prediction accuracy improves, but data processing complexity increases
Solution Approach 1:
The patent segments the genome into predefined functional categories of genes (e.g., immune response, cell cycle, metabolism) and aggregates read counts within each category. This segmentation transforms complex individual gene expression data into manageable functional profiles, reducing processing complexity while preserving the biological context needed for accurate treatment response prediction
Solution Approach 2:
The patent merges individual gene expression measurements into category-level summaries by aggregating read counts across multiple genes within each functional category. This merging process consolidates redundant information and presents a simplified yet comprehensive view of tumor biology, making data processing more efficient while maintaining predictive accuracy
Data Source
AI summary
The method includes compressing numbers of reads data for targeted genes of a gene expression assay performed on a test sample. The targeted genes are organized into categories. Each category represents a functional context associated with the targeted genes in that category. The numbers of reads corresponding to targeted genes each category is compressed to form a compressed value for the category. The compressed value is compared to a baseline value for the category to determine an enrichment or a loss of a signature corresponding to the functional context of the category. The method may include analyzing information from multiple assays performed on the test sample, assigning a score value to each assay result and predicting a response to immune-oncology treatment based on the assigned scores.


