Custom SARJ File Generation for Genomic Data Standardization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current genomic sequencing technologies generate diverse and duplicative data outputs, making it cumbersome to manage and analyze genomic information and sequence variant data from various files, which hinders efficient downstream genomic analyses.
Innovation Solution
A computer-implemented method generates a custom Sample Analysis Results JSON (SARJ) file by structuring and aggregating nucleic acid sequencing analysis files from multiple formats into a standardized format, using a schema to identify and store relevant data objects, and calculating checksums for authentication and validation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If multiple different data outputs from various sequencing assays are maintained, then comprehensive genomic information is preserved, but data management becomes clunky and duplicative
Solution Approach 1:
The patent merges multiple different sequencing assay data outputs into a single standardized file format. The system consolidates genomic information from various sequencing assays (whole genome sequencing, exome sequencing, targeted gene panels, RNA sequencing, etc.) into one unified file structure, eliminating the need to manage multiple separate files and reducing duplicative data management operations.
Solution Approach 2:
The standardized file format serves as a universal container that can accommodate data from multiple different sequencing assay types. The file structure is designed to be multi-functional, supporting whole genome sequencing data, exome sequencing data, targeted gene panel data, RNA sequencing data, and other genomic data types within a single standardized framework.
2Adaptability or versatility
If various file formats are used to store different sequencing data types, then data-specific requirements are met, but downstream analysis becomes cumbersome
Solution Approach 1:
The patent applies homogeneity by standardizing the file format structure across all sequencing assay types. The standardized file format uses consistent data structures, field names, and organization principles for all genomic data, making downstream analysis uniform and simplifying bioinformatics processing regardless of the original sequencing assay type.
Solution Approach 2:
The standardized file format is segmented into distinct sections for different data types (genomic data, transcriptomic data, clinical data, etc.), allowing each section to maintain its specific requirements while the overall structure provides uniformity. This segmentation enables easy extraction and processing of specific data types without handling the entire file complexity.
3Loss of information
If genomic data from multiple files is aggregated, then complete analysis information is obtained, but data processing time increases
Solution Approach 1:
The system performs preliminary aggregation and standardization of genomic data from multiple sequencing files into a single standardized format before downstream analysis. By consolidating data from whole genome sequencing, exome sequencing, targeted gene panels, and other assays into one pre-organized file structure, the system eliminates the need for repeated data collection and formatting operations during subsequent analysis steps, significantly reducing processing time.
Data Source
AI summary
Methods and systems are disclosed which can gather large data sets from nucleic acid sequencing technologies and devices, filter relevant genomic information and sequence variant information of biological samples from files of various formats, generate a custom data file having only relevant information in a standardized format, and provide the generated information to downstream analysis for personalized medicine use.


