Custom SARJ File Generation for Genomic Data Standardization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current genomic sequencing technologies generate diverse and duplicative data outputs, making it cumbersome to manage and analyze genomic information and sequence variant data from various files, which hinders efficient downstream genomic analyses.

Innovation Solution

A computer-implemented method generates a custom Sample Analysis Results JSON (SARJ) file by structuring and aggregating nucleic acid sequencing analysis files from multiple formats into a standardized format, using a schema to identify and store relevant data objects, and calculating checksums for authentication and validation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If multiple different data outputs from various sequencing assays are maintained, then comprehensive genomic information is preserved, but data management becomes clunky and duplicative

Engineering Contradiction:
Improvecomprehensive genomic informationVSAvoiddata management complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent merges multiple different sequencing assay data outputs into a single standardized file format. The system consolidates genomic information from various sequencing assays (whole genome sequencing, exome sequencing, targeted gene panels, RNA sequencing, etc.) into one unified file structure, eliminating the need to manage multiple separate files and reducing duplicative data management operations.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The standardized file format serves as a universal container that can accommodate data from multiple different sequencing assay types. The file structure is designed to be multi-functional, supporting whole genome sequencing data, exome sequencing data, targeted gene panel data, RNA sequencing data, and other genomic data types within a single standardized framework.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If various file formats are used to store different sequencing data types, then data-specific requirements are met, but downstream analysis becomes cumbersome

Engineering Contradiction:
Improvedata-specific requirementsVSAvoiddownstream analysis ease
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent applies homogeneity by standardizing the file format structure across all sequencing assay types. The standardized file format uses consistent data structures, field names, and organization principles for all genomic data, making downstream analysis uniform and simplifying bioinformatics processing regardless of the original sequencing assay type.

Inventive Principle:
Principle #33Homogeneity

Solution Approach 2:

The standardized file format is segmented into distinct sections for different data types (genomic data, transcriptomic data, clinical data, etc.), allowing each section to maintain its specific requirements while the overall structure provides uniformity. This segmentation enables easy extraction and processing of specific data types without handling the entire file complexity.

Inventive Principle:
Principle #1Segmentation

3Loss of information

If genomic data from multiple files is aggregated, then complete analysis information is obtained, but data processing time increases

Engineering Contradiction:
Improveanalysis information completenessVSAvoiddata processing time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The system performs preliminary aggregation and standardization of genomic data from multiple sequencing files into a single standardized format before downstream analysis. By consolidating data from whole genome sequencing, exome sequencing, targeted gene panels, and other assays into one pre-organized file structure, the system eliminates the need for repeated data collection and formatting operations during subsequent analysis steps, significantly reducing processing time.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20220084640A1Custom data files for personalized medicine
Publication Date: 2022.03.17 ILLUMINA INC
  • US20220084640A1 patent drawing
  • US20220084640A1 patent drawing
  • US20220084640A1 patent drawing

AI summary

Methods and systems are disclosed which can gather large data sets from nucleic acid sequencing technologies and devices, filter relevant genomic information and sequence variant information of biological samples from files of various formats, generate a custom data file having only relevant information in a standardized format, and provide the generated information to downstream analysis for personalized medicine use.