Systems and methods for correcting noise and systematic variation in sequencing data
Long-read sequencing data is used to correct noise and systematic variation in genotyping, addressing the challenges of genotyping repeat-rich genes by accurately identifying protein-truncating variants and improving the reliability of genotyping results.
Patent Information
- Application Number
- JP2025541734
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-01-26
- Filing Date
- 2024-01-25
- Publication Date
- 2026-02-18
AI Technical Summary
Existing genotyping methods, particularly those using next-generation sequencing (NGS), struggle with accurately genotyping genes containing large repeat structures due to short read lengths and increased sequencing error rates, leading to erroneous assemblies and unreliable results.
A method utilizing long-read sequencing data, such as PacBio®, to correct for noise and systematic variation by length-sorting reads, removing artifacts, and identifying protein-truncating variants, enabling accurate genotyping of repeat-rich genes.
The method provides accurate genotyping of high-repeat genes by transforming and correcting sequencing data, reducing noise and systematic variation, and improving the reliability of genotyping results.
Smart Images

Figure 2026505717000001 
Figure 2026505717000002 
Figure 2026505717000003
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 441,308, filed January 26, 2023, the entire contents of which are incorporated herein by reference.
[0002] The present disclosure relates generally to correcting noise and systematic variation in sequencing data, and more particularly to systems and methods for correcting noise and systematic variation in sequencing data caused by repeat regions. [Background technology]
[0003] Genotyping allows researchers to determine an individual's DNA sequence at specific loci in their DNA. Advances in gene sequencing technology now allow researchers to investigate common variants in individual genes, thereby enabling the investigation of disease-causing variants across populations.
[0004] Currently, genotyping is performed by traditional genotyping methods using predefined single nucleotide polymorphisms or by direct sequencing.
[0005] Traditional genotyping typically requires predefined single nucleotide polymorphism ("SNP") markers, which must be discovered and validated in advance. These markers are often population-specific and therefore must be detected for each population sample before the genotype can be assessed. SNPs are typically detected by hybridization or individual SNP-specific PCR-based assays. Thus, establishing the necessary SNP markers across an entire database can be inefficient due to the significant effort required to define SNPs prior to genotyping.
[0006] Genotyping-by-sequencing (GBS) technology allows for the detection of a wider range of polymorphisms than PCR-based assays, including insertions or deletions ("indels") in addition to SNPs. The main advantage of sequencing over genotyping arrays is that arrays are fixed in content, whereas sequencing can discover novel variants. GBS allows researchers to detect greater variation among individuals in a dataset, thereby enabling more accurate assessment of individuals within a population. GBS technology eliminates the need for prior polymorphism discovery and validation required for traditional genotyping. Therefore, GBS significantly reduces the effort required for genotyping and can be used for any polymorphic species and any segregating population.
[0007] However, GBS has several drawbacks. First, it requires double-stranded adapters for site-specific ligation to amplification primers. The methods required to ligate adapters to target sequences prior to sequencing require strict control of the template-to-adapter concentration ratio during adapter ligation. As a result, precisely quantified, high-quality input DNA is required as starting material (see, for example, Elshire et al.). Second, these methods comprehensively analyze hundreds of thousands of sites, necessitating a large number of sequencing reads to ensure sufficient coverage at each site in each sample. Therefore, although GBS is less labor-intensive than conventional sequencing, it still requires considerable effort to implement.
[0008] The widespread availability of next-generation sequencing technologies ("NGS") provides tremendous opportunities for detailed genotyping of individuals and the discovery of novel genetic variations.
[0009] However, despite advances in NGS, certain genes remain difficult to genotype, particularly those containing large repeat structures in the nucleic acid code, due to the inherent limitations of NGS technology.
[0010] Genotyping requires sequencing data generated by a sequencing platform as input. The quality of the sequencing data varies significantly across platforms and technologies, which directly impacts the quality of genotyping. Therefore, a basic understanding of sequencing is necessary.
[0011] Currently, the standard platform for sequencing utilizes next-generation sequencing (NGS), typified by Illumina® sequencing-by-synthesis technology. NGS technology generates "short read" data, i.e., sequencing data for small fragments of DNA, typically approximately 150 nucleotides in length. To perform sequencing, NGS technology divides the DNA sequence in a sample into numerous small fragments, a step called "library preparation." Once the library is prepared, the DNA fragments are replicated into millions of copies in a process called "amplification." Finally, the amplified fragments are recorded by a sequencer through a biochemical process, generating a digital record of each nucleotide in each fragment. Each digital record of a fragment is called a "sequencing read" or simply a "read."
[0012] Once sequencing reads are generated, software is required to "assemble" the sample's DNA sequence by mapping the reads to each other and reconstructing the larger sequence. Assembling large repeat structures is challenging due to the short read length of NGS data. NGS sequencing reads average 150 nucleotides in length. Therefore, gene assembly using NGS data requires extensive read depth to generate sufficient overlap to pinpoint the precise location of each read within the genome. However, for genes with large repeat regions, extensive read depth may not overcome the limitations of short-read sequencing and may even make the process more difficult. This is because short reads covering the repeat region map to multiple locations within the repeat. Therefore, pinpointing the exact location of a given read within the sequence is difficult, often resulting in erroneous assemblies that do not represent the sample being sequenced.
[0013] Long-read data, such as those generated by Pacific Biosciences (Menlo Park, CA) sequencing (hereafter referred to as "PacBio"), offers a solution to the assembly challenges posed by NGS data. Reads from long-read data methods such as PacBio® exceed the average length of repeat sequences, allowing for proper anchoring of assemblies to precise locations within a sample. See, e.g., U.S. Patent No. 9,165,109, the entire contents of which are incorporated by reference.
[0014] While certain aspects of the methods described herein focus on the alignment of reads generated by a PacBio sequencing instrument, such as the PacBio® RS instrument, the characteristics of sequence reads generated by the PacBio® RS instrument are likely common to other single-molecule sequencing methods. The key differences in the read characteristics of single-molecule sequencing reads are that they are one to two orders of magnitude longer than those generated by NGS, and that sequencing errors occur at a higher overall error rate than NGS sequencing and tend to be biased toward insertions and deletions. In contrast to NGS methods, signals observed in single-molecule sequencing do not suffer from the problem of phasing or prephasing, in which signals are observed from different positions within an amplified template in the same cycle (see Kao, et al. (2009) Genome Research 10:1884-1895). Because the fluorescent signal does not decay throughout the length of the read, longer reads can be obtained and there is no positional bias in sequencing errors. The read length on the PacBio® RS platform is limited by the processivity of the polymerase and can be approximated by an exponential process. Sequencing is based on real-time measurements of the fluorescent analog signal during the incorporation reaction. Very occasionally, base incorporation proceeds too quickly for detection, resulting in deletions in the determined base sequence. In other cases, a base is present in the observation region but is not incorporated into the growing strand, emitting a signal that results in an insertion in the read. Incorrect base incorporation is rare, and the incidence of substitutions is very low. As a result, the error types in single-molecule sequencing are biased toward insertions and deletions. While raw data accuracy is lower than that of NGS sequencing, consensus accuracy improves rapidly with increasing coverage due to the absence of positional bias in sequencing errors (see Travers, et al. (2010) Nuc. Ac. Res. 38(15):e159).
[0015] In the context of genotyping, sequence assembly of long-read data presents several data quality issues due to increased sequencing error rates, bad reads, and other technical artifacts. [Prior art documents] [Patent documents]
[0016] [Patent Document 1] U.S. Patent No. 9,165,109 [Non-patent literature]
[0017] [Non-Patent Document 1] Elshire et al. [Non-patent document 2] Kao, et al. (2009) Genome Research 10:1884-1895 [Non-patent document 3] Travers,et al.(2010)Nuc.Ac.Res.38(15):e159 Summary of the Invention
[0018] Therefore, the disclosed methods and systems aim to provide accurate genotyping of high-repeat genes using long-read sequence data.
[0019] Aspects of the present invention relate to systems and methods for generating genotypes for one or more samples using long-read sequencing data. The genotyping techniques transform and correct for systematic noise in sequencing data due to repeat regions of repeat-rich genes.
[0020] In one embodiment, the present disclosure provides a method for genotyping high-repeat genes of individuals across a dataset using long-read sequence data.
[0021] In some embodiments, the parameters in the method are related to noise and other sources of systematic variation. These parameters can then be used to convert a dataset of long-read sequencing data into a dataset of normalized allele representations that accurately represent the genotypes of the samples in the dataset. In various embodiments, systematic variation can arise from instrument artifacts, chemical artifacts, process artifacts, operational artifacts, sequencing artifacts, and assay artifacts. Correcting for noise and systematic variation allows for sample genotyping. These and other features of the present disclosure are disclosed herein.
[0022] In some embodiments of the present disclosure, genotyping of an individual is performed by: length-sorting reads based on the expected length of the alleles and removing noise; determining the repeat structure of the remaining reads and removing noise based on the structure of known variants; sorting the remaining reads by allele; identifying protein-truncating variants; and identifying novel alleles with unknown copy numbers.
[0023] In some embodiments, denoising the input data may be achieved by creating a first dataset of reads by classifying the reads based on length and removing reads with lengths that do not match known alleles, thereby excluding those reads from the dataset.
[0024] In some embodiments, denoising the first dataset to generate the second dataset can be achieved by mapping repeat sequence references to each long-read sequence read and generating a structural readout. A non-limiting example is Minimap2: Pairwise Alignment of Nucleotide Sequences (Li H., Bioinformatics, Vol. 34, Issue 18, 15 September 2018, Pages 3094-3100).
[0025] In some embodiments, denoising the first sequencing data to create the second dataset may include removing reads from the dataset that are determined to have an improper structure as an artifact.
[0026] In some embodiments, the reads of the second dataset can be transformed into a third dataset by sorting by allele.
[0027] In some embodiments, the creation of the third dataset may include correction for PCR imbalances and include only reads that adequately represent the individual's alleles.
[0028] In some embodiments, the reads of the second dataset may be converted into a third dataset, where the third dataset includes a classification based on copy number homogeny for each sample.
[0029] In some embodiments, converting the second dataset to a third dataset involves classifying individual samples into either copy number heterogeneous ("CN-HET") or copy number homogeneous ("CN-HOM") individuals based on the size ratio between the most dominant read group and the next most dominant read group. If the size ratio between the most dominant read groups is larger than expected, the read is classified as CN-HOM. If the size ratio closely matches expected, the read is classified as CN-HET. If the length of the first dominant read group is larger than the length of the second dominant read group, the sample is classified based on whether the length of the first read group is larger or smaller than the length of the third read group.
[0030] In some embodiments, the third dataset is further transformed into a fourth dataset by generating a consensus sequence for each allele in the third dataset.
[0031] In some embodiments, converting the third dataset to the fourth dataset can include generating a consensus sequence across each allele in the third dataset. Generating a consensus sequence can be accomplished using software such as SPOA, as described in "Fast and accurate de novo genome assembly from long uncorrected reads" by Vaser R, Sovic I, Nagarajan N, and Sikic M (Genome Res. 2017 May;27(5):737-746. Epub 2017 Jan 18.).
[0032] In some embodiments, the consensus sequence of the fourth dataset can be aligned to a customized reference to identify all nucleic acid variations.
[0033] In some embodiments, the consensus sequences of the fourth dataset can be further converted into a fifth dataset by translating the consensus sequences into corresponding amino acid codes.
[0034] In some embodiments, the fourth data set is converted to the corresponding amino acid code via the longest open reading frame ("ORF"). Such conversion can be performed, for example, using a software program such as the getorf utility included in the EMBOSS package (described in EMBOSS: The European Molecular Biology Open Software Suite (2000), Rice, P., Longden, I. and Bleasby, A., Trends in Genetics 16(6), pp. 276-277).
[0035] Certain embodiments may be utilized for genotyping the filaggrin gene, which is the most frequently mutated gene associated with atopic dermatitis (AD).
[0036] Filaggrin is a repeat-rich gene with 10-12 perfect repeat sequences within its third exon. The length of each repeat sequence ranges from 972-975 nucleotides. The complex structure of filaggrin makes sequencing using conventional short-read techniques highly unreliable and inaccurate. Long-read sequencing is also unreliable using conventional calling algorithms. The current approach described herein provides a novel method for achieving accurate genotyping of long-read sequencing data for genes with complex structures (e.g., filaggrin). [Brief explanation of the drawings]
[0037] [Figure 1] FIG. 1 illustrates a block diagram of an exemplary computer system according to one embodiment of the present disclosure. [Figure 2] FIG. 2 shows a block diagram of an exemplary genotyping computing device that can be used with the exemplary computer system shown in FIG. 1. [Figure 3] 3 shows a block diagram of an artificial intelligence (AI) / deep learning (DL) module that may be used with the genotyping computing device shown in FIG. 2. [Figure 4] FIG. 1 is a schematic diagram showing the structure of the filaggrin (FLG) gene. [Figure 5] FIG. 1 is a schematic representation of PCR bias observed after haplotype amplification. [Figure 6] FIG. 1 is a schematic diagram showing PCR amplification of haplotypes. [Figure 7] FIG. 1 is a flowchart for mapping repeat sequences obtained after PacBio sequencing using reads measuring two alleles. [Figure 8] This flowchart shows the workflow from identifying reads measuring two alleles, to correlating them with baseline atopic dermatitis (AD) severity, and performing validation analyses such as Hardy-Weinberg (HWE) for common alleles. [Figure 9] FIG. 1 is a schematic diagram showing an approach to exploiting the structure of repeat sequences to remove sequencing noise from true variants, according to one embodiment of the present disclosure. [Figure 10] FIG. 1 shows the generation of two consensus sequences, followed by identification of the longest open reading frame (ORF) and translation of the ORF into an amino acid sequence, according to one embodiment of the present disclosure. [Figure 11] According to one embodiment of the present disclosure, the difference in ORF (open reading frame) length of the detected haplotypes is plotted against its observed frequency compared to the CN=10 allele (GRCh38 reference allele). [Figure 12] FIG. 1 shows a global alignment of generated read lengths with reference to the third exon of the Genome Reference Consortium Human Genome Build 38 ("GRCh38"), according to one embodiment of the present disclosure. [Figure 13] A shows the order of alignment between a series of reads, and B shows the size of these reads, comparing their size to the actual read length. [Figure 14A] FIG. 1 is a schematic diagram illustrating the structure of first and second frequent groups in an individual sample, according to one embodiment of the present disclosure. [Figure 14B] 1 is a histogram showing the distribution of QC-passed reads for approximately 3,000 individuals after excluding the third and subsequent frequent groups, according to one embodiment of the present disclosure. [Figure 15] FIG. 1 is a schematic diagram showing two exemplary scenarios for the analysis of imbalance rates and allele lengths of the most frequent read populations when an individual is determined to be CN-HOM, according to one embodiment of the present disclosure. [Figure 16]FIG. 1 is a schematic diagram showing how two different alleles identified within a read group indicate a heterozygous genotype, according to one embodiment of the present disclosure. [Figure 17] 1 is a table showing the protein truncating mutations (PTV or pLoF) and dosage of any PTV (pLoF load) observed within this study, according to one embodiment of the present disclosure. [Figure 18] FIG. 1 is a schematic diagram showing an association test performed between variables representing FLG genotypes and baseline disease phenotypes of atopic dermatitis (AD), according to one embodiment of the present disclosure. [Figure 19A] 1 is a box plot showing baseline clinical phenotypes of body surface area in relation to functional repeat copy number, according to one embodiment of the present disclosure. [Figure 19B] 1 is a box plot showing baseline clinical phenotype of Eczema Area and Severity Index (EASI) in relation to functional repeat copy number, according to one embodiment of the present disclosure. [Figure 19C] 1 is a box plot showing baseline clinical phenotype of SCORing Atopic Dermatitis (SCORAD) scores in relation to functional repeat copy number, according to one embodiment of the present disclosure. [Figure 20] 1 is a graph showing the relationship between age at diagnosis of atopic dermatitis ("AD") and copy number, according to one embodiment of the present disclosure. [Figure 21A] 1 is a graph showing the relationship between the history of asthma and the number of alleles having PTV of the FLG gene in AD patients. [Figure 21B] 1 is a graph showing the relationship between a history of food allergy and the number of alleles with PTVs of the FLG gene in AD patients, according to one embodiment of the present disclosure. [Figure 22A] 1 is a graph showing the relationship between the history of AD and the number of alleles with PTV of the FLG gene in patients with asthma. [Figure 22B] 1 is a graph showing the relationship between a history of food allergy and the number of alleles having PTVs of the FLG gene, according to one embodiment of the present disclosure. [Figure 23] 1 is a graph showing known and newly validated PacBio PTVs in exon 3 according to one embodiment of the present disclosure. [Figure 24] 1 is a graph showing predicted copy numbers and new copy numbers according to one embodiment of the present disclosure. [Figure 25] FIG. 1 is a graph showing an example of artifact-introducing amplification of repeat sequences during PacBio sequencing, visualized by aligning an unusually long (19 kb) read to the FLG reference sequence (GRCh38), according to one embodiment of the present disclosure. [Figure 26] 1 is a flowchart of a process for converting an input dataset into genotypes of individual samples according to one embodiment of the present disclosure. [Figure 27] 2 is a block diagram of a server computing device that can be used with the computer system shown in FIG. 1 according to one embodiment of the present disclosure. [Figure 28] 2 is a block diagram of a user computing device that may be used with the computer system shown in FIG. 1 according to one embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0038] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, specific methods and materials are described herein.
[0039] Next-generation sequencing (NGS) has brought about significant advances in genotyping technologies, contributing to a better understanding of disease-causing mutations in the human genome.
[0040] However, despite the advances in NGS, it has become clear that certain genes remain difficult to genotype, particularly those containing large repeat structures in their DNA sequences, due to the inherent limitations of NGS technology.
[0041] NGS technology is a short-read platform. The average read length of sequencing data generated by an NGS platform is 150 nucleotides. Due to the short read length nature of NGS data, genome or gene assembly using NGS data requires extensive coverage of the genome or gene of interest to ensure sufficient overlap between reads and identify the correct genomic location of each read.
[0042] However, for genes with large repeat regions, this process is complicated because reads covering the repeat region map to multiple regions of the gene sequence, making it difficult to identify the correct sequence location of the read, often leading to erroneous assemblies that are not representative of the sequenced sample.
[0043] Long-read data, such as those obtained from PacBio sequencing, provide a means to resolve assembly issues in NGS data. Because the reads from long-read data methods such as PacBio® exceed the average length of repeat sequences, assembly can be properly anchored to the correct location in the genome.
[0044] However, several data quality issues have been reported in read assembly for long-read sequencing due to increased sequencing error rates, bad reads, and other technical artifacts.
[0045] The disclosed method is capable of accounting for errors contained in long-read sequencing reads, enabling accurate assembly of the genome.
[0046] In one aspect, the present disclosure provides a method for genotyping an individual using long-read sequencing data, such as that provided by PacBio®.
[0047] In another aspect, the present disclosure provides a method for denoising sequencing data provided as input, allowing for accurate assembly of the DNA sequence of an individual sample.
[0048] In another aspect, the present disclosure provides a method for determining an amino acid sequence of an individual based on an assembly of the individual's DNA sequence.
[0049] In another aspect, the disclosure provides a method for identifying protein-truncating variants by analyzing the amino acid sequences of individuals across a population based on comparison to a reference amino acid sequence, where individuals in the population who have a gene with an amino acid sequence that is shorter in length than the reference amino acid sequence are determined to have a protein-truncating variant of the gene.
[0050] The term "a" should be understood to mean "at least one," and the terms "about" and "approximately" should be understood to allow for normal variation understood by one of ordinary skill in the art. Furthermore, when numerical ranges are given, they are intended to be inclusive of the endpoints. As used herein, the terms "include," "includes," and "including" are used in an open-ended sense and are understood to be synonymous with "comprise," "comprises," and "comprising," respectively.
[0051] The term "sequencing read" or "read" should be understood to mean the DNA sequence of a DNA fragment generated by a sequencer.
[0052] The term "read length" should be understood to mean the length of the sequencing read in nucleotides.
[0053] The term "noise" should be understood to mean statistical anomalies in a dataset that are not representative of the sample. In the context of DNA sequencing and genotyping, noise can include sequencing artifacts such as insertions or deletions, or errors in the balance of reads for each allele resulting from the amplification reaction.
[0054] The term "amplified" should be understood to mean a DNA sequence that has been replicated for sequencing using artificial methods. Samples are amplified before DNA sequencing to increase the number of sequences that can be sequenced. This also increases the number of reads present in the sequencing data generated by the sequencer.
[0055] The term "repeat region" should be understood to mean a region of a DNA sequence that contains a pattern of nucleotides that appears multiple times within that region. Repeat sequences can consist of two to several thousand consecutively repeated nucleotides, and are estimated to account for approximately 30% of the human genome.
[0056] The term "DNA segment" should be understood to mean a length of genomic DNA that is identifiable by sequencing using a DNA probe.
[0057] The term "genome" refers to the DNA or RNA genetic material present in a cell and should be understood to include the chromosomal / nuclear genome of the cell, as well as any mitochondrial and / or plasmid genomes. The nuclear genome may include protein-coding genes, non-coding genes, non-coding DNA, or other functional regions containing RNA, and junk DNA or RNA, if present.
[0058] The term "protein-coding gene" should be understood to mean the sequence of a nucleic acid molecule (RNA or DNA molecule) that comprises a nucleotide sequence that encodes a protein. The coding sequence can further comprise initiation and termination signals operably linked to regulatory elements, including a promoter and polyadenylation signal, capable of directing expression in the cells of an individual or mammal to which the nucleic acid is administered.
[0059] The term "protein-altering variant" should be understood to mean a mutation that affects the protein structure. The term "protein-truncating variant" should be understood to mean a genetic variant that is predicted to result in a truncation of the amino acid chain after translation. The amino acid sequence may be truncated by the addition of a stop codon or by a frameshift or missense mutation.
[0060] The term "PacBio" should be understood to mean the PacBio sequencing method, which is the trademarked sequencing method of Pacific Biosciences.
[0061] The term "long read" should be understood to mean a sequencing read that is greater than 500 nucleotides in length. Long read data is typically generated by third- or next-generation sequencing platforms, with an average read length greater than 10,000 nucleotides (or 10 kb), compared to NGS sequencing, which has an average read length of 150 nucleotides.
[0062] The term "genotyping" should be understood to mean determining an individual's DNA sequence at specific loci.
[0063] The term "phasing" should be understood to mean inferring haplotypes from genotype data.
[0064] The term "artifact" should be understood to mean errors introduced into the representation of data by the equipment used.
[0065] The term "sequence artifact" should be understood to mean variation caused by factors other than biological processes.
[0066] The term "structural artifact" should be understood to mean a sequence whose copy number is unknown compared to the exon or subunit structure of a reference gene. The unknown copy number may include an unknown exon or subunit structure (e.g., when there are multiple exons corresponding to each domain). The reference gene may include an exon structure derived from one or more known alleles of the gene.
[0067] The term "detecting read structure by alignment" should be understood to mean mapping the exon structure of a read to a reference exon structure of a reference gene (e.g., mapping exons representing repeats).
[0068] The term "variant allele frequency" should be understood to mean the frequency with which sequence reads matching a particular DNA variant are observed.
[0069] The term "barcoded amplicon" should be understood to be a unique identifier that is ligated to each individual sample in order to index it during sample preparation prior to sequencing, thereby preserving the identity of the sample.
[0070] The term "CN-HOM" should be understood to mean an individual that contains a pair of alleles that have the same length and the same number of repeats or identical repeat structure in the sequencing data for the gene being genotyped.
[0071] The term "CN-HET" should be understood to mean an individual that contains a pair of alleles in the sequencing data for the gene being genotyped that are different in length and number of repeats, or have the same number of repeats but different repeat structure.
[0072] The term "historical data" should be understood to mean training data, which includes dataset X and allele classification information for samples within the dataset. For example, in some embodiments, a machine learning model is trained to associate datasets with allele classifications. More specifically, the model finds a relationship between input (dataset) and output (allele classification). Allele classification can include, for example, assigning different alleles based on heterozygous SNPs and / or indels in the case of CN-HOM, or assigning different alleles based on CN structure in the case of CN-HET.
[0073] Sequence identity can be calculated using algorithms such as the Needleman-Wunsch algorithm for global alignments (Needleman and Wunsch 1970, J. Mol. Biol. 48:443-453) or the Smith-Waterman algorithm for local alignments (Smith and Waterman 1981, J. Mol. Biol. 147:195-197). Another suitable algorithm is described in Dufresne et al., 2002, Nature Biotechnology (vol. 20, pp. 1269-71), and this algorithm is used in the GenePAST software (GQ Life Sciences, Inc., Boston, Massachusetts).
[0074] Read assembly can be calculated using software such as SPOA, as described in "Fast and accurate de novo genome assembly from long uncorrected reads" by Vaser R, Sovic I, Nagarajan N, and Sikic M (Genome Res. 2017 May;27(5):737-746 Epub 2017 Jan 18.).
[0075] Software such as BEDTools, as described in "BEDTools: a flexible suite of utilities for comparing genomic features" by Aaron R. Quinlan and Ira M. Hall (Bioinformatics, Volume 26, Issue 6, March 15, 2010, Pages 841-842), can be used to analyze sequencing data, align, and calculate genomic features.
[0076] Therefore, the disclosed methods and systems aim to provide accurate genotyping of high-repeat genes using long-read sequencing data.
[0077] Disclosed are systems and methods for generating genotyping calls for one or more samples using long-read sequencing data, which transform and correct for noise and / or systematic variation in sequencing data caused by repeat regions in repeat-rich genes.
[0078] In one embodiment of the present disclosure, a method is provided for genotyping repeat-rich genes of individuals in a dataset using long-read sequencing data.
[0079] In some embodiments, the parameters of the method relate to the removal of noise and systematic variation.
[0080] In some embodiments, the present disclosure genotypes individuals by sorting reads by length and removing noise based on the expected length of the allele; identifying the repeat structure of the remaining reads and removing noise based on the structure of known variants; sorting the remaining reads by allele; and identifying protein-truncating mutations.
[0081] In some embodiments, the input sequencing data is sorted by length to perform noise removal based on the expected length of the gene. The read length is compared with the length of the known allele of the gene of interest. Read lengths that do not match the length of the known allele are considered to contain sequencing errors, such as insertions or deletions introduced by the sequence, and noise removal is performed to create a first dataset from which reads containing sequence errors are removed. Sequencing artifacts may refer to artifacts such as repeat regions in sequencing reads caused by errors in the sequencing reaction, amplification reaction, and / or instrument errors, resulting in reads containing DNA sequence structures that do not represent the individual's actual genome. Reads whose lengths match those of the known alleles are used to create the first dataset.
[0082] In some embodiments, the noise removal process of input data is realized by classifying reads based on length, and removing the length that does not match the known allele length of the gene, thereby creating a first data set of reads by removing the reads from the data set.If a read is outside the known allele length range of the gene that is to be genotyped, the read is considered to be inconsistent with the known allele length of the gene.Therefore, these reads are determined to contain artifacts, and are removed from the data set as noise.
[0083] In some embodiments, generating a second dataset by denoising the first dataset can be performed by mapping a repeat sequence reference to each long-read sequencing read and creating a map of the structure of each read. The reference sequence is mapped using software to identify the position of each repeat in the DNA sequence within the read and the number of repeats within the read. Reads that show a structure that does not match known alleles of the gene are determined to contain artifacts and are removed from the dataset as noise.
[0084] In some embodiments, mapping of reference repeat sequences to long-read sequencing reads can be performed using software such as BEDTools, which is described in Aaron R. Quinlan, Ira M. Hall, "BEDTools: a flexible suite of utilities for comparing genomic features," Bioinformatics, Volume 26, Issue 6, March 15, 2010, Pages 841-842. BEDTools includes a variety of sequence analysis tools, such as intersect, merge, count, complement, and shuffle tools.
[0085] In some embodiments, denoising the first sequencing dataset to create the second dataset can include culling reads determined to have an inappropriate structure as an artifact from the dataset. The structure of the reads can be first determined by mapping repeat sequences to the reads using a software tool such as BEDTools. The number and position of the repeats are used as the structure of the read. If a read contains a repeat combination or sequence that does not exist in known alleles of the gene, the read is culled as having an "inappropriate structure".
[0086] In some embodiments, the reads of the second dataset may be converted into a third dataset of allelically classified reads by correcting for PCR imbalance, determining individual copy numbers, and removing groups of reads determined to be noise.
[0087] PCR is often used to prepare DNA samples for sequencing to ensure sufficient DNA sequence for assembly. The ideal result after PCR is a dataset containing reads in the same ratio as those present in the sample before amplification. However, the mechanism of PCR creates an artificial imbalance, where short reads are overproduced. This imbalance in PCR is an unavoidable consequence of the known fact that PCR amplifies short sequences at a faster rate than long sequences. Therefore, PCR results will always contain an artificially increased amount of short sequences compared to the amount of long sequences.
[0088] PCR imbalance is particularly problematic when there are heterogeneous alleles: the shorter alleles are replicated at a higher amplification rate, creating an artificial imbalance in the sequence ratio between the alleles.
[0089] In genes with repeat regions, individuals' alleles often differ in the number of copies of the repeat region. Individuals with homogeneous copy number ("CN-HOM") have the same number of repeats in both alleles. Individuals with heterogeneous copy number ("CN-HET") have alleles with different numbers of repeats.
[0090] In CN-HET individuals, PCR imbalance causes the proportion of readings corresponding to the allele with lower copy number in their sequencing data to be imbalanced.This is because this allele is amplified more frequently during PCR.Therefore, instead of the approximately equal distribution of alleles in sequencing data, there is often an imbalance in which the readings corresponding to short alleles are significantly more than the readings of long alleles.
[0091] In some embodiments, the computer system can utilize machine learning and / or artificial intelligence to acquire and integrate sequencing data and determine genotypes. The computer system can include appropriate data storage capabilities, such as cloud storage, to access and / or store any of the above data. In this manner, the computer system can access and analyze historical and / or current data (e.g., real-time or near-real-time data). In some embodiments, the computer system can include or be implemented via one or more local or remote processors, servers, transceivers, memory units, mobile devices, wearable devices, smart watches, smart contact lenses, smart glasses, augmented reality (AR) glasses, virtual reality (VR) headsets, mixed reality (MR) or extended reality (XR) glasses or headsets, voice or chat bots, AI bots (including generative AI bots), and / or other electronic or electrical components, which may be connected to each other via wired or wireless communication.
[0092] In an exemplary embodiment, the computer system includes at least one computing device configured to perform the functions generally described and / or attributed herein as being performed by the computer system as a whole.
[0093] In particular, in an exemplary embodiment, the computing device may be capable of communicating with one or more computing devices associated with a user, which may include personal mobile computing devices such as smartphones, tablets, and the like.
[0094] The computing device may receive portions of the sequencing data and / or any other type of data from other computing devices, such as sequencing devices. Additionally or alternatively, the computing device may access portions of the data from one or more databases or other storage devices. The computing device may be configured to aggregate, combine, integrate, analyze, compare, and / or otherwise process the data to determine genotypes, as described in more detail herein.
[0095] The computing device may store the received, acquired, and / or accessed data in one or more databases, and may store the determined genotypes and / or other generated data in the one or more databases. The databases may be any suitable storage location and, in some embodiments, may include cloud storage such that multiple computing devices (e.g., multiple user analysis computing devices, third-party computing devices, etc.) can access the database. The databases may be integrated into the computing device or may be remote from the computing device.
[0096] The technical effects of the systems and processes described herein may be achieved by performing at least one of the following steps: (i) receiving an input set of genetic sequencing data; (2) measuring at least one repeat structure of at least one sequence read of the genetic sequencing data and removing sequence reads having an aberrant repeat structure to generate a first dataset, the first dataset excluding sequence reads having sequence artifacts; (3) filtering the sequence reads of the first dataset to generate a second dataset, the second dataset excluding sequence reads having structural artifacts. (4) applying input data including the sequence reads of the second dataset to the model to generate a third dataset, the third dataset including sequence reads classified by allele; (5) generating a consensus sequence for each allele in the third dataset to generate a fourth dataset of consensus reads corresponding to all alleles in the third dataset; (6) translating each allele in the fourth dataset to generate a fifth dataset, the fifth dataset including amino acid sequences corresponding to each allele; and (7) determining a protein mutation variant status based on the amino acid sequences of the fifth dataset. The protein mutation variant status may include whether the variant affects protein structure and / or the effect of the variant on the protein structure. For example, the protein mutation variant status may indicate whether the sequence includes a protein-truncating variant (PTV or pLoF) and / or the observed dosage of PTV (pLoF burden).
[0097] At least one technical challenge addressed by the systems and methods disclosed herein may include: (i) the time-consuming, laborious, and costly process of determining genotypes; (ii) limitations in the data utilized in determining genotypes; (iii) limitations in qualitative and quantitative analysis of genotypes in the drug discovery candidate compound selection process; and (iv) noise and systematic variation in sequence read data significantly impact the accuracy and functionality of genotype analysis.
[0098] The resulting technical effects may include, for example: (i) more accurate genotyping; (ii) the ability to receive data from multiple different data sources and convert the data into a format for further analysis; (iii) reduced processing time and costs associated with genotyping; (iv) a significant reduction in noise and systematic variation in sequence read data; and (v) more accurate disease diagnosis.
[0099] An exemplary computer system for genotyping is disclosed herein. For example, FIG. 1 shows a schematic diagram of an exemplary computer system 100. The computer system 100 is configured to determine genotypes. In one embodiment, the computer system 100 may include and / or facilitate communication between a genotyping computing device 110 and one or more user computing devices 130 (which may also be referred to as "mobile devices"), and / or between the genotyping computing device 110 and one or more third-party devices 140 and / or a genotyping server 150.
[0100] The genotyping computing device 110 may be implemented as a server computing device with artificial intelligence and deep learning capabilities. Alternatively, the genotyping computing device 110 (and / or the user computing device 130) may be implemented as any device capable of interconnecting with the Internet, including a mobile computing device or "mobile device" such as a smartphone, a "phablet," or other web-enabled appliance or mobile device (such as one or more local or remote processors, servers, transceivers, memory units, mobile devices, wearables, smart watches, smart contact lenses, smart glasses, augmented reality (AR) glasses, virtual reality (VR) headsets, mixed reality (MR) or extended reality (XR) glasses or headsets, voice chatbots, AI bots (including generative AI bots), and / or other electronic or electrical components, which may be capable of communicating with each other via wired or wireless connections).
[0101] The genotyping computing device 110 can communicate with one or more user computing devices 130, third party devices 140, and / or genotyping servers 150, such as via wireless communication or transmitting and receiving data over one or more radio frequency links or wireless communication channels. In an exemplary embodiment, the components of the computer system 100 may be communicatively connected to the Internet through a number of interfaces, which may include, but are not limited to, the Internet, a local area network (LAN), a wide area network (WAN), or a network such as an integrated services digital network (ISDN), a dial-up connection, a digital subscriber line (DSL), a cellular communication connection (e.g., a 3G, 4G, 5G, etc. connection), a cable modem, and a BLUETOOTH connection.
[0102] Computer system 100 also includes one or more database(s) 120 that contain information about a variety of matters. For example, database 120 may contain information such as sequencing data (e.g., long-read sequencing data) and / or other information used, received, and / or generated by computer system 100 and / or its components, including information described herein. In one embodiment, database 120 may include a cloud storage facility, such that information stored therein is securely stored while remaining accessible by one or more components of computer system 100, such as genotyping computing device 110, user computing device 130, and / or genotyping server 150. In some embodiments, database 120 may be stored on genotyping computing device 110. In other embodiments, database 120 may be stored remotely from genotyping computing device 110 and may be decentralized.
[0103] In some embodiments, the user computing device 130 may include a computer with a web browser or software application that allows the user computing device 130 to access the functionality of the genotyping computing device 110 over the Internet or a direct connection, such as a cellular network connection. The user computing device 130 may be any device capable of accessing the Internet, including, but not limited to, a desktop computer, a mobile device (e.g., laptop, personal digital assistant (PDA), mobile phone, smartphone, tablet, phablet, netbook, notebook, smartwatch or bracelet, smart glasses, wearable electronics, pager, virtual reality headset, augmented reality glasses, voice or chat bot, wearable, etc.), or other web-enabled device.
[0104] The user computing device 130 may be used to access, for example, via a user interface 132, a data management application 112 managed by the genotyping computing device 110, when the data management application 112 executes on the user computing device 130. A user may use the data management application 112 to provide input to the genotyping computing device 110, view genotypes determined by the genotyping computing device 110, and perform other operations, including those described elsewhere herein.
[0105] The third party device 140 may be a computing device associated with an external data source. The genotyping computing device 110 may request, receive, and / or otherwise access data from the third party device 140. The third party device 140 may be any device capable of connecting to the internet, including a server computing device, a mobile computing device or "mobile device" such as a smartphone, or other web-enabled appliance or mobile device.
[0106] Examples of computing devices for user analysis are disclosed herein. For example, FIG. 2 illustrates a genotyping computing device 110 (see FIG. 1) according to one embodiment. In some embodiments, the genotyping computing device 110 may include a processor 202, a memory 204 (which may include something similar to the database 120 shown in FIG. 1), a communication interface 206, and a storage interface 208. The processor 202 is configured to execute instructions that may be stored in the memory 204. The processor 202 may include one or more processing units (e.g., a multi-core configuration) and be configured to execute multiple modules.
[0107] In some embodiments, processor 202 is capable of executing an artificial intelligence / deep learning (AI / DL) module 210, a genotyping module 212, and a module 214 that maintains the functionality of data management application 112 (shown in FIG. 1 ). Modules 210, 212, and 214 may include dedicated instruction sets and / or coprocessors. Database 120 and / or memory 204 may store any data and / or instructions necessary for modules 210, 212, 214 to function as described herein. In an exemplary embodiment, database 120 may store sequencing data (e.g., long-read sequencing data) 220, a first dataset 222 (e.g., a dataset that excludes reads containing sequence artifacts), a second dataset 224 (e.g., a dataset that excludes sequences containing structural artifacts), a third dataset 226 (e.g., a dataset that includes sequence reads sorted by allele), a fourth dataset 228 (e.g., a dataset that includes consensus reads for each allele), a fifth dataset 230 (e.g., a dataset that includes amino acid sequences corresponding to each allele), and / or other information used, received, and / or generated by genotyping computing device 110 (e.g., a sixth dataset of protein truncation variants).
[0108] The AI / DL module 210 may perform artificial intelligence and / or deep learning functions on behalf of the genotyping module 212. In particular, the AI / DL module 210 may include any rules, algorithms, training datasets / programs, and / or other suitable data and / or executable instructions to enable the genotyping computing device 110 to perform genotyping using artificial intelligence and / or deep learning.
[0109] 3 is a diagram illustrating the AI / DL module 210 (shown in FIG. 2) according to one embodiment. In some embodiments, the AI / DL module 210 may include a training set construction module 302 that is programmed to send one or more queries to the database 120 (shown in FIGS. 1 and 2) to retrieve data and / or subsets of data and to construct a training data set using these data subsets to generate a predictive model 308.
[0110] In exemplary embodiments, the training set construction module 302 is programmed to obtain training datasets from subsets of the obtained data. Each training dataset may correspond to past data, which may include one or more second datasets 224 (e.g., datasets excluding sequences containing structural artifacts) and / or other information used, received, and / or generated by the genotyping computing device 110, and may further include corresponding allele classification information determined previously rather than in real time at the time of acquisition by the training set construction module 302. Each training dataset may include model input data and result data representing allele classifications. The model input data may represent factors that are predicted to correlate with allele classifications or that may be unexpectedly discovered during model training. In some embodiments, the model input data may include one or more second datasets 224 (e.g., datasets excluding sequences containing structural artifacts) and / or any other information used, received, and / or generated by the genotyping computing device 110.
[0111] After the training set construction module 302 generates the training datasets, it passes the training datasets to a model training module 304. The model training module 304 is programmed to apply the model input data fields of each training dataset as inputs to one or more machine learning models. Each of the one or more machine learning models is programmed to generate, for each training dataset, at least one output that corresponds to or is intended to "predict" the value of at least one outcome data field of the training dataset. Machine learning techniques can be used to train models to identify and recognize patterns in existing data, thereby facilitating predictions for subsequent new input data. For example, support vector machines (SVMs), neural networks, and multilayer perceptron (MLP) classifiers may be used to train and optimize the models.
[0112] The model training module 304 is programmed to compare at least one output of the model with at least one outcome data field of the training dataset for each training dataset and apply a machine learning algorithm to adjust model parameters to reduce the difference, or "error," between the at least one output and the corresponding at least one outcome data field. In this manner, the model training module 304 trains a machine learning model to accurately determine allele classifications. In other words, the model training module 304 repeatedly applies one or more machine learning models to the training dataset, adjusting model parameters until the error between the at least one output and a predetermined allele classification is below a predetermined threshold, and then uploads the at least one trained machine learning model to the predictive model module 308 for application to new sequencing data. For example, the model training module 304 may be programmed to compare, for each training dataset, the allele classification determined by the model with the allele classification previously determined using a conventional testing method. The model learning module 304 may then use a machine learning algorithm, such as a backpropagation algorithm, to adjust one or more weights to reduce the difference between the allele classification determined by the model and the allele classification determined by conventional testing techniques.
[0113] In some embodiments, the one or more machine learning models may include one or more neural networks, such as a convolutional neural network, a deep learning neural network, or the like. The neural network may have one or more layers of nodes, and the model parameters adjusted during training may be individual weight values applied to one or more inputs for each node to generate a node output. In other words, the nodes in each layer may receive one or more inputs and apply weights to each input to generate a node output. The node inputs to the first layer may correspond to model input data fields, and the node outputs from the final layer may correspond to at least one output of the model, which may be intended to predict at least one outcome data field. For example, the node inputs to the first layer may correspond to sequencing data (e.g., long-read sequencing data), and the node output from the final layer may correspond to at least one allele classification. One or more intermediate layer nodes may be connected between the nodes in the first layer and the nodes in the final layer. As the model training module 304 sequentially processes the training data set, the model training module 304 applies a suitable backpropagation algorithm to adjust the weights of each node layer to minimize the error between at least one output (e.g., an allele classification determined by the model) and a corresponding outcome data field (e.g., a previously determined allele classification). In this manner, the machine learning model is trained to generate one or more outputs that reliably determine allele classifications. Alternatively, the machine learning model may have any suitable structure. In some embodiments, the model trainer module 304 provides the advantage of automatically detecting complex second-order, third-order, and / or other nonlinear interconnections between the model input data fields and at least one output and applying appropriate weights. Without the machine learning model, such connections would be unexpected and / or undetectable by a human analyst.
[0114] Additionally or alternatively, one or more machine learning models may include one or more multilayer perceptron (MLP) classifiers. An MLP classifier may have an input layer, an output layer, and one or more hidden layers each having a large number of neurons stacked thereon. For example, an MLP classifier according to the present disclosure may have an input layer containing sequencing data and an output layer containing allele classification.
[0115] Additionally or alternatively, the one or more machine learning models may include one or more support vector machines (SVMs). SVMs are supervised learning models with associated learning algorithms that analyze data for classification and regression analysis. More specifically, SVMs construct a hyperplane or set of hyperplanes in a high- or infinite-dimensional space that can be used for classification, regression, or other tasks such as outlier detection. For example, in some embodiments, training datasets may each be marked as belonging to one of two categories based on a predetermined allele classification. The SVM maps the training dataset to a point in the space and maximizes the width of the gap between the two categories. A new dataset is then mapped to the same space and predicted to belong to one of the categories based on which side of the gap the dataset lies on.
[0116] In some embodiments, the predictive model module 308 compares known allele classifications with outputs from the trained model and sends the comparison results to the model update module 306 of the AI / DL module 210. The model update module 306 is programmed to derive correction signals from the comparison results and provide the correction signals to the model training module 304, thereby enabling updating or "retraining" at least one machine learning model to improve its performance. For example, one or more new weight values may be derived from the comparison results, and the correction signal may adjust the weight values applied to one or more inputs. The retrained machine learning model may be periodically re-uploaded to the predictive model module 308.
[0117] In some embodiments, the model learning module 304 can further improve the accuracy of the computational model by updating the training dataset by creating one or more new historical records that include the new data and retraining the computational model with the updated training dataset.
[0118] The genotyping module 212 can utilize the AI / DL module 210 to determine allele classifications using a trained model. More specifically, the genotyping module 212 can determine allele classifications using output from the trained model. The determined allele classifications and other data can be viewable via the data management application 112.
[0119] The application module 214 is configured to enable maintenance of the data management application 112 and provision of its functionality to a user. The app module 214 may store instructions that enable downloading and / or execution of the data management application 112 on the user computing device 130. The app module 214 may store instructions regarding the user interface, controls, commands, settings, etc., and may format data in a format suitable for transmission to the user computing device 130 for display.
[0120] In some embodiments, the processor 202 is operatively connected to a communications interface 206, which enables the genotyping computing device 110 to communicate with remote device(s), such as the user computing device 130, the third party device 140, and / or the genotyping server 150 (all shown in FIG. 1 ), via a wired or wireless connection. For example, the communications interface 206 can receive sequencing data (e.g., long-read sequencing data) from the user computing device 130 and / or the third party device 140. The communications interface 206 can include, for example, a wired or wireless network adapter and / or a wireless data transceiver for use in a cellular communications network.
[0121] The processor 206 may be operatively connected to the database 120 (and / or other storage devices) via the storage interface 208. The database 120 may be any computer-controlled hardware suitable for storing and / or retrieving data. In some embodiments, the database 120 may be integrated into the genotyping computing device 110. For example, the genotyping computing device 110 may include one or more hard disk drives as the database 120. In other embodiments, the database 120 is external to the genotyping computing device 110 and is accessed by multiple computing devices. For example, the database 120 may include a storage area network (SAN), a network-attached storage (NAS) system, an inexpensive redundant array of disks (RAID) configuration using multiple storage units (e.g., hard disks and / or solid-state disks), a cloud storage device, and / or other suitable storage devices.
[0122] Storage interface 208 is any element capable of providing processor 202 with access to database 120. Storage interface 208 may include, for example, an Advanced Technology Attachment (ATA) adapter, a Serial ATA (SATA) adapter, a Small Computer System Interface (SCSI) adapter, a RAID controller, a SAN adapter, a network adapter, and / or any other element that provides processor 202 with access to database 120.
[0123] Processor 202 may execute computer-executable instructions to carry out aspects of the present disclosure. In some embodiments, processor 202 may be converted into a special-purpose microprocessor by executing computer-executable instructions or by being otherwise programmed. For example, processor 202 may be programmed with instructions such as those shown in FIG. 26.
[0124] The memory 204 may include, but is not limited to, random access memory (RAM), such as dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), and non-volatile random access memory (NVRAM). The above memory types are merely examples and are not intended to limit the types of memory that may be used to store computer programs.
[0125] Specific embodiments of the methods disclosed herein may be utilized for genotyping the filaggrin gene.
[0126] Filaggrin is a highly charged cationic protein that promotes the aggregation of keratin filaments and the subsequent formation of disulfide bonds. Filaggrin is a large (4400 kD) phosphorylated precursor derived from profilaggrin and is expressed as keratohyalin granules in the granular layer of the epidermis. During the transition from the granular layer to the stratum corneum, profilaggrin is converted to filaggrin by site-specific proteolysis and dephosphorylation. In addition to the processing of profilaggrin to filaggrin, the transition from granular cells to keratinocytes is characterized by the degradation of nuclei and other organelles, the formation of the cornified membrane, and the reorganization of the keratin intermediate filament network into a two-dimensional sheet.
[0127] Filaggrin plays a key role in the formation and maintenance of a flexible, hydrated stratum corneum, and its hydrolysis is carefully regulated to generate free amino acids, the major components of natural moisturizing factors (NMFs). The transition from the granular precursor profilaggrin to a diffusely distributed protein occurs rapidly at the granular-to-stratum corneum transition in response to a yet-to-be-identified initiation signal. The expression of profilaggrin as a precursor rather than as a mature protein suggests that filaggrin expression must be regulated to prevent cytotoxic effects.
[0128] Many inflammatory skin diseases are characterized by a weakening of the granular layer and the resulting residual keratinization (e.g., retention of keratinocyte nuclei in the stratum corneum). Although the signals that inhibit terminal differentiation in these inflammatory diseases may vary, a common final feature is the loss of the granular layer followed by incomplete terminal differentiation. When profilaggrin is reduced, as in atopic dermatitis, or nearly absent, as in ichthyosis vulgaris, the quality of the stratum corneum is impaired because the stratum corneum, depleted of natural moisturizing factors (NMFs), is unable to retain moisture under the drying effects of the environment.
[0129] Filaggrin is a filament-binding protein that plays an important role in the skin. Loss-of-function mutations in filaggrin have recently been reported to be correlated with several skin diseases, such as ichthyosis vulgaris, atopic eczema, asthma, and other allergic diseases.
[0130] As shown in Figure 4, filaggrin is composed of three exons. It is a repeat-rich gene, with 10 to 12 perfect repeats in the third exon of the gene. Each repeat is 972 to 975 nucleotides long.
[0131] Figure 4 further shows the variable length of the third exon of filaggrin. The third exon of filaggrin varies in length from 12,753 to 14,697 reads depending on the genotype of the individual. Currently, as known in the art, filaggrin alleles are classified into four genotypes containing 10, 11, or 12 repeats in the third exon.
[0132] Figure 5 shows PCR bias in a sequencing dataset for a CN-HET individual. The paternal allele 510 has 10 copies and an average read length of 13,000 nucleotides. Meanwhile, the maternal allele 520 has 12 copies and an average read length of 15,000 nucleotides. Due to PCR bias, the sequencing data for this individual contains approximately 60% of reads corresponding to the short paternal allele 510, while the long maternal allele 520 appears in only approximately 30% of the data.
[0133] Figure 6 shows PCR amplification of haplotypes in an exemplary dataset of CN-HET individuals, demonstrating an even distribution of target read lengths. The copy number of 10 repeats in the 13 kb read on the paternal chromosome allele 610 is amplified comparably to the larger repeat (12 repeats) on the maternal chromosome 620 or the 15 kb read. No artifacts are observed after aligning the PacBio sequencing data.
[0134] Due to the mechanism of DNA fragment amplification in PCR, imbalances in PCR are more pronounced when copy number differences are large. This is because a logarithmic increase in amplification is observed for short samples compared to long samples. To accurately genotype these individuals, it may be necessary to correct the data to appropriately reflect their genotype.
[0135] In FLG, each repeat is approximately 1 kb long, so a difference in copy number of just one repeat can significantly increase the degree of imbalance due to the large size of the increment. For example, each repeat in the FLG gene is approximately 1 kb, so a difference in repeat number of even 1 repeat can have a large impact on the observed amplification imbalance.
[0136] In some embodiments, the reads in the second dataset can be adjusted for PCR imbalance by indexing the sequencing reads into unique groups based on sequence structure and analyzing the most frequent of the groups to create a third dataset containing reads identifying the individual's alleles. The reads are organized into groups based on structure and ordered according to their frequency of occurrence in the dataset. The sample is classified as CN-HOM or CN-HET by analyzing the reads in the most frequent group of reads ("first frequency group") and, if present, the second most frequent group of reads ("second frequency group"). If the reads are separated into only a single frequency group in terms of copy number and structure, the sample is determined to be CN-HOM, and only the first frequency group is indexed as the individual's allele. If the reads are separated into two or more frequency groups, the two largest frequency groups are analyzed for sequencing noise.
[0137] In some embodiments, the two largest frequency groups may have the same copy number but different structures. The ratio of their frequencies may be compared. If the ratio is close to 1, the individual is classified as CN-HET, and both the first and second frequency groups are indexed and retained in the third dataset. If the ratio is greater than 1, the individual is classified as CN-HOM, and only the first frequency group is indexed and retained for inclusion in the third dataset.
[0138] In some embodiments, the reads in the second dataset are analyzed for sequencing noise through analysis of their frequency of occurrence. The reads are classified into groups based on their structure and ordered according to their frequency of occurrence in the dataset. The reads in the most frequent group ("first frequency group") and the next most frequent group ("second frequency group") are identified. A null distribution for the degree of PCR bias is created based on the length difference between the first and second frequency groups. The frequency distributions of the first and second frequency groups are then plotted on a logarithmic scale. Samples whose plot shows outliers from the expected frequency distribution are classified as CN-HOM, and samples whose frequency distribution aligns with the expected frequency distribution are classified as CN-HET. For CN-HOM samples, only the first frequency group is indexed and retained in the third dataset. For CN-HET samples, both the first and second frequency groups are indexed and retained in the third dataset.
[0139] In some embodiments, the reads in the second dataset are subjected to a sequencing noise analysis through copy number analysis. If the copy number of the second frequency group is greater than the copy number of the first frequency group, the individual is determined to be CN-HET, and the first and second frequency groups are indexed as alleles for the individual. If the copy number of the second frequency group is less than the copy number of the first frequency group, the individual is determined to be CN-HOM, and only the first frequency group is indexed as an allele for the individual.
[0140] In some embodiments, converting the second dataset to a third dataset involves classifying the sample as copy number heterogeneous (CN-HET) or copy number homogeneous (CN-HOM) based on the frequency ratio of the first and second most common read populations. If the frequency ratio is greater than expected, the reads are classified as CN-HOM. This is because if the frequency of the second frequency population is significantly lower than the frequency of the first frequency population, the reads in the second frequency population are likely to be sequencing artifacts. If the frequency ratio is approximately as expected, they are marked as CN-HET. This is because both read populations are observed more frequently than would be expected if one read population were composed of sequencing artifacts. If the length of the first primary read population is greater than the length of the second primary read population, the sample is classified based on whether the length of the first read population is greater or less than the length of the third read population.
[0141] For example, sequencing reads, which may have been filtered to remove sequencing and / or structural artifacts as described above, can then be classified by allele. This can be done by: (i) mapping sequences from one or more known individual repeat references to the sequencing reads to identify the structure of the reads; (ii) grouping the reads based on their structure; (iii) ranking the groups by read frequency; (iv) identifying reads in a first frequency group and, if present, a second frequency group; and / or (v) classifying individuals into CN-HOM or CN-HET based on the difference in length between the first and second frequency groups and / or the frequency of each group in the dataset. This length can include the total length of the gene or the copy number (i.e., the number of repeat exons).
[0142] The classification of an individual into CN-HOM or CN-HET in the above step (v) may be performed by applying one or more of the following rules (1) to (5): (1) If the reads are separated into only one frequency group based on copy number and structure, the sample is determined to be CN-HOM. (2) The read frequencies of each group are compared, and if the frequency ratio is higher than expected, the read is classified as CN-HOM. If the frequency ratio matches the expected ratio, the read is classified as CN-HET. The frequency distribution may be calculated by analyzing PCR bias based on the method described above. For example, a null distribution for the degree of PCR bias is created based on the difference in length between the first frequency group and the second frequency group. The frequency distribution between the first frequency group and the second frequency group is plotted on a logarithmic scale. Samples whose plot shows outliers from the expected frequency distribution are classified as CN-HOM, while samples whose frequency distribution matches the expected distribution are classified as CN-HET. (3) If the copy number of the second frequency group is smaller than that of the first frequency group, the individual is determined to be CN-HOM. (4) If the copy number of the second frequency group is greater than that of the first frequency group, the individual is determined to be CN-HET. The above rules (1) to (5) may be implemented by a trained model or an untrained model. For example, in some embodiments, rules (1) to (5) may be implemented by an untrained classification model. If the model is a trained model, the training data may include reads that have been filtered to remove sequencing artifacts or structural artifacts as described above to form the second dataset. The output for the training is allele classification. That is, allele classification in CN-HOM is performed by assigning heterozygous SNPs and / or indels to alleles. Allele classification in CN-HET is performed by assigning alleles based on copy number.
[0143] In some embodiments, reads classified as CN-HOM are further analyzed to determine the individual's alleles. The reads of the first frequency group are aligned to a customized reference to detect heterozygous SNPs and indels. The reads of the first frequency group are then separated into groups corresponding to each allele using the k-means clustering algorithm described in Lu, Y., Lu, S., Fotouhi, F. et al., "Incremental genetic K-means algorithm and its application in gene expression data analysis," BMC Bioinformatics 5, 172 (2004). https: / / doi.org / 10.1186 / 1471-2105-5-172.
[0144] FIG. 7 illustrates a flowchart of a process 700 for converting a second dataset to a third dataset, according to one embodiment. First, in 702, a reference sequence for each individual repeat is mapped to the reads of the second dataset, and a structure for each individual read is generated based on the repeats contained in the read and their order. In 704, the reads are grouped based on the sequence structure and ordered by frequency. That is, reads showing the same sequence structure are grouped together, and each group is then ordered based on the proportion of the structure appearing in all reads. Next, in 706, the reads are classified as CN-HOM or CN-HET based on the size ratio of the first and second most common read groups. In 708, samples determined to be CN-HOM are aligned to a customized reference using the first frequency group, and the reads are further classified into reads corresponding to each allele. The reads may be classified into reads corresponding to each allele using an AI / ML model. In a further embodiment, the AI / ML model includes a single-stage classifier and / or a multi-stage classifier. In some embodiments, the AI / ML model includes a k-means clustering algorithm. Samples are analyzed for heterozygous SNPs at 710. These SNPs are then used to cluster reads by allele based on the genotypes of the heterozygous SNPs at 712. The final output of the framework is a third dataset of reads grouped into alleles obtained from the first and second frequency groups for CN-HET individuals, or from the clusters determined by SNP genotypes for CN-HOM individuals, as shown at 714.
[0145] In some embodiments, the third dataset is further transformed into a fourth dataset by generating a consensus sequence for each allele of the third dataset.
[0146] In some embodiments, converting the third dataset to the fourth dataset can include generating a consensus sequence across each allele of the third dataset. Generating the consensus sequence can be accomplished using software such as SPOA, as described in "Fast and accurate de novo genome assembly from long uncorrected reads" by Vaser R, Sovic I, Nagarajan N, Sikic M (Genome Res. 2017 May;27(5):737-746. Epub 2017 Jan 18).
[0147] In some embodiments, the consensus sequence of the fourth data set can be aligned with a customized reference sequence to identify all nucleic acid mutations. The sequence of the consensus sequence can be mapped to the reference sequence of the target gene. The variant nucleotides between the reference sequence and the consensus sequence can be identified and further analyzed to detect the effect on codons.
[0148] In some embodiments, the consensus sequences of the fourth dataset may be further converted into a fifth dataset by translating them into the corresponding amino acid codes.
[0149] In some embodiments, the fourth dataset is converted to the corresponding amino acid code via the longest open reading frame ("ORF") using a software program such as the getorf utility in the EMBOSS package (described in EMBOSS: The European Molecular Biology Open Software Suite (2000), Rice, P., Longden, I., and Bleasby, A., Trends in Genetics 16, (6) pp276-277). The getorf utility in the EMBOSS package can detect the longest open reading frame of a consensus sequence by locating the start codon in the consensus sequence and determining the start codon that results in the longest translated region.
[0150] Figure 8 shows a flowchart of a process for analyzing atopic dermatitis ("AD") protein-truncating variants identified in a reference amino acid sequence 800, according to one embodiment. The process begins with identifying reads separated into two alleles at 802. A consensus sequence for each group of reads is generated at 804. At 806, open reading frames ("ORFs") are evaluated, the longest ORF from position 1 is selected, and each allele is translated into its corresponding amino acid sequence. At 808, samples with fewer than 10 reads are excluded from this process. At 808, the amino acid sequence is matched against the reference amino acid sequence to identify protein-altering variants. At 810, the protein-truncating variants are validated by two methods: (1) validation by internal whole exome sequencing ("WES") at 812, and (2) validation by an external database such as gnomAD™ (see Karczewski et al., 2020) at 814. gnomAD™ is a genetics dataset containing exome and genome sequence data from 141,456 multi-ethnic individuals.
[0151] At 816, for WES validation, the call set is further analyzed based on the entire exon 3 ORF. Protein-altering variants from this set are evaluated based on mutation type, such as frameshift mutations at position 818, stop-gain mutations at position 820, or missense mutations at position 822. At 824, frameshift and stop-gain mutations are advanced for concordance confirmation of common potential loss-of-function mutations ("pLoFs"). At 826, missense mutations are analyzed based on Hardy-Weinberg equilibrium (HWE) analysis with common variants and marked as likely false positives. This process includes filtering reads with fewer than 10 repeat units, aligning the longest ORF, translating DNA sequences into amino acid sequences (interpretation), and validating the variants. According to one embodiment, validation includes integrating variants with sequencing data, evaluating them using HWE, and correlating them with clinical traits.
[0152] The exon 3 ORF is analyzed, and amino acid coding changes are classified as frameshift, stop-gain, or missense mutations. Missense mutations are analyzed using the Hardy-Weinberg equilibrium equation and then flagged as likely false positives in 828. In 824, stop-gain and frameshift mutations are analyzed after confirming their concordance with common pLoF mutations. Next, samples with discordant genotypes ("GT") in two common pLoFs are further filtered in 830. In 832, structural genotypes are verified by confirming HWE of the common allele. The structural variants and PTVs are then integrated into a genotype matrix, as shown in 834. Next, samples are analyzed for association with AD severity 842 based on copy number subtype 836, pLoF dosage 838, and repeat number in the amino acid sequence 840.
[0153] Figure 9 shows an example of analyzing the sequence structure of repeats to remove sequencing noise from useful reads. In one embodiment, the sequences of exon 2, introns, and 5' and 3' unique regions are used to identify sequencing artifacts. As shown in Figure 9, reads lacking structural features found in known references, such as 5' and 3' unique regions, are excluded as noise. Similarly, reads containing structural features that do not match known variants, such as the triple repeat of Repeat 3 in the read shown at the bottom of Figure 9, are also discarded as noise.
[0154] Figure 10 is a flowchart illustrating a process for identifying protein-truncating variants in a dataset 1000, according to one embodiment. At 1002, the dataset begins with reads measuring two alleles. At 1004, the reads measuring the two alleles are translated into a consensus DNA sequence using software such as SPOA, generating a new dataset of consensus sequences corresponding to each allele. At 1006, ORFs in the consensus sequence are then identified using a software tool such as the getorf program in the EMBOSS software suite, creating a dataset of the longest ORF starting at position 1. Figure 11 shows a graph illustrating the difference in ORF length values generated for filaggrin data, according to one embodiment. The graph shows the full protein length as well as known disease-correlated variants and rare variants indicative of pLoF mutations.
[0155] Figure 12 illustrates unexpected noise in PacBio sequencing data. This figure shows a global alignment of generated read lengths with the Genome Reference Consortium Human Build 38 ("GRCh38") exon 3 reference, according to one embodiment. Multiple variants are observed, and there are clear patterns of misalignment, as indicated by gaps in the alignment data to the left and right of the reads. These gaps represent insertions or deletions due to sequencing artifacts, and therefore, reads containing these should be excluded as sequence artifacts, or noise.
[0156] Figures 13A and 13B show the alignment of a set of random reads from a single individual to the FLG region of GRCh38 using MUMmer 3.23, according to one embodiment. Figure 13A shows the alignment order between the sets of reads. Several different read length groups were observed due to insertions or deletions that occurred during sequencing of the 13 kb long reads. Figure 13B shows the size of these reads relative to the actual read length. The horizontal axis 1302 shows the difference in length between the read and the reference. The vertical axis 1304 shows the number of reads of that size.
[0157] Figures 14A and 14B illustrate filtering of sequencing noise based on structure. Figure 14A shows the structure of first and second frequency groups in a sample of CN-HET individuals, according to one embodiment. Reads are grouped based on their common structure, shown as box and hexagon shapes. Figure 14B shows a graph of reads, showing the number of reads in each group, ordered from smallest to largest, based on the frequency of the read group.
[0158] Figure 15 illustrates an analysis to detect CN-HOM samples by examining imbalance rates and allele lengths, according to one embodiment. Figure 15 illustrates two exemplary scenarios for analyzing imbalance rates and allele lengths of the most frequent read groups, representing cases where an individual is determined to be CN-HOM. Case 1 (1501) illustrates a scenario in which 95% of an individual's sequencing reads belong to a first frequency group and there is no clear second frequency group. Because there is no clear second frequency group, the individual is determined to be CN-HOM. Case 2 (1502) illustrates a scenario in which approximately 90% of an individual's sequencing reads belong to a first frequency group and approximately 8% belong to a second frequency group. Because the second frequency group contains fewer CNs than the first frequency group, the individual is determined to be CN-HOM.
[0159] 16 illustrates a process for identifying alleles of a CN-HOM individual through analysis against a customized reference, according to one embodiment. After aligning the sequence to the individually customized reference 1602, the alleles are separated through analysis of heterozygous SNPs. These reads are then separated into their corresponding alleles based on the alignment results. The two larger alleles are identified relative to the smaller allele.
[0160] Figure 17 illustrates three distinct variants 1702, 1704, and 1706, along with the overall PTV / pLoF burden 1708 observed in this study, according to one embodiment. More specifically, Figure 17 compares the identified potential loss-of-function mutations to previously analyzed mutations in the same gene. The frequency ratio of the variants was approximately three times higher in the patient population compared to the general population. Stop-gains are mutations that result in de novo premature termination of translation, leading to premature termination of translation. Two stop-gain mutations were observed: 1) an arginine stop codon at position 501 (e.g., variant 1702) and 2) an arginine stop codon at position 2447 (e.g., variant 1706). Frameshift mutations are insertions or deletions of nucleotide bases that are not a multiple of three. One frameshift mutation was observed at position 761 (e.g., variant 1704). This mutation results in the translation of serine to cysteine. pLoF load 1708 is a predicted loss-of-function mutation.
[0161] 18 shows a summary of the results of genotypic variables tested for filaggrin and atopic dermatitis ("AD"), according to one embodiment. Tests were conducted to test the hypothesis that FLG genotype, as measured by pLoF load 1802, dose of recurrent pLoF variants 1804, structural alleles 1806, copy number of functional repeats in DNA 1808, and / or copy number of functional repeats in amino acid sequence 1810, is associated with baseline AD disease severity 1812. Additionally, multiple regression analyses were performed between genotypic variables 1814 and phenotypic variables 1812, with age, sex, and the first 10 genetic PCs as covariates 1816.
[0162] Figures 19A-19C are box plots showing the association between functional repeat copy number (counting repeats fully translated as amino acid sequence and summing across both alleles) 1908 and baseline clinical phenotype, according to one embodiment. More specifically, BSABL 1902 (shown in Figure 19A) represents baseline values for body surface area, EASITSBSL 1904 (shown in Figure 19B) represents baseline values for Eczema Area and Severity Index, and SCORADBL 1906 (shown in Figure 19C) represents baseline values for SCORing Atopic Dermatitis Score. Copy number variables range from 0 to 10 (1910), 10 to 20 (1912), 20 to 22 (1914), and 22 to 26 (1916). The associations observed here provide evidence that the genotyping calling algorithm is accurate and functional. Patients with low functional copy numbers, as expected, presented with more severe disease.
[0163] Figure 20 is a graph showing the relationship between age at diagnosis 2002 and copy number 2004 for AD, according to one embodiment. Copy number variables ranged from 0-10 (2010), 10-20 (2012), 20-22 (2014), and 22-26 (2016), with age ranges from 0 to 30s. The associations observed here provide evidence that the genotyping calling algorithm is accurate and functional. Patients with lower functional copy numbers develop AD earlier.
[0164] 21A and 21B are graphs showing the relationship between pLoF and FLG genotype 2106 in relation to a history of asthma 2102 or a history of food allergies 2104 in AD patients, according to one embodiment. The observed associations are evidence that the genotyping calling algorithm is accurate and functional. Patients with low functional copy number are at increased risk for developing asthma and food allergies, both of which are known comorbidities in AD. As shown in FIGS. 21A and 21B, the genotype number corresponds to the occurrence of PTV in the alleles of each sample. Individuals were classified as genotype 0 (2110) if none of the alleles contained PTV, genotype 1 (2112) if one allele contained PTV, and genotype 2 (2114) if both alleles contained PTV.
[0165] Figures 22A and 22B are graphs showing the relationship of FLG genotype 2206 with pLoF associated with a history of AD 2202 or a history of food allergy 2204 in asthma patients, according to one embodiment. The associations observed here provide evidence that the genotyping calling algorithm is accurate and functional. Asthma patients with low functional copy number are at increased risk for developing AD and food allergies. As shown in Figures 22A and 22B, the genotype numbers correspond to the occurrence of PTV present in the alleles of each individual sample. Individuals were classified as genotype 0 (2210) if they did not contain PTV in both alleles, genotype 1 (2212) if they contained PTV in one allele, and genotype 2 (2214) if they contained PTV in both alleles.
[0166] Figure 23 shows known and novel PTVs mapped to their positions on exon 3 as frameshift mutations, pLoF mutations, and stop-gain mutations according to one embodiment of the present invention. The known PTVs were validated using internal WES data analysis and gnomAD™. This dataset was evaluated by HWE, and structural alleles were not excluded. However, due to the low frequency of the PTVs, they were not evaluated by HWE. The novel PTVs have not been previously reported.
[0167] Figure 24 shows the identification of known and novel copy number alleles validated based on the correlation between ORFs and DNA copy number, according to one embodiment. Novel, previously unreported structural alleles were identified at both lower and higher than expected copy numbers. The novel structural alleles were enriched in samples of African, European, or East Asian ancestry.
[0168] FIG. 25 is a graph showing an example of a structural artifact resulting in repeat amplification during PacBio sequencing, visualized by aligning an unusually long (19 kb) read to the reference sequence FLG (GRCh38), according to one embodiment of the present disclosure.
[0169] FIG. 26 shows a flow diagram 2600 of a process for converting an input dataset into a genotype for an individual sample, according to one embodiment of the present disclosure. At 2602, an input set of sequencing data is received. Next, at 2604, the input set of sequencing data is denoised to generate a first dataset that excludes reads containing sequence artifacts. Then, at 2606, the reads of the first dataset are filtered to generate a second dataset that excludes sequences containing structural artifacts. Next, at 2608, the reads in the second dataset are classified by allele to generate a third dataset. In some embodiments, the reads may be classified into reads corresponding to each allele using an AI / ML model. In further embodiments, the AI / ML model includes a single-stage classifier and / or a multi-stage classifier. For example, in some embodiments, the AI / ML model may include a k-means clustering algorithm. At 2610, a consensus sequence is generated for each allele in the third dataset to generate a fourth dataset. At 2612, each allele in the fourth dataset is translated to generate a fifth dataset. The fifth dataset includes an amino acid sequence corresponding to each allele. Then, at 2614, the individual's genotype is determined based on the amino acid sequences in the fifth dataset. [Example]
[0170] Example 1 Patient population The analysis included data from seven clinical trials of dupilumab, including two asthma trials and five atopic dermatitis trials.
[0171] asthma DRI12544 was a randomized, double-blind, placebo-controlled, parallel-group, pivotal Phase 2b clinical trial. The study enrolled patients with asthma aged 18 years or older. Patients were randomly assigned (1:1:1:1:1) to receive subcutaneous dupilumab 200 mg or 300 mg every 2 weeks or every 4 weeks, or placebo.
[0172] Study EFC13579 was a phase 3, double-blind, placebo-controlled, parallel-group study (NCT02414854) in patients aged 12 years and older with persistent asthma. Patients were randomized (2:2:1:1) to receive either 200 mg or 300 mg of dupilumab every 2 weeks or placebo.
[0173] The primary efficacy outcomes in both asthma studies were absolute change from baseline to week 12 in pre-bronchodilator forced expiratory volume in one second (FEV1) and annualized rate of serious exacerbation events.
[0174] Atopic dermatitis Phase 2 trials: Study 1307 was a placebo-controlled, double-blind, phase 2 study. Patients aged 18 years and older were randomly assigned 1:1 to receive dupilumab 200 mg once weekly or placebo. Study 1021 was a placebo-controlled, double-blind, phase 2b study in patients aged 18 years and older. Patients were randomly assigned (1:1:1:1:1:1) to receive dupilumab 300 mg once weekly, 300 mg every 2 weeks, 200 mg every 2 weeks, 300 mg every 4 weeks, 100 mg every 4 weeks, or placebo once weekly. The primary endpoint in both phase 2 studies was the mean percentage change in the EASI score from baseline to week 16.
[0175] Phase 3 trials: Study 1334 and Study 1416 were identical, placebo-controlled, phase 3 trials (SOLO 1 and SOLO 2) enrolling patients aged 18 years or older with moderate-to-severe atopic dermatitis inadequately controlled with topical treatments. Patients were randomized (1:1:1) to receive dupilumab (300 mg) once weekly, placebo once weekly, or the same dose of dupilumab every other week, alternating with placebo in between. The primary endpoint was the proportion of patients achieving an Investigator's Global Assessment score of 0 or 1 (clearly improved or almost clearly improved) and a 2-point or greater reduction from baseline at week 16.
[0176] Study 1224 was a placebo-controlled, phase 3 trial in adults aged 18 years or older with moderate to severe atopic dermatitis who had an inadequate response to topical corticosteroids. Patients were randomized (3:1:3) to receive dupilumab 300 mg once weekly (qw), dupilumab 300 mg every 2 weeks (q2w), or placebo. Concomitant topical corticosteroids were administered in all study arms. The primary endpoint of the IGA was defined similarly to other atopic dermatitis studies, with an additional coprimary endpoint of 75% improvement from baseline in the Eczema Area and Severity Index (EASI-75) at week 16.
[0177] Example 2 DNA samples were processed to amplify a 13.5 kb region encompassing exons 2 and 3 of filaggrin. This amplification was performed using a two-step PCR method, generating 96 barcoded amplicons, which were then subjected to multiplex sequencing using the PacBio® Sequel II system. The filaggrin locus was amplified using target-specific sequences ligated with universal adapter sequences. In the second PCR reaction, each amplicon was barcoded using primers ligated with unique barcode sequences ligated to the universal sequence. The barcoded amplicons were analyzed by automated electrophoresis using an Agilent® ZAG DNA Analyzer System for the presence or absence of 1 kb bands in the 1-18 kb range. Agilent® ProSize software was used to perform smear analysis in the 9-18 kb range, and the molar concentrations of the desired (9-18 kb) and undesired (1-9 kb) portions of each amplicon were quantified.
[0178] Samples were sequenced in 96-sample increments and equimolarly mixed based on smear analysis of 9–18 kb. To remove excess 1–9 kb regions, the multiplexed amplicon pool was subjected to size separation using the SageELF electrophoresis system. All samples were fractionated into 12 fractions ranging in size from 1–18 kb. The estimated width of each fraction was determined by the SageELF instrument. The sizes of the 12 fractions were determined by automated electrophoresis using an Agilent® ZAG DNA Analyzer system. Three fractions with sizes greater than 9 kb were pooled together and used for PacBio® SMRTbell library preparation for sequencing using the PacBio® Sequel II, according to the manufacturer's recommendations.
[0179] Sequencing data were generated by CCS software, part of the Pacbio® SMRT® Link package.
[0180] Example 3 FLG genotyping pipeline 3.1 Identifying technical artifacts in sequencing data Based on the four known structural alleles, PacBio read lengths were expected to be approximately 13 kb, 14 kb, or 15 kb. For each individual, either one read length (derived from two alleles with the same repeat number) (referred to as "CN-HOM") or two read lengths (derived from two alleles with different copy numbers) (referred to as "CN-HET") were expected to exist, with a roughly 1:1 ratio. However, for many individuals, the distribution of read lengths showed a wide range of possible values, some of which were greater than 15 kb or less than 13 kb.
[0181] Alternatively, we attempted to cluster read lengths around 13 kb, 14 kb, and 15 kb. We observed that size ratios between any two read groups were rarely nearly equal. Possible technical artifacts include: 1) PCR bias: In CN-HET individuals, PCR over-amplifies short FLG alleles relative to long alleles, resulting in imbalances in sequencing depth between alleles. 2) Sequencing-induced insertions / deletions: Due to the repeat structure of exon 3 and the characteristics of long-read sequencing technology, extra insertions / deletions were introduced in some reads. Visualization of this artifact is shown in Figure 11, where a random set of reads from each read length group for an individual was aligned to the FLG reference of GRCh38 using MUMmer 3.23.
[0182] 3.2 Removing sequencing noise by measuring repeat structure Examination of the repeat structure sequences revealed that the sequences of repeats 1 through 9 and the three forms of repeat 10 were highly homologous to each other but sufficiently different to allow for unique alignment to the reference sequence using the default parameters of the local aligner. The structural information obtained by local alignment was used to remove sequencing noise. The reference sequence was divided into regions representing exon 2, introns, the exon 3-specific region near the 5' end, repeats 1 through 9, repeats 10a, 10b, and 10c, and the exon 3-specific region near the 3' end. These 16 reference fragments were aligned to all PacBio reads using the all-to-all alignment function in minimap2. Next, each read was sorted by start position based on the alignment results, and the order of the reference fragments aligned to that read was determined. For each individual, the structures of all reads were summarized using BEDTools software, and reads with identical sequence structures were grouped together. The read groups were arranged by frequency (based on the number of reads in each group) from most to least frequent. Under the assumption that sequencing noise is random and sparse, the reads belonging to the most frequent group of reads ("first frequency group") and the second most frequent group of reads ("second frequency group") were retained, and the remaining reads were considered to contain sequence variants derived from artifacts.
[0183] 3.3 Read classification by allele As a next step in the pipeline, the useful reads were classified based on the alleles they represented, a classification similar to the concept of so-called "phasing."
[0184] As described above in Example 3.1, samples with the same copy number of alleles are CN-HOM, and samples with different copy numbers of alleles are CN-HET. For CN-HET, samples in the first frequency group measured the first allele, and samples in the second frequency group measured the second allele. For CN-HOM samples, both alleles were included in the first frequency group, and the second frequency group was likely to be sequencing noise. CN-HOM and CN-HET samples are distinguished by ML and / or AI. For example, in some embodiments, CN-HOM and CN-HET samples are distinguished by an AI / ML model. In yet another embodiment, the AI / ML model includes a single-stage classifier and / or a multi-stage classifier.
[0185] Given the over-amplification of short alleles, we expected that the first frequency group would be shorter than the second frequency group in CN-HET samples. If this was not the case, the second frequency group was likely to be rare sequencing noise, and samples containing long alleles in the first frequency group were likely to be CN-HOM samples. Therefore, individuals in which the length of the first frequency group was greater than the length of the second frequency group were classified as CN-HOM samples.
[0186] For samples with the same copy number but different structural alleles, the frequencies of the first and second frequency groups were expected to be similar (ratio close to 1). If instead the first frequency group was observed to be significantly more prevalent (ratio ≫ 1), the second frequency group was flagged as noise and the sample was classified as CN-HOM.
[0187] For samples in which the first frequency group represented the shorter allele, a null distribution ("expected value") for the degree of PCR bias was derived based on the difference in the lengths of the two alleles. PCR bias was measured by subtracting the copy number of the first frequency group from the copy number of the second frequency group. For example, when the copy number of the first frequency group was one copy less than that of the second frequency group (CNsecond - CNfirst = 1), the distribution of the frequency ratio (Freqfirst / Freqsecond) was plotted on a logarithmic scale. Even after accounting for sampling imbalance due to PCR bias, outliers were assumed to be CN-HOM samples in which the second frequency group was too rare to be considered a true signal. The same procedure was repeated for the subset of samples in which (CNsecond - CNfirst = 2). In this case, the imbalance rate was higher. Some samples with CNsecond - CNfirst > 2 were classified as CN-HET due to the small number of observations, making it difficult to derive a null distribution.
[0188] After identifying the CN-HOM samples, reads belonging to the first frequency group were phased into two distinct alleles. Because structural variants provided no further information, the two read groups were clustered based on the genotypes of smaller variants, SNPs and indels. To improve alignment accuracy, we first created a customized reference sequence for each sample that mimicked the individual's two allele structures. Next, we applied a global alignment (Smith-Waterman algorithm) between the first frequency group reads and the customized reference sequence to detect heterozygous (HET) SNPs / indels (defined as mismatches with a VAF of approximately [0.25, 0.8]). As a final step, we applied a k-means (k = 2) clustering algorithm to the genotype matrix of HET SNPs and indels to classify the reads into two allele groups.
[0189] 3.4 Detection of protein truncation variants If the sequencing depth for both alleles was sufficient (for group 2, only samples with ≥10 reads were included), we used SPOA software to generate two consensus sequences representing the DNA sequence of the FLG gene for each individual in the dataset. These consensus sequences were aligned to a customized reference sequence to identify all nucleic acid variants. However, given that repeats typically form multiple filaggrin monomers that support skin function, we focused on protein-truncating variants (PTVs), which are likely to affect phenotypes and are translated into amino acid sequences.
[0190] Each consensus sequence was translated into the longest open reading frame (ORF) using the getorf utility in the EMBOSS package, provided that the starting sequence was identical or similar to the reference. The translated ORF was then compared to the expected ORF length based on the number of repeats in exon 3. If the observed sequence was shorter, it was identified as a PTV. The truncated ORF was then aligned to the reference amino acid sequence and classified as either a stop-gain variant or a frameshift variant. To identify the location of a specific PTV, the position of the last aligned base and the corresponding amino acid were recorded. This method was performed on both alleles of all samples to annotate and determine the allele frequencies of recurrent PTVs.
[0191] 3.5 Call Set Quality Evaluation The PTVs were validated by internal whole exome sequencing ("WES") data obtained on the same set of samples and the external database gnomAD™, a genetics dataset containing exome and genome sequence data for 141,456 multiethnic individuals (see Karczewski et al., 2020).
[0192] The quality of copy number genotypes for FLG repeats was assessed by HWE testing. HWE was calculated for each genetically determined ancestry subgroup (African, East Asian, Mixed American, European, and South Asian) when appropriate. Significant deviations from HWE (p<1×10) in any ancestry group were not associated with a significant HWE error. -6 ) was not observed. Thus, the methods disclosed herein accurately determined copy number genotypes for FLG repeats.
[0193] Further identification of known and novel copy number alleles was verified by correlation of ORFs with DNA copy number. Novel structural alleles were identified at both lower and higher than expected copy numbers. The novel structural alleles were enriched in samples with African, European, or East Asian ethnic backgrounds. Known and novel PTVs were mapped via frameshift, pLoF, and stopgain sequences corresponding to their positions on exon 3. Known PTVs were verified by internal WES data analysis and gnomAD™.
[0194] An exemplary server computing device is disclosed herein. For example, FIG. 27 is a schematic diagram illustrating the configuration of an example server computing device 2700, according to some embodiments of the present disclosure. A server computing device having an architecture similar to server computing device 2700 may be used to implement one or more computing systems, such as genotyping computing device 110, shown in FIG. 1. In an exemplary embodiment, server computing device 2700 includes a processor 2705 for executing instructions (not shown) stored in memory 2710. In one embodiment, processor 2705 may include one or more processing units (e.g., in a multi-core configuration). The instructions may be executed on a variety of different operating systems, such as UNIX®, LINUX® (LINUX is a registered trademark of Linux Torvalds), Microsoft Windows®, etc. It should also be understood that various instructions may be executed during an initialization process at the start of a computer-based method. Some operations may be necessary to perform one or more processes described herein, while other operations may be more general or specific to a particular programming language (e.g., C, C#, C++, Java, or other suitable programming language).
[0195] In an exemplary embodiment, the processor 2705 is operably connected to a communications interface 2715, which enables the server computing device 2700 to communicate with remote devices, such as a user's or system administrator's computing system (not shown) or other server computing devices 2700.
[0196] In an exemplary embodiment, processor 2705 is also operably connected to storage device 2730, which may be, for example, a computer-operable hardware unit suitable for storing or retrieving data. In some embodiments, storage device 2730 is integrated into server computing device 2700. For example, device 2700 may include one or more hard disk drives as storage device 2730. In other embodiments, storage device 2730 may be external to device 2700 and accessed by multiple server computing devices 2700. For example, storage device 2730 may include multiple storage units, such as hard disks or solid-state disks, in a redundant array of inexpensive disks (RAID) configuration. Storage device 2730 may include a storage area network (SAN) or network-attached storage (NAS) system. Storage device 2730 may be used as a repository for one or more databases or other data structures for storing various data elements received, processed, and / or generated by genotyping computing device 110. For example, storage device 2730 may be used to implement database 120 (shown in FIGS. 1 and 2).
[0197] In some embodiments, processor 2705 is operably connected to storage device(s) 2730 via optional storage interface 2720. Storage interface 2720 may include, for example, components capable of providing processor 2705 with access to storage device(s) 2730. In one embodiment, storage interface 2720 further includes one or more of an Advanced Technology Attachment (ATA) adapter, a Serial ATA (SATA) adapter, a Small Computer System Interface (SCSI) adapter, a RAID controller, a SAN adapter, a network adapter, or components with similar functionality, to provide processor 2705 with access to storage device(s) 2730.
[0198] Memory area 2710 may include, but is not limited to, random access memory (RAM), such as dynamic RAM (DRAM) or static RAM (SRAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), non-volatile RAM (NVRAM), and magnetoresistive random access memory (MRAM). The above memory types are exemplary only and are not intended to limit the types of memory that may be used to store computer programs.
[0199] An exemplary user computing device is disclosed herein. For example, FIG. 28 illustrates an exemplary configuration of a user computing device 2802. The user computing device 2802 includes a processor 2804 for executing instructions. In some embodiments, executable instructions are stored in a memory area 2806. The processor 2804 may include one or more processing units (e.g., in a multi-core configuration). The memory area 2806 is any device that allows for the storage and retrieval of information, such as executable instructions and / or other data. The memory area 2806 may include one or more computer-readable media.
[0200] The user computing device 2802 also includes at least one media output component 2808 for presenting information to the user. The media output component 2808 is any component capable of communicating information to a user. In some embodiments, the media output component 2808 includes an output adapter, such as a video adapter and / or an audio adapter. The output adapter is operably connected to the processor 2804 and is operably connectable to an output device, such as a display device (e.g., a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a cathode ray tube (CRT), or an "electronic ink" display), or an audio output device (e.g., speakers or headphones). For example, a user can view genotyping information via the media output component 2808.
[0201] In some embodiments, user computing device 2802 includes input devices 2810 for receiving input from a user. Input devices 2810 may include, for example, a keyboard, a pointing device, a mouse, a stylus, a touch-sensitive panel (e.g., a touchpad or touchscreen), a camera, a gyroscope, an accelerometer, a position detector, and / or an audio input device. A single component, such as a touchscreen, may function as both an output device for media output component 2808 and as input device 2810.
[0202] The user computing device 2802 may also further include a communications interface 2812 that is communicatively connectable to a remote device, such as a server system or a web server operated by a biometric registry. The communications interface 2812 may include, for example, a wired or wireless network adapter or a wireless data transceiver intended for use with a cellular network (e.g., Global System for Mobile Communications (GSM), 3G, 4G, or Bluetooth) or other mobile data network (e.g., Worldwide Interoperability for Microwave Access (WIMAX)).
[0203] Disclosed herein are methods of machine learning. The computer-implemented methods described herein may include additional, fewer, or alternative actions, including those described elsewhere herein. These methods may be performed via computer-executable instructions stored on one or more local or remote processors, transceivers, servers, and / or persistently computer-readable medium(s).
[0204] In addition, the computer systems described herein may include additional, fewer, or alternative functionality, including functionality described elsewhere herein. The computer systems described herein may include or be executed by computer-executable instructions stored on a persistent computer-readable medium(s).
[0205] The processor or processing element can learn using supervised or unsupervised machine learning, and the machine learning program may employ a neural network, which may be a convolutional neural network, a deep learning neural network, a reinforcement learning or reinforcement learning module or program, or a hybrid learning module or program that trains in two or more fields or areas of interest. Machine learning may identify and recognize patterns in existing data to facilitate prediction of subsequent data. Models may be created based on example inputs to make valid and reliable predictions for new inputs.
[0206] Additionally or alternatively, the machine learning program can be trained by inputting sample datasets or specific data into the program, including sequencing data (e.g., long-read sequencing data) 220, a first dataset 222 (e.g., a dataset excluding reads containing sequence artifacts), a second dataset 224 (e.g., a dataset excluding sequences containing structural artifacts), a third dataset 226 (e.g., a dataset including sequencing reads classified by allele), a fourth dataset 228 (e.g., a dataset including consensus reads for each allele), a fifth dataset 230 (e.g., a dataset including amino acid sequences corresponding to each allele), and / or any other information used, received, and / or generated by the genotyping computing device 110. The machine learning program can primarily utilize deep learning algorithms that may focus on pattern recognition and can be trained after processing multiple examples. Machine learning programs may include, individually or in combination, Bayesian Program Learning (BPL), speech recognition and synthesis, image or object recognition, optical character recognition, and / or natural language processing. Machine learning programs may also include natural language processing, semantic analysis, automated reasoning, and / or machine learning.
[0207] Supervised and unsupervised machine learning techniques may be used. In supervised machine learning, a processing element is provided with example inputs and their associated outputs and discovers general rules that map the inputs to outputs, so that when a subsequent new input is provided, the processing element can accurately predict the correct output based on the discovered rules. In unsupervised machine learning, a processing element may be required to discover its own structure within unlabeled example inputs. In some embodiments, machine learning techniques may be used to extract data about specific alleles based on sequencing reads (e.g., sequencing reads that are free of sequence and / or structural artifacts).
[0208] In some embodiments, the voicebots or chatbots referred to herein may be configured to utilize ML and / or AI techniques. For example, the voicebots or chatbots may be AI chatbots. The voicebots or chatbots may use supervised or unsupervised machine learning techniques, followed by and / or in conjunction with reinforcement learning techniques.
[0209] As can be appreciated from the foregoing description, the above-described embodiments of the present disclosure can be implemented using computer programming or engineering techniques, including computer software, firmware, hardware, or any combination or subset thereof. The resulting program having computer-readable code means can be embodied or provided in one or more computer-readable media, thereby creating a computer program product, i.e., an article of manufacture, in accordance with the embodiments discussed in this disclosure. The computer-readable medium can be, for example, but not limited to, a fixed (hard) drive, a diskette, an optical disk, a magnetic tape, a semiconductor memory such as a read-only memory (ROM), an SD card, a memory device, and / or any transmission / reception medium, such as the Internet or other communications network or link. An article of manufacture containing the computer code can be created and / or used by executing the code directly from one medium, by copying the code from one medium to another, or by transmitting the code over a network.
[0210] These computer programs (also known as programs, software, software applications, "apps," or code) include machine instructions for a programmable processor and may be implemented in high-level procedural and / or object-oriented programming languages and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or equipment (e.g., magnetic disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, such as a machine-readable medium that receives machine instructions as machine-readable signals. However, "machine-readable medium" and "computer-readable medium" do not include transitory signals. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0211] As used herein, a processor may include any programmable system, including systems that use microcontrollers, reduced set of circuits (RISC), application specific integrated circuits (ASIC), logic circuits, and other circuits or processors that perform the functions described herein. The above examples are merely examples and are not intended to limit in any way the definition and / or meaning of the term "processor."
[0212] As used herein, "software" and "firmware" are used interchangeably and include any computer program stored in memory for execution by a processor, such as RAM memory, ROM memory, EPROM memory, EEPROM memory, and non-volatile RAM (NVRAM) memory. The above memory types are exemplary only and are not intended to limit the types of memory that can be used to store computer programs.
[0213] In some embodiments, a computer program is provided, the program being recorded on a computer-readable medium. In an exemplary embodiment, the system runs on a single computer system and does not require connection to a server computer. In a further embodiment, the system runs in a Windows® (Windows is a registered trademark of Microsoft Corporation, Redmond, Washington) environment. In yet another embodiment, the system runs in a mainframe environment and a UNIX® (UNIX is a registered trademark of X / Open Company Limited, Reading, Berkshire, England) server environment. The application is flexible and designed to run in a variety of environments without compromising its primary functionality.
[0214] In some embodiments, a system includes multiple elements distributed across multiple computing devices. One or more elements may be in the form of computer-executable instructions embodied on a computer-readable medium. The system and processes are not limited to the specific embodiments described herein. In addition, each system element and each process can be performed independently and independently of other elements and processes described herein. Each element and process can also be used in combination with other assembly packages and processes. The embodiments may enhance the functionality and capabilities of a computer and / or computer system.
[0215] In this specification, the use of the words "a" or "an" in the singular does not exclude a plurality of elements or steps, unless the context clearly dictates otherwise. Furthermore, references to an "exemplary embodiment" or "one embodiment" of the present disclosure are not intended to be interpreted as excluding the existence of additional embodiments that incorporate the recited features.
[0216] The claims set forth at the end of this document are not intended to be construed under 35 U.S.C. § 112(f) unless traditional functional language is expressly recited in the claim(s), such as by expressly reciting "means for" or "step for" language.
[0217] This written description uses the disclosed examples to disclose the best mode, and also to enable any person skilled in the art to practice the disclosure, including making and using any devices or systems, and practicing the incorporated methods. The patentable scope of the disclosure is defined by the claims, and may include other examples that occur to those skilled in the art. If the embodiment has structural elements that do not differ from the literal language of the claims, or if the embodiment includes equivalent structural elements that are not substantially different from the literal language of the claims, then such other embodiments are intended to be within the scope of the claims.
Claims
1. 1. A system for determining the genotype of a gene comprising at least one repeat region, comprising: at least one memory storing computer-executable instructions; at least one processor in communication with the at least one memory, the at least one processor executing the computer-executable instructions to: receiving an input set of genetic sequencing data; measuring at least one repeat structure of at least one sequence read of the gene sequencing data and removing sequence reads having an aberrant repeat structure to generate a first dataset, wherein the first dataset excludes sequence reads having sequence artifacts; filtering the sequence reads of the first dataset to generate a second dataset, the second dataset excluding sequences with structural artifacts; applying input data to a model to generate a third dataset, the input data comprising sequence reads of the second dataset, and the third dataset comprising sequence reads classified by allele; generating a consensus sequence for each allele included in the third dataset and generating a fourth dataset of consensus reads for all alleles included in the third dataset; translating each allele of the fourth dataset to generate a fifth dataset, the fifth dataset comprising an amino acid sequence corresponding to each allele; determining a protein-altering variant state based on the amino acid sequences of the fifth dataset; The system is configured to:
2. 2. The system of claim 1, wherein the at least one processor is further configured to execute the computer-executable instructions to compare amino acid sequence lengths of the fifth dataset with expected lengths of an amino acid reference sequence to generate a sixth dataset of protein truncation variants.
3. 3. The system of claim 2, wherein the at least one processor is further configured to execute the computer-executable instructions to classify the protein-truncating variants in the sixth dataset as stop-gain variants or frameshift variants.
4. The system of any one of claims 1 to 3, wherein the at least one processor is further configured to execute the computer-executable instructions to remove sequencing artifacts by detecting read structure by alignment, grouping sequence reads based on structure, sorting reads by frequency, and identifying novel alleles with previously unknown copy numbers.
5. The system of claim 4 , wherein the expected length is a size range determined by the range of lengths of variants of the gene that are known to exist.
6. 4. The system of claim 1, wherein the at least one processor is further configured to execute the computer-executable instructions to filter out structural artifacts by mapping sequences from known references for each individual repeat to the sequence reads to determine the structure of the reads, and to filter out reads containing repeat structures that do not represent known mutations of the gene and are therefore considered sequencing artifacts.
7. 7. The system of claim 1, wherein at least one output of the model comprises a classification of sequence reads of the second dataset as either copy number heterogeneous or copy number homogeneous.
8. The at least one processor further executes the computer-executable instructions to classify the sequence reads of the second dataset as copy number heterogeneous or copy number homogeneous: clustering the sequencing reads of the second dataset by structure; ordering the groups by frequency; Identifying a first group of frequent reads and, if present, a second group of frequent reads; classifying the genes as copy number homogeneous or copy number heterogeneous based on the difference in length of the first and second frequent groups and the difference in frequency of each group in the second dataset; The system of claim 7 , configured to classify by:
9. 9. The system of claim 8, wherein the at least one processor is further configured to execute the computer-executable instructions to, during execution of the model, separate sequence reads classified as copy number uniform by analyzing heterozygous SNPs and indels in the DNA sequence of the sequence reads for each allele.
10. 10. The system of claim 9, wherein the at least one processor is further configured to execute the computer-executable instructions to separate reads by analyzing, for each allele, heterozygous SNPs and indels in their DNA sequences using a k-means clustering algorithm.
11. 1. A computer-implemented method for genotyping a repeat-rich gene, the method being performed using a system comprising a computing device including a processor communicatively connected to a memory device, the method comprising: receiving an input set of genetic sequencing data; measuring at least one repeat structure of at least one sequence read of the gene sequencing data and removing sequence reads having an aberrant repeat structure to generate a first dataset, wherein the first dataset excludes sequence reads having sequence artifacts; filtering the sequence reads of the first dataset to generate a second dataset, the second dataset excluding sequences with structural artifacts; applying input data to a model to generate a third dataset, the input data comprising sequence reads of the second dataset, and the third dataset comprising sequence reads classified by allele; generating a consensus sequence for each allele included in the third dataset and generating a fourth dataset of consensus reads for all alleles included in the third dataset; translating each allele of the fourth dataset to generate a fifth dataset, the fifth dataset comprising an amino acid sequence corresponding to each allele; determining a protein-altering variant state based on the amino acid sequences of the fifth dataset.
12. 12. The computer-implemented method of claim 11, further comprising comparing the amino acid sequence lengths of the fifth dataset with expected lengths of an amino acid reference sequence to generate a sixth dataset of protein truncation variants.
13. 13. The computer-implemented method of claim 12, further comprising classifying the protein-truncating variants in the sixth dataset as stop-gain variants or frameshift variants.
14. 14. The computer-implemented method of any one of claims 11 to 13, further comprising filtering out sequencing artifacts by detecting read structure by alignment, grouping sequence reads based on structure, sorting reads by frequency, and identifying novel alleles with previously unknown copy numbers.
15. 15. The computer-implemented method of claim 14, wherein the expected length is a size range determined by the range of lengths of variants of the gene that are known to exist.
16. 14. The computer-implemented method of any one of claims 11 to 13, wherein structural artifacts are filtered out by mapping sequences from known references of each individual repeat to the sequence reads to determine the structure of the reads, and filtering out reads containing repeat structures that do not represent known variants of the gene and are therefore considered sequencing artifacts.
17. 17. The computer-implemented method of any one of claims 11 to 16, wherein at least one output of the model comprises a classification of sequence reads of the second dataset as either copy number heterogeneous or copy number homogeneous.
18. Classifying sequence reads as copy number heterogeneous or copy number homogeneous involves: clustering the sequencing reads of the second dataset by structure; ordering the groups by frequency; Identifying a first group of frequent reads and, if present, a second group of frequent reads; and classifying the gene as copy number homogeneous or copy number heterogeneous based on the difference in length of the first and second frequent groups and the difference in frequency of each group in the second dataset.
19. 20. The computer-implemented method of claim 18, further comprising, during execution of the model, for each allele, separating reads by analyzing heterozygous SNPs and indels in their DNA sequences using a k-means clustering algorithm.
20. At least one persistent computer-readable storage medium having computer-executable instructions stored thereon, the computer-executable instructions, when executed by at least one processor, causing the at least one processor to: receiving an input set of genetic sequencing data; measuring at least one repeat structure of at least one sequence read of the gene sequencing data and removing sequence reads having an aberrant repeat structure to generate a first dataset, wherein the first dataset excludes sequence reads having sequence artifacts; filtering the sequence reads of the first dataset to generate a second dataset, the second dataset excluding sequences with structural artifacts; applying input data to a model to generate a third dataset, the input data comprising sequence reads of the second dataset, and the third dataset comprising sequencing reads grouped by allele; generating a consensus sequence for each allele in the third dataset and generating a fourth dataset of consensus reads for all alleles included in the third dataset; translating each allele of the fourth dataset to generate a fifth dataset, the fifth dataset comprising an amino acid sequence corresponding to each allele; determining a protein-altering variant state based on the amino acid sequences of the fifth dataset; the at least one persistent computer-readable storage medium causing
Citation Information
Patent Citations
Sequence assembly and consensus sequence determination
US9165109B2