Genotyping variable number tandem repeats

JP2024522702A5Pending Publication Date: 2025-06-23ILLUMINA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023577216
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-06-14
Filing Date
2022-06-13
Publication Date
2025-06-23

AI Technical Summary

Technical Problem

Accurate detection of variable number tandem repeats (VNTRs) in genome sequencing is complicated by their low complexity and repetitive nature, leading to low detection power in existing short read pipelines.

Method used

A method and system for determining VNTR status using long sequence reads to construct a database of common haplotypes, followed by realigning short sequence reads to these haplotypes, and applying probability indices to determine the most likely VNTR genotypes, utilizing a processor to generate a user interface for representation.

Benefits of technology

Improves VNTR genotyping accuracy from 16% to 78% by leveraging long reads to build a database and optimizing short read realignment, enabling precise haplotype and genotype determination.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Disclosed herein is a system, device and method for determining variable number tandem repeat (VNTR) status.The haplotype of VNTR can be determined using the long sequence read of the reference sample aligned to the VNTR in the reference.The short read of the test sample of the test subject can be aligned to the haplotype determined using the long sequence read to determine the VNTR status of the test subject (e.g., one or more haplotypes or genotypes of the test subject) based on the probability indicator of the haplotype.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] Related Applications This application claims priority under 35 USC § 119(e) to U.S. Provisional Application No. 63 / 210,294, filed June 14, 2021. The contents of the related application are incorporated herein by reference in their entirety.

[0002] Sequence Listing Reference This application has been filed with a sequence listing in electronic format. The sequence listing is provided as a file entitled Sequence_Listing_47 CX-311979-WO, created May 29, 2022, which is 1 kilobyte in size. The information in the electronic format of the sequence listing is incorporated herein by reference in its entirety.

[0003] The present disclosure relates generally to the field of processing sequencing data, and more specifically to genotyping variable number tandem repeats. [Background technology]

[0004] Variable nucleotide tandem repeats (VNTRs) account for a significant proportion of intergenomic variation. Accurate detection of VNTRs has long been complicated by the low complexity nature of the regions and the length of the repeated sequences. The detection power of VNTRs in existing short read pipelines requires improvement. Summary of the Invention

[0005] Disclosed herein is a method for determining variable number tandem repeat (VNTR) status, such as genotyping VNTR. In some embodiments, the method for determining VNTR status is under the control of a processor (e.g., a hardware processor or a virtual processor) and includes receiving a plurality of long sequence reads generated from a plurality of first samples obtained from a plurality of first subjects. The method may include determining a plurality of haplotypes of the VNTR using a long sequence read from the plurality of long sequence reads aligned to the VNTR in a reference (e.g., a reference human genome sequence, e.g., hg19 or hg38). The method may include receiving a plurality of short sequence reads generated from a second sample obtained from a second subject. The method may include rearranging the short sequence reads from the plurality of short sequence reads aligned to the VNTR into haplotypes to generate rearrangements for each of the plurality of haplotypes of the VNTR. The method may include determining a probability index for each of a plurality of haplotypes of the VNTR for the second subject using a realignment of the short sequence reads realigned to the haplotypes. The method may include determining a VNTR status of the second subject based on the probability index for each of the plurality of haplotypes. In some embodiments, the method includes generating a user interface (UI) that includes a UI element that represents or includes a status of the VNTR.

[0006] In some embodiments, the haplotypes of the VNTR are associated with a disease (e.g., bipolar disorder or monogenic diabetes). In some embodiments, determining the VNTR haplotypes includes constructing or creating a database including the VNTR haplotypes. In some embodiments, determining the VNTR haplotypes includes, for each of the plurality of first samples, extracting long sequence reads from the plurality of long sequence reads of the first sample aligned to the VNTR in the reference. Determining the VNTR haplotypes may include rearranging the extracted long sequence reads into a left flanking region and a right flanking region of the VNTR to determine aligned long sequence reads. Determining the VNTR haplotypes may include determining the haplotypes of the plurality of haplotypes based on the aligned long sequence reads each having an alignment score exceeding an alignment threshold. At least one of the long sequence reads of the first sample is aligned to the VNTR and / or realigned to the left and right flanking regions that span the VNTR. In some embodiments, determining the haplotypes of the haplotypes of the VNTR comprises trimming the sequences of the aligned long sequence reads that each have an alignment score that exceeds an alignment threshold, and aligning to the left and right flanking regions to generate trimmed long sequence reads. Determining the haplotypes of the haplotypes of the VNTR may comprise determining the haplotypes of the haplotypes based on the trimmed long sequence reads.

[0007] In some embodiments, the first sample is homozygous for the VNTR. Determining the haplotypes of the multiple haplotypes may include determining only one haplotype of the multiple haplotypes based on the trimmed long sequence reads. Determining the only one haplotype may include clustering the trimmed long sequence reads into only one cluster. Clustering the trimmed long sequence reads into only one cluster may include clustering the trimmed long sequence reads into only one cluster based on the length of the trimmed long sequence reads. The clustering may include k-means clustering. Determining the only one haplotype may include determining only one haplotype based on the trimmed long sequence reads.

[0008] In some embodiments, the first sample is heterozygous for the VNTR. Determining the haplotypes of the plurality of haplotypes may include determining two haplotypes of the plurality of haplotypes of the VNTR based on the trimmed long sequence reads. Determining the two haplotypes may include clustering the trimmed long sequence reads into two clusters. Clustering the trimmed long sequence reads into two clusters may include clustering the trimmed long sequence reads into two clusters based on the length of the trimmed long sequence reads. The clustering may include k-means clustering. Determining the two haplotypes may include determining a first haplotype of the two haplotypes based on the trimmed long sequence reads in a first cluster of the two clusters. Determining the two haplotypes may include determining a second haplotype of the two haplotypes based on the trimmed long sequence reads in a second cluster of the two clusters. In some embodiments, the trimmed long sequence reads include a first plurality of trimmed long sequence reads and a second plurality of trimmed long sequence reads having different lengths. The different lengths differ by at least 5,000 base pairs. The first cluster may include all, substantially all, or a majority of the first plurality of trimmed long sequence reads. The second cluster may include all, substantially all, or a majority of the second plurality of trimmed long sequence reads.

[0009] In some embodiments, determining the haplotypes of the multiple haplotypes of the VNTR comprises determining a consensus sequence of the trimmed long sequence reads. In some embodiments, determining the consensus sequence of the trimmed long sequence reads comprises, for each position of the trimmed long sequence reads, a base that is not the most frequent base in the trimmed long sequence reads at that position: modifying the trimmed long sequence reads at that position using each of a plurality of operations (deletion operation, insertion operation, and substitution operation) independently; and determining the sum of distances (e.g., edit distances) between (i) the modified trimmed long sequence reads resulting from the operation on the trimmed long sequence reads at that base and (ii) the trimmed long sequence reads other than the trimmed long sequence read being modified. Determining a consensus sequence of the trimmed long sequence reads may include correcting the trimmed long sequence at bases using a plurality of operations that results in a minimum sum of distances (e.g., edit distance) among the plurality of operations, or replacing the trimmed long sequence read with a corrected trimmed long sequence read that corresponds to the minimum sum of distances (e.g., edit distance).

[0010] In some embodiments, determining the consensus sequence of the trimmed long sequence reads includes, for each corresponding position of the trimmed long sequence reads, determining the most frequent base among the bases of the trimmed long sequence reads at that position. Determining the consensus sequence of the trimmed long sequence reads may include, for each trimmed long sequence read having a base at a position that is not the most frequent base at that position, determining the sum of distances (e.g., edit distances) between (i) the modified trimmed long sequence read resulting from each of a plurality of independent operations (e.g., deletion operation, insertion operation, and substitution operation) on the trimmed long sequence read and (ii) the trimmed long sequence read other than the modified trimmed long sequence read. Determining the consensus sequence of the trimmed long sequence reads may include determining the minimum sum of distances (e.g., edit distances) among the sums of distances (e.g., edit distances). Determining the consensus sequence of the trimmed long sequence reads may include modifying the trimmed long sequence reads at the base with an operation that results in a minimum distance sum (e.g., edit distance), or replacing the trimmed long sequence read with the modified trimmed long sequence read that corresponds to the minimum distance sum (e.g., edit distance). In some embodiments, the multiple operations include deleting a base of the trimmed long sequence at the position. The multiple operations may include inserting the most frequent base at the position into the trimmed long sequence at the position. The multiple operations may include replacing a base of the trimmed long sequence at the position with the most frequent base at the position.

[0011] In some embodiments, the quality of the long sequence reads among the plurality of long sequence reads aligned to the VNTR in the reference meets a quality criterion. The quality of the plurality of haplotypes can meet a quality criterion.

[0012] In some embodiments, the status of the VNTR comprises a haplotype status of the VNTR. The haplotype status may comprise a haplotype, a haplotype length, and / or a confidence interval of the haplotype length. The status of the VNTR may comprise a genotype status of the VNTR. The genotype status may comprise a genotype, a haplotype length of the genotype, and / or a confidence interval of the length of each haplotype of the genotype. The confidence interval may comprise a minimum haplotype length and a maximum haplotype length.

[0013] In some embodiments, determining the VNTR status of the second subject comprises determining two or more haplotypes among the plurality of haplotypes whose probability index meets a probability criterion. Determining the VNTR status of the second subject may comprise determining the length of the determined two or more haplotypes. The shortest length of the haplotype may be the shortest length of the determined two or more haplotypes. The longest length of the haplotype may be the longest length of the determined two or more haplotypes. In some embodiments, the accuracy of the VNTR status is at least 60%.

[0014] In some embodiments, the probability index for each of the plurality of haplotypes of the VNTR comprises a probability for each of the plurality of haplotypes of the VNTR. The probability criterion can include a probability threshold.

[0015] In some embodiments, the plurality of long sequence reads comprises sequence reads each having a length of about 10,000 base pairs to about 20,000 base pairs. The plurality of long sequence reads can be generated by targeted sequencing or whole genome sequencing (WGS). The WGS can be clinical WGS (cWGS). The plurality of first subjects can include human subjects.

[0016] In some embodiments, the plurality of short sequence reads may include sequence reads each of about 100 base pairs to about 1000 base pairs in length. The plurality of short sequence reads may include paired-end sequence reads. The plurality of short sequence reads may include single-end sequence reads. The plurality of short sequence reads may be generated by targeted sequencing or whole genome sequencing (WGS). The WGS may be clinical WGS (cWGS). The second subject may include a human subject.

[0017] In some embodiments, the plurality of first subjects includes a second subject. The plurality of first samples can include a second sample. In some embodiments, the plurality of first samples and / or the second samples include cells, cell-free DNA, cell-free fetal DNA, amniotic fluid, a blood sample, a biopsy sample, or a combination thereof. The plurality of first samples can include at least 50 samples.

[0018] In some embodiments, each haplotype of the plurality of haplotypes of the VNTR comprises multiple copies of a repeat unit. The repeat unit may be more than 6 base pairs in length. The number of multiple copies may be at least three. In some embodiments, the sequences of two copies of the multiple copies of the repeat unit of a haplotype of the plurality of haplotypes differ at one or more differentiation positions. The sequences of two copies of the multiple copies of the repeat unit of a haplotype of the plurality of haplotypes have at least 80% sequence identity. The sequences of two copies of the multiple copies of the repeat unit of a haplotype of the plurality of haplotypes may be identical. In some embodiments, two haplotypes of the plurality of haplotypes of the VNTR comprise different numbers of copies of a repeat unit. In some embodiments, two haplotypes of the plurality of haplotypes of the VNTR comprise the same number of copies of a repeat unit. In some embodiments, the sequence of the copy of the repeat unit of one of the two haplotypes and the sequence of the copy of the repeat unit of the other of the two haplotypes differ at one or more differentiation positions. The sequences may have at least 80% sequence identity. The sequence of the copy of the repeat unit of one of the two haplotypes and the sequence of the copy of the repeat unit of the other of the two haplotypes can be identical.

[0019] Disclosed herein is a system for determining variable number tandem repeat (VNTR) status, such as genotyping VNTRs. In some embodiments, the system for determining VNTR status includes a non-transitory memory configured to store executable instructions and a plurality of haplotypes of VNTRs. The system can include a processor (e.g., a hardware processor or a virtual processor) in communication with the non-transitory memory, the processor being programmed with executable instructions to receive a plurality of short sequence reads generated from a test sample obtained from a test subject. The processor can be programmed with executable instructions to realign short sequence reads of the plurality of short sequence reads aligned to the VNTR to a haplotype for each of the plurality of haplotypes of the VNTR to generate a realignment. The processor can be programmed with executable instructions to determine a probability of each of the plurality of haplotypes for the test subject using the realignment of the short sequence reads realigned to the haplotypes. The processor can be programmed with executable instructions to determine a VNTR status of the test subject. In some embodiments, the processor is programmed with the executable instructions to determine a user interface (UI) that includes UI elements that represent or include a state of the VNTR.

[0020] In some embodiments, the haplotypes of the VNTR are associated with a disease (e.g., bipolar disorder or monogenic diabetes). In some embodiments, the haplotypes of the VNTR are determined using long sequence reads of a plurality of long sequence reads aligned to the VNTR in a reference (e.g., a reference human genome sequence, e.g., hg19 or hg38). In some embodiments, the plurality of long sequence reads can be generated from a plurality of reference samples obtained from a plurality of reference subjects. The plurality of haplotypes of the VNTR can be determined by extracting, for each of the plurality of samples, a long sequence read of a plurality of long sequence reads of a test sample aligned to the VNTR in the reference. The plurality of haplotypes of the VNTR can be determined by rearranging the extracted long sequence reads into the left flanking region and the right flanking region of the VNTR to determine aligned long sequence reads. The plurality of haplotypes of the VNTR can be determined by determining haplotypes of a plurality of haplotypes based on aligned long sequence reads each having an alignment score exceeding an alignment threshold. At least one long sequence read of the multiple long sequence reads of the test sample can be aligned to the VNTR. At least one long sequence read of the multiple long sequence reads of the test sample can be rearranged to the left flanking region, and the right flanking region spans the VNTR. In some embodiments, the haplotypes of the multiple haplotypes of the VNTR are determined by trimming the sequences of the aligned long sequence reads, each of which has an alignment score exceeding an alignment threshold, and aligning them to the left flanking region and the right flanking region to generate trimmed long sequence reads. The haplotypes of the multiple haplotypes of the VNTR can be determined by determining the haplotypes of the multiple haplotypes based on the trimmed long sequence reads.

[0021] In some embodiments, the reference sample is homozygous for the VNTR. The haplotypes of the multiple haplotypes of the VNTR can be determined to include one haplotype of the multiple haplotypes based on the trimmed long sequence reads. The single haplotype can be determined by clustering the trimmed long sequence reads into a single cluster. Clustering the trimmed long sequence reads into a single cluster can include clustering the trimmed long sequence reads into a single cluster based on the length of the trimmed long sequence reads. The clustering can include k-means clustering. The single haplotype can be determined by determining a single haplotype based on the trimmed long sequence reads.

[0022] In some embodiments, the reference sample is heterozygous for the VNTR. The haplotypes of the multiple haplotypes of the VNTR can be determined based on the trimmed long sequence reads to include two haplotypes of the multiple haplotypes. The two haplotypes can be determined by clustering the trimmed long sequence reads into two clusters. Clustering the trimmed long sequence reads into two clusters can include clustering the trimmed long sequence reads into two clusters based on the length of the trimmed long sequence reads. The clustering can include k-means clustering. The two haplotypes can be determined by determining a first haplotype of the two haplotypes based on the trimmed long sequence reads in a first cluster of the two clusters. The two haplotypes can be determined by determining a second haplotype of the two haplotypes based on the trimmed long sequence reads in a second cluster of the two clusters. In some embodiments, the trimmed long sequence reads include a first plurality of trimmed long sequence reads and a second plurality of trimmed long sequence reads having different lengths. The different lengths differ by at least 5,000 base pairs. The first cluster may include all, substantially all, or a majority of the first plurality of trimmed long sequence reads. The second cluster may include all, substantially all, or a majority of the second plurality of trimmed long sequence reads.

[0023] In some embodiments, a consensus sequence of the trimmed long sequence read is determined to determine the haplotypes of multiple haplotypes of VNTR. In some embodiments, the consensus sequence of the trimmed long sequence read is determined by having a base that is not the most frequent base in the trimmed long sequence reads at each position of the trimmed long sequence read, modifying the trimmed long sequence read at the position using each of a plurality of operations (e.g., deletion operation, insertion operation, and substitution operation), and determining the sum of edit distances between (i) the modified trimmed long sequence read resulting from the operation on the trimmed long sequence read at the base and (ii) the trimmed long sequence read other than the trimmed long sequence read being modified, and includes modifying the trimmed long sequence at the base using the operation of the plurality of operations that results in the minimum sum of edit distances among the plurality of operations, or by replacing the trimmed long sequence read with the modified trimmed long sequence read that corresponds to the minimum sum of edit distances.

[0024] In some embodiments, the consensus sequence of the trimmed long sequence reads is determined by: for each corresponding position of the trimmed long sequence reads, determining the most frequent base among the bases of the trimmed long sequence reads at that position; and for each of the trimmed long sequence reads having a base that is not the most frequent base at that position, for each of a plurality of operations (e.g., a deletion operation, an insertion operation, and a substitution operation), determining a sum of edit distances between (i) a modified trimmed long sequence read resulting from the operation on the trimmed long sequence read and (ii) a trimmed long sequence read other than the trimmed long sequence read being modified; determining a minimum sum of edit distances among the sums of edit distances; and modifying the trimmed long sequence read at the base by an operation that results in the minimum sum of edit distances or replacing the trimmed long sequence read with the modified trimmed long sequence read corresponding to the minimum sum of edit distances. In some embodiments, the manipulations include deleting a base of the trimmed long sequence at that position, inserting the most frequent base at that position into the trimmed long sequence at that position, and replacing the base of the trimmed long sequence at that position with the most frequent base at that position.

[0025] In some embodiments, the quality of the long sequence reads among the plurality of long sequence reads aligned to the VNTR in the reference meets a quality criterion. The quality of the plurality of haplotypes can meet a quality criterion.

[0026] In some embodiments, the status of the VNTR comprises a haplotype status of the VNTR. The haplotype status may comprise a haplotype, a length of the haplotype, and / or a confidence interval of the length of the haplotype. The status of the VNTR may comprise a genotype status of the VNTR, and the genotype status may comprise a genotype, a length of the haplotype of the genotype, and / or a confidence interval of the length of each haplotype of the genotype. The confidence interval may comprise a minimum length of the haplotype and a maximum length of the haplotype.

[0027] In some embodiments, determining the haplotype status of the VNTR of the subject comprises determining two or more haplotypes among the plurality of haplotypes whose probability index meets a probability criterion. Determining the haplotype status of the VNTR of the test subject may comprise determining the length of the determined two or more haplotypes. The shortest length of the haplotype may be the shortest length of the determined two or more haplotypes. The longest length of the haplotype may be the longest length of the determined two or more haplotypes. In some embodiments, the accuracy of the haplotype status is at least 60%.

[0028] In some embodiments, the probability index for each of the plurality of haplotypes of the VNTR comprises a probability for each of the plurality of haplotypes of the VNTR. The probability criterion can include a probability threshold.

[0029] In some embodiments, the plurality of long sequence reads comprises sequence reads each having a length of about 10,000 base pairs to about 20,000 base pairs. The plurality of long sequence reads can be generated by targeted sequencing or whole genome sequencing (WGS). The WGS can be clinical WGS (cWGS). The plurality of reference subjects can include human subjects.

[0030] In some embodiments, the plurality of short sequence reads comprises sequence reads each of which is between about 100 base pairs and about 1000 base pairs in length. The plurality of short sequence reads may comprise paired-end sequence reads. The plurality of short sequence reads may comprise single-end sequence reads. The plurality of short sequence reads may be generated by targeted sequencing or whole genome sequencing (WGS). The WGS may be clinical WGS (cWGS). The subject may comprise a human subject. The first sample may comprise cells, cell-free DNA, cell-free fetal DNA, amniotic fluid, a blood sample, a biopsy sample, or a combination thereof. The first subject may be a human subject.

[0031] In some embodiments, the plurality of reference subjects comprises a test subject. The plurality of reference samples can comprise a test sample. In some embodiments, the plurality of reference samples and / or test samples comprises cells, cell-free DNA, cell-free fetal DNA, amniotic fluid, a blood sample, a biopsy sample, or a combination thereof. The plurality of reference samples can comprise at least 50 samples.

[0032] In some embodiments, each haplotype of the plurality of haplotypes of the VNTR comprises multiple copies of a repeat unit. The repeat unit may be more than 6 base pairs in length. The number of multiple copies may be at least three. In some embodiments, the sequences of two copies of the multiple copies of the repeat unit of a haplotype of the plurality of haplotypes differ at one or more differentiation positions. The sequences of two copies of the multiple copies of the repeat unit of a haplotype of the plurality of haplotypes have at least 80% sequence identity. The sequences of two copies of the multiple copies of the repeat unit of a haplotype of the plurality of haplotypes may be identical. In some embodiments, two haplotypes of the plurality of haplotypes of the VNTR comprise different numbers of copies of a repeat unit. In some embodiments, two haplotypes of the plurality of haplotypes of the VNTR comprise the same number of copies of a repeat unit. In some embodiments, the sequence of the copy of the repeat unit of one of the two haplotypes and the sequence of the copy of the repeat unit of the other of the two haplotypes differ at one or more differentiation positions. The sequences may have at least 80% sequence identity. The sequence of the copy of the repeat unit of one of the two haplotypes and the sequence of the copy of the repeat unit of the other of the two haplotypes can be identical.

[0033] The details of one or more implementations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will become apparent from the specification, drawings, and claims. Neither this summary nor the following detailed description is intended to define or limit the scope of the inventive subject matter. [Brief description of the drawings]

[0034] [Figure 1] FIG. 1 shows a non-limiting exemplary illustration of the VNTRs in the reference sequence and five samples. [Diagram 2] FIG. 1 shows a non-limiting, exemplary schematic of construction of a VNTR database from long reads. [Figure 3A]FIG. 1 is a non-limiting, illustrative schematic diagram of generating haplotypes from multiple long reads. [Figure 3B] FIG. 1 is a non-limiting, illustrative schematic diagram of generating haplotypes from multiple long reads. [Figure 4] FIG. 1 is a non-limiting, exemplary schematic diagram of genotyping VNTRs on short reads. [Diagram 5] FIG. 1 is a flow diagram showing an exemplary method for determining VNTR status (eg, VNTR haplotype or genotype). [Figure 6] FIG. 1 is a block diagram of an exemplary computing system configured to perform a determination of VNTR status (e.g., VNTR haplotype or genotype).

[0035] Throughout the drawings, reference numbers may be reused to indicate correspondence between referenced elements. The drawings are provided to illustrate example embodiments described herein and are not intended to limit the scope of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0036] In the following detailed description, reference is made to the accompanying drawings, which form a part of this specification. In the drawings, like symbols typically identify like components unless the context dictates otherwise. The exemplary embodiments described in the detailed description, drawings, and claims are not intended to be limiting. Other embodiments may be utilized, and other changes may be made, without departing from the spirit or scope of the subject matter presented herein. It is readily understood that aspects of the present disclosure, as generally described herein and illustrated in the drawings, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations, all of which are expressly contemplated herein and make part of the disclosure herein.

[0037] All patents, published patent applications, other publications, and sequences from GenBank and other databases referenced herein are hereby incorporated by reference in their entirety with respect to the relevant art.

[0038] Disclosed herein is a method for determining variable number tandem repeat (VNTR) status, such as genotyping VNTR. In some embodiments, the method for determining VNTR status is under the control of a processor (e.g., a hardware processor or a virtual processor) and includes receiving a plurality of long sequence reads generated from a plurality of first samples obtained from a plurality of first subjects. The method may include determining a plurality of haplotypes of the VNTR using a long sequence read from the plurality of long sequence reads aligned to the VNTR in a reference (e.g., a reference human genome sequence, e.g., hg19 or hg38). The method may include receiving a plurality of short sequence reads generated from a second sample obtained from a second subject. The method may include rearranging the short sequence reads from the plurality of short sequence reads aligned to the VNTR into haplotypes to generate rearrangements for each of the plurality of haplotypes of the VNTR. The method may include determining a probability index for each of a plurality of haplotypes of the VNTR for the second subject using a realignment of the short sequence reads realigned to the haplotypes. The method may include determining a VNTR status of the second subject based on the probability index for each of the plurality of haplotypes. In some embodiments, the method includes generating a user interface (UI) that includes a UI element that represents or includes a status of the VNTR.

[0039] Disclosed herein is a system for determining variable number tandem repeat (VNTR) status, such as genotyping VNTRs. In some embodiments, the system for determining VNTR status includes a non-transitory memory configured to store executable instructions and a plurality of haplotypes of VNTRs. The system can include a processor (e.g., a hardware processor or a virtual processor) in communication with the non-transitory memory, the processor being programmed with executable instructions to receive a plurality of short sequence reads generated from a test sample obtained from a test subject. The processor can be programmed with executable instructions to realign short sequence reads of the plurality of short sequence reads aligned to the VNTR to a haplotype for each of the plurality of haplotypes of the VNTR to generate a realignment. The processor can be programmed with executable instructions to determine a probability of each of the plurality of haplotypes for the test subject using the realignment of the short sequence reads realigned to the haplotypes. The processor can be programmed with executable instructions to determine a VNTR status of the test subject. In some embodiments, the processor is programmed with the executable instructions to determine a user interface (UI) that includes UI elements that represent or include a state of the VNTR.

[0040] Genotyping variable number tandem repeats Disclosed herein are genotyping agents that significantly improve variable number tandem repeat (VNTR) genotyping performance on short read sequencing data (e.g., sequencing data generated by sequencing methods such as sequencing by synthesis). For example, the improvement is made by utilizing a pre-built VNTR database. As another example, the improvement is made by optimizing current genotyping methods for low complexity regions. The present disclosure also provides a workflow that can build a population VNTR database from, for example, Pacific Biosciences of California, Inc. (PacBio, Menlo Park, CA) HiFi data.

[0041] VNTRs can be repeat sequences with repeats of more than 6 base pairs (bps) and repeat regions of more than 80% pure (less than 20% mismatch for exact repeats). Structural variations (SVs) of VNTRs include insertions / deletions of repeated sequences. Mutations can be highly population-specific. Some VNTRs are known to cause genetic diseases such as bipolar disorder and monogenic diabetes. VNTRs account for a significant proportion of sample-to-sample variation. Approximately half of all SVs per individual (>10k) can be classified as VNTRs. On average, an individual has about 2.2 megabase pairs (Mbps) of deletion sequence and about 5.7 Mbps of insertion sequence in VNTRs.

[0042] FIG. 1 shows a non-limiting exemplary illustration of the VNTRs in the reference sequence and five samples. The VNTRs in the reference human genome GRCh38 are located at chr1:3428147-3428340 (FIG. 1, top left panel). The repeat unit has a length of 48 bps. The reference sequence of the repeat unit is

[0043] [Table 1] The different copies of the repeat unit in the VNTR (intra- or inter-haplotype) may differ in particular in the three bases that are bold and underlined. The three bases may be G, G, and A, respectively, in the first type or sequence of the repeat unit, G, G, and G, respectively, in the second type or sequence of the repeat unit, A, G, and A, respectively, in the third type or sequence of the repeat unit, and G, A, and G, respectively, in the fourth type or sequence of the repeat unit (Figure 1, top right panel). The VNTR contains four copies of the repeat unit in GRCh38 (Figure 1, bottom panel). The four copies include two copies of the first type followed by two copies of the second type (Figure 1, bottom panel). The five samples contained 3, 5, 7, 7, and 10 copies of the repeat unit, respectively. In the case of sample NA19240, an African subject, the VNTR contained one copy of the first type followed by two copies of the second type. For European subject sample NA12878, the VNTR contained one copy of the first type, three copies of the second type, and one copy of the first type. For European subject sample NA24385, the VNTR contained one copy of the first type, one copy of the second type, two copies of the first type, two copies of the second type, and one copy of the third type. For East Asian subject sample HG00597, the VNTR contained three copies of the second type, one copy of the first type, and three copies of the second type. For African subject sample HG03453, the VNTR contained one copy of the first type, two copies of the second type, one copy of the fourth type, one copy of the first type, one copy of the second type, one copy of the fourth type, and three copies of the second type.

[0044] VNTR genotyping is missing in short read pipelines. Short reads often fail to cover the full length of most VNTRs. Short reads are also referred to herein as short sequence reads. Approximately 29% of VNTRs have additional repeats with a total length of 150 bps or more in an individual. Due to the repetitive nature of VNTRs, it is extremely difficult to correctly reconstruct the haplotype of VNTRs from short reads. VNTR detection power is very low in short read pipelines. For example, DRAGEN v3.4 detection power for VNTRs is less than 20%.

[0045] Disclosed herein is a VNTR genotyping method that includes one or more of the following: First, the method may include (or may be performed by a VNTR genotyping analyst) constructing a database of common VNTR haplotypes in a population. Long reads (e.g., PacBio HiFi reads), which may be of high accuracy, may be used to construct a database of common VNTR haplotypes. In this specification, long reads are also referred to as long sequence reads. Second, the method may include extracting short reads from a target VNTR region that are generated by a sequencing method that includes sequencing by synthesis, such as short reads generated using a sequencing instrument from Illumina, Inc. (San Diego, CA). These extracted short reads may be realigned to each haplotype sequence in the database. Third, the method may include deriving the most likely VNTR haplotype (and therefore genotype) from the realignment. VNTRs usually have differences between the repeat units of different haplotypes. The differences between the repeat units may be referred to as differentiation bases in this specification. The most likely haplotypes (and therefore genotypes) can be determined from these differentiation positions (see, for example, the top right panel of Figure 1 for distinguishing bases within and between haplotypes of VNTRs).

[0046] The method may include constructing a VNTR database from long reads (e.g., PacBio HiFi reads), which may be of high accuracy. PacBio HiFi reads are long enough (average 15 kb) to span the entire length of most VNTRs. Long read sequencing is limited by DNA input and cost and cannot be performed on a large scale. However, as described herein, it is possible to sequence several samples (e.g., hundreds of samples) to construct a database.

[0047] FIG. 2 shows an example of constructing a VNTR database from long reads. For each sample, long reads (e.g., PacBio HiFi reads) can be extracted from the target VNTR region. The reads can be aligned to the left and right flanking regions of the VNTR reads with good alignment to the flanking regions on both sides. The flanking regions can be trimmed from the reads. It can be distinguished whether the trimmed reads are derived from one haplotype or two haplotypes. For example, if the reads can be clustered into two clusters (e.g., clustered k-means), the sample is heterozygous. Otherwise, the sample is homozygous. Haplotype(s) can be assembled from the differentiated reads. For example, if the reads can be clustered into two clusters, the reads within each cluster can be assembled into haplotypes. The reads within the two clusters can be assembled into two haplotypes. If the reads cannot be clustered into two clusters, the reads can be assembled into haplotypes. The resulting database of recurrent haplotypes can include "star alleles." For example, the haplotypes in the database may contain distinguishing bases that can be used to distinguish the haplotypes. The repeat units in the haplotypes may contain distinguishing bases that can be used to distinguish the haplotypes. The three haplotypes shown in Figure 2 have 4, 5, and 6 copies of the repeat unit, respectively.

[0048] At low sequencing error rates (e.g., less than 1%), it should be rare to observe different bases at each position. Thus, the method illustrated with reference to Figures 3A-3B can be used to correct sequencing errors and assemble haplotypes. For each position, the base with the highest proportion (most common) among these reads (e.g., trimmed reads) is labeled as the "consensus base" (also referred to herein as the "true base"). For each read (e.g., trimmed reads) that has a base different from the "consensus base", the following three actions (or operations) are performed independently: delete a base, add a "consensus base", or change a base to the "consensus base". The distance (e.g., edit distance) between the corrected read (e.g., trimmed read to corrected read) generated by each action (or operation) and each other read (e.g., trimmed read) can be calculated, and these distances can be summed. The read can be corrected with the action (or operation) that has the smallest sum of distances. The sum of the distances of an action and the sum of the distances of another action may be the same (tie), although this is unlikely. If the sum of the distances for an action and the sum of the distances for the other action are the same, one of the two actions can be selected, for example, randomly. This process can be repeated for each read, then each position (or each position, then each read) until all reads have the same sequence. With a low sequencing error rate (e.g., less than 1%), only a small fraction of the bases at each position should differ from the "consensus base" at that position.

[0049] 3A-3B show an example of generating haplotypes from multiple long reads. Scan from the beginning of each long read. If at a position the base is not 100% the same among all reads, assume the most prevalent base among all reads as the "consensus base" or "true base". Then try to fix the reads with different bases as the "consensus base". Continue scanning and fixing bases until the end of each read is reached. In the example shown in FIG. 3A-3B, three reads have the sequences ATCG, ATCT, and ATTCG. The prevalent base at the third position is cytosine (C). The third base in read 3 (bold and underlined) is thymine (T) and is different from the "consensus base". The following three actions (or operations) can be performed independently on the third base in read 3: delete the base, add the "consensus base", or change the base to the "consensus base". If the third base is deleted, the distances (e.g., edit distances) between the modified read 3 with a sequence of ATCG and reads 1 and 2 are 0 and 0, respectively. Thus, the sum of the distances for the action of deleting a base is 0. If the third base is changed to a "consensus base", the modified read 3 has a sequence of ATCCG. The distances between the modified read 3 and reads 1 and 2 are 1 and 1, respectively. Thus, the sum of the distances for the action of changing a base to a "consensus base" is 2. If a "consensus base" is added or inserted at the third position of read 3, the modified read 3 has a sequence of ATCTCG. The distances between the modified read 3 and reads 1 and 2 are 2 and 2, respectively. Thus, the sum of the distances for the action of adding a "consensus" base is 4. The resulting sum of the distances is the smallest for the action of deleting a base among the three actions, so that action is selected. The sum of the distances of an action and the sum of the distances of another action may be the same (tie), although this is unlikely. If the sum of the distances for the action is the same as the sum of the distances for the other actions, one of the two actions can be selected, for example randomly. Modified lead 3 is fixed in the third position. This process is repeated until the end of all leads is reached.

[0050] Figure 4 shows an example of genotyping VNTR on short reads. Short reads (e.g., Illumina read pairs) can be extracted from the BAM file at the location of the VNTR. Each read can be realigned to each of the haplotypes in the VNTR haplotype database. In some embodiments, no gaps are allowed in the realignment. Each haplotype / read pair combination can be scored. The VNTR genotyping model used in some embodiments for scoring is as follows: Lead R with L base i For a given haplotype H1, its probability is:

[0051]

number

[0052]

number

[0053] In some pure repeats that do not have "star alleles", more than one best genotype may be observed. Fragment length information may help narrow down the possible genotypes, but cannot completely eliminate this ambiguity. In some embodiments, confidence intervals (CIs) are reported as estimates of the VNTR length. A minimal set that can cover all these equally best genotypes can first be derived. Using this minimal set, CIs can be reported for each haplotype as [shortest allele, longest allele]. For example, a VNTR can be genotyped as having two possible genotypes with lengths of 50 / 60 and 50 / 80, and reported as having CIs of [50,50] and [60,80].

[0054] VNTR genotyping accuracy. Improved VNTR genotyping accuracy was obtained using the genotyping method described herein (Table 1). 60 samples were sequenced on a PacBio HiFi. An Illumina NovaSeq 6000 was used to test the genotyping accuracy of VNTRs. A total of 1,000 VNTRs were tested in this analysis. The genotyping method described herein had an accuracy of 62%, 71%, and 78% as measured by correct genotype, repeat length, and repeat length CI. Dragen v3.4 large variant detection accuracy as measured by repeat length was 16%. In paragraph (Chen, S., et al. Paragraph: a graph-based structural variant genotyper for short-read sequence data. Genome Biol 20, 291 (2019); the contents of which are incorporated herein by reference in their entirety), the genotyping accuracy of large variants, as measured by whether the repeat presence was correctly genotyped, was 38%.

[0055] [Table 2]

[0056] "Star alleles" are absent for some pure repeats. PacBio HiFi data can have low quality in regions with large homopolymers. VNTR genotyping accuracy was improved by limiting the analysis to a whitelist containing high-quality assembled haplotypes (Table 2). Filtering criteria used to generate the whitelist included homopolymer length, repeat unit purity, haplotype assembly quality, and repeat variability in the population. The performance shown in Table 2 was based on a whitelist covering 60% of the VNTRs originally tested. The improvement shown is relative to the performance of the originally tested VNTRs shown in Table 1. With the whitelist, VNTR genotyping performance was improved.

[0057] [Table 3]

[0058] Determination of VNTR status FIG. 5 is a flow diagram illustrating an exemplary method 500 of determining a VNTR status (e.g., a VNTR haplotype or genotype), such as genotyping a VNTR. The method 500 may be embodied in a set of executable program instructions stored on a computer-readable medium, such as one or more disk drives, of a computing system. For example, a computing system 600, illustrated in FIG. 6 and described in more detail below, may execute a set of executable program instructions to perform the method 500. When the method 500 is initiated, the executable program instructions may be loaded into a memory, such as a RAM, and executed by one or more processors of the computing system 600. Although the method 500 is described with respect to the computing system 600 illustrated in FIG. 6, the description is merely exemplary and not intended to be limiting. In some embodiments, the method 500, or portions thereof, may be executed by multiple computing systems, serially or in parallel.

[0059] After the method 500 begins at block 504, the method 500 proceeds to block 508, where a computing system (e.g., computing system 600 described with reference to FIG. 6) receives a plurality of long sequence reads generated from a plurality of first samples (or reference samples) obtained from a plurality of first subjects (or reference subjects). The long sequence reads are also referred to herein as long reads. The long sequence reads can be, for example, PacBio HiFi reads. The long sequence reads can be, for example, 5 kilobase pairs (kbps), 6 kbps, 7 kbps, 8 kbps, 9 kbps, 10 kbps, 11 kbps, 12 kbps, 13 kbps, 14 kbps, 15 kbps, 20 kbps, 25 kbps, 30 kbps, or more. For example, the plurality of long sequence reads includes sequence reads that are between about 10 kbps and about 20 kbps. Each one, one or more of the plurality of long sequence reads (or long sequence reads aligned to the left and right flanking regions of the VNTR in block 512) may have a high accuracy, such as 95%, 96%, 97%, 98%, 99%, or more. The plurality of first samples may include at least 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 1000, or more samples. The plurality of long sequence reads may be generated by targeted sequencing or whole genome sequencing (WGS). The WGS may be clinical WGS (cWGS). The plurality of first subjects may include human subjects.

[0060] From block 508, the method 500 proceeds to block 512, where the computing system determines a plurality of haplotypes of the VNTR (or a database of haplotypes of the VNTR) using the long sequence reads aligned to the VNTR in the reference among the plurality of long sequence reads (see FIG. 2 and accompanying description). The reference can be, for example, a reference human genome sequence such as hg19 or hg38. One haplotype of the plurality of haplotypes of the VNTR is associated with the disease. Non-limiting examples of the disease include bipolar disorder, MCKD1, stroke, CAD, FSHD, ADHD, Parkinson's disease, diffuse panbronchiolitis (DPB), monogenic diabetes, T1D, T2D, obesity, OCD, ADHD, osteochondritis dissecans, Kawasaki disease, ATF in stroke, BPSD, Alzheimer's disease, OCD, anxiety, schizophrenia, metastatic colorectal cancer, Kawasaki disease, or progressive myoclonic epilepsy 1A. The VNTR may be in a coding or non-coding region. The VNTR may be in a 5' untranslated region (UTR), promoter, intron, or 3'UTR. The gene containing or affected by a VNTR can be, for example, PER3, MUC1, IL1RN, DUX4, DAT1, MUC21, CEL, INS, DRD4, ACAN, ZFHX3, GP1BA, SERT, SERT, HIC1, MMP9, CSTB, or MAOA.

[0061] Each haplotype of the multiple haplotypes of the VNTR can include multiple copies of a repeat unit, which can be (or can be at least or more than) 6 bps, 7 bps, 8 bps, 9 bps, 10 bps, 11 bps, 12 bps, 13 bps, 14 bps, 15 bps, 16 bps, 17 bps, 18 bps, 19 bps, 20 bps, 30 bps, 40 bps, 50 bps, 60 bps, 70 bps, 80 bps, 90 bps, 100 bps, 150 bps, 200 bps, or more in length. The number of multiple copies can be (or can be at least or more than) 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 300, 400, 500 or more. The pathogenic copy number can be equal to, greater than, or less than the copy number in the reference.

[0062] The two copies of the repeat unit of the haplotype may include differentiation of bases at a specific position (herein referred to as differentiation position). For example, the sequences of the two copies of the multiple copies of the repeat unit of the haplotype of the multiple haplotypes differ at one or more differentiation positions (e.g., 2, 3, 4, 5, 10, 20 or more positions). The star alleles of the haplotype may include differentiation of bases at these positions. The star alleles may include positions that can help distinguish two or more haplotypes from each other. The sequences of the two copies of the multiple copies of the repeat unit of the haplotype have (or have at least) 70%, 75%, 80%, 85%, 90%, 95%, 99% or more sequence identity. The sequences of the two copies of the multiple copies of the repeat unit of the haplotype of the multiple haplotypes may be identical. In some embodiments, the two haplotypes of the multiple haplotypes of the VNTR include different numbers of copies of the repeat unit.

[0063] The copies of the repeat unit in each of the two haplotypes may contain base differentiation at certain positions (referred to herein as differentiation positions). For example, two haplotypes of a VNTR include the same number of copies of the repeat unit. The sequence of the copy of the repeat unit in one of the two haplotypes and the sequence of the copy of the repeat unit in the other of the two haplotypes may differ at one or more differentiation positions. Star alleles of a haplotype may contain base differentiation at these positions. The sequences of the two copies may have (or at least have) 70%, 75%, 80%, 85%, 90%, 95%, 99% or more sequence identity. The sequence of the copy of the repeat unit in one of the two haplotypes and the sequence of the copy of the repeat unit in the other of the two haplotypes may be identical.

[0064] To determine multiple haplotypes of the VNTR, the computing system can construct or create a database including multiple haplotypes of the VNTR. To determine multiple haplotypes of the VNTR, the computing system can extract, for each of multiple first samples, long sequence reads from multiple long sequence reads of the first sample aligned to the VNTR in the reference. The computing system can realign the extracted long sequence reads to the left flanking region and the right flanking region of the VNTR to determine aligned long sequence reads. The aligned long sequence reads can be long sequence reads aligned with the left flanking region and the right flanking region. The aligned long sequence reads can be, for example, long sequence reads having related alignments to the left flanking region and the right flanking region. The computing system can determine the haplotypes of the multiple haplotypes based on the aligned long sequence reads each having an alignment score exceeding an alignment threshold (e.g., 80%, 85%, 90%, 95%, 99% or 100% sequence identity). The alignment threshold can be pre-determined. In some embodiments, the alignment threshold is determined using a large number of samples, such as 100, 200, 300, 400, 500, 1000, 2000, 3000, 4000, 5000, or more or less samples. At least one long sequence read of the multiple long sequence reads of the first sample can be aligned to the VNTR. At least one long sequence read of the multiple long sequence reads of the first sample can be rearranged to the left flanking region, and the right flanking region spans the VNTR.In some embodiments, to determine the haplotypes of the multiple haplotypes of the VNTR, the computing system can trim the sequences aligned to the left flanking region and the right flanking region of the aligned long sequence reads, each of which has an alignment score exceeding an alignment threshold, to generate trimmed long sequence reads. The computing system can determine the haplotypes of the multiple haplotypes based on the trimmed long sequence reads.

[0065] The first sample may be heterozygous for the VNTR. The plurality of trimmed long sequence reads may be clustered into two clusters. To determine the haplotypes of the plurality of haplotypes, the computing system may determine the two haplotypes of the plurality of haplotypes of the VNTR based on the trimmed long sequence reads. To determine the two haplotypes, the computing system may cluster the trimmed long sequence reads into two clusters. The computing system may cluster the trimmed long sequence reads into two clusters based on the length of the trimmed long sequence reads. The computing system may cluster the trimmed long sequence reads into two clusters using a clustering method. The clustering method may include k-means clustering (e.g., k is equal to 2). The clustering method may include hierarchical clustering. The clustering method may be performed using, for example, a connectivity model, a centroid model, a distribution model, or a density model. The computing system may determine a first haplotype of the two haplotypes based on the trimmed long sequence reads in a first cluster of the two clusters. The computing system can determine a second haplotype of the two haplotypes based on the trimmed long sequence reads in a second cluster of the two clusters. In some embodiments, the trimmed long sequence reads include a first plurality of trimmed long sequence reads and a second plurality of trimmed long sequence reads having different lengths. The cluster can have a length (e.g., an average length of the trimmed long sequence reads in a cluster) of about 1 kilobase pair (kbps), 2 kbps, 3 kbps, 4 kbps, 5 kbps, 10 kbps, 15 kbps, 20 kbps, 30 kbps, 40 kbps, 50 kbps, 100 kbps, or more.The lengths of the two clusters (e.g., the average length of the trimmed long sequence reads in each cluster) may differ by about, or at least, 1 kbps, 2 kbps, 3 kbps, 4 kbps, 5 kbps, 10 kbps, 15 kbps, 20 kbps, 30 kbps, 40 kbps, 50 kbps, 100 kbps, or more. For example, the length of one cluster may be about 5 kbps and the length of the other cluster may be about 30, and the lengths of the two clusters may differ by about 25 kbps. The first cluster may include all, substantially all (e.g., 90%, 95%, 99%, or more), or a majority (e.g., 51%, 60%, 70%, 80%, or more) of the first plurality of trimmed long sequence reads. The second cluster may include all, substantially all, or a majority of the second plurality of trimmed long sequence reads.

[0066] The first sample may be homozygous for the VNTR. The first sample may be heterozygous for the VNTR. The multiple trimmed long sequence reads may not be clustered into two clusters. To determine a haplotype among the multiple haplotypes, the computing system may determine only one haplotype among the multiple haplotypes based on the trimmed long sequence reads. To determine only one haplotype, the computing system may cluster the trimmed long sequence reads into only one cluster. For example, the separation between the trimmed long sequence reads may be small enough such that the trimmed long sequence reads are not clustered into two clusters and / or are clustered into only one cluster. The computing system may cluster the trimmed long sequence reads into only one cluster based on the length of the trimmed long sequence reads. The computing system may cluster the trimmed long sequence reads into only one cluster using a clustering method. The clustering method may include k-means clustering (e.g., k is equal to 2). For example, the length difference of the trimmed long sequence reads can be small enough so that the trimmed long sequence reads are not clustered into two clusters and / or are clustered into only one cluster using k-means clustering where k is 2. The clustering method can include hierarchical clustering. The clustering method can be performed using, for example, a connectivity model, a centroid model, a distribution model, or a density model. The computing system can determine only one haplotype based on the trimmed long sequence reads.

[0067] To determine the haplotypes of multiple haplotypes of the VNTR, the computing system can determine a consensus sequence of the trimmed long sequence reads (see FIG. 2 and accompanying description). In some embodiments, to determine the consensus sequence of the trimmed long sequence reads, the computing system can, for each position of each of the trimmed long sequence reads, use a base that is not the most frequent base in the trimmed long sequence reads at that position (traversing the (corresponding) positions of all the trimmed long sequence reads before proceeding to the next position of all the trimmed long sequence reads, or traversing each position of the trimmed long sequence read before traversing each position of another trimmed long sequence read). The computing system can independently use each of the multiple operations (e.g., a deletion operation, an insertion operation, and a substitution operation) to modify the trimmed long sequence read at the position and determine a sum of distances (e.g., an edit distance) between (i) the modified trimmed long sequence read resulting from the operation on the trimmed long sequence read at the base and (ii) the trimmed long sequence read other than the trimmed long sequence read being modified. The computing system can modify the trimmed long sequence at the base using an operation of the multiple operations that results in a minimum sum of distances (e.g., an edit distance) among the multiple operations. Alternatively, the computing system can replace the trimmed long sequence read with the modified trimmed long sequence read that corresponds to the minimum sum of distances (e.g., an edit distance). In some embodiments, the multiple operations include deleting a base of the trimmed long sequence at the position. The multiple operations can include inserting the most frequent base at the position into the trimmed long sequence at the position. The multiple manipulation may include replacing the base of the trimmed long sequence at that position with the most frequent base at that position.

[0068] In some embodiments, to determine a consensus sequence of the trimmed long sequence reads, the computing system can perform the following for each position of the trimmed long sequence read (traversing the (corresponding) positions of all trimmed long sequence reads before proceeding to the next position of all trimmed long sequence reads, or traversing each position of the trimmed long sequence read before traversing each position of another trimmed long sequence read). The computing system can determine the most frequent base among the bases of the trimmed long sequence reads at that position. For each trimmed long sequence read having a base at a position that is not the most frequent base at that position, the computing system can determine the sum of distances (e.g., edit distances) between (i) the corrected trimmed long sequence read resulting from each of the multiple operations on the trimmed long sequence read and (ii) the trimmed long sequence read other than the trimmed long sequence read being corrected. The computing system can determine the minimum sum of distances (e.g., edit distances) among the sums of distances (e.g., edit distances). The computing system can modify the trimmed long sequence reads read at the base by the operation that results in the minimum sum of distances (e.g., edit distance). Alternatively, the computing system can replace the trimmed long sequence reads with the modified trimmed long sequence reads that correspond to the minimum sum of distances (e.g., edit distance).

[0069] In some embodiments, long sequence reads and / or haplotypes can be filtered out or discarded based on quality criteria such as sequencing quality, homopolymer length, repeat unit purity, haplotype ensemble quality, and repeat variability in the population. The quality of the multiple long sequence reads (before or after the long sequence reads are aligned to the VNTRs in the reference) can meet one or more quality criteria. Long sequence reads that do not meet one or more quality criteria (or filtering criteria) can be filtered out or discarded from use to determine haplotypes of the multiple haplotypes. Quality criteria can include sequencing quality (e.g., base call accuracy such as Phred quality score) and homopolymer length. For example, long sequence reads can have low quality in regions with large homopolymers. Such low quality long sequence reads can be discarded from determining haplotypes of the multiple haplotypes. In some embodiments, the quality of the multiple haplotypes meets one or more quality criteria (or filtering criteria). Haplotypes that do not meet one or more quality criteria can be filtered out or discarded. The quality criteria may include, for example, homopolymer length, purity of repeat units, quality of haplotype collection, and / or repeat variability within the population. The remaining haplotypes may be a whitelist of haplotypes. The whitelist of haplotypes may be used in one or more subsequent blocks of method 500. The whitelist of haplotypes may include, for example, about 50%, 60%, 70%, or 80% of all haplotypes initially determined (including both the whitelist of haplotypes and the excluded or discarded haplotypes). The plurality of haplotypes may include the whitelist of haplotypes but not the excluded haplotypes.

[0070] In some embodiments, instead of receiving a plurality of long sequence reads at block 508 and determining a plurality of haplotypes of the VNTR using the plurality of long sequence reads at block 512, the computing system receives a plurality of haplotypes of the VNTR (or a database of haplotypes of the VNTR). Alternatively or additionally, the plurality of haplotypes of the VNTR (or a database of haplotypes of the VNTR) is stored in a memory of the computing system. The plurality of haplotypes can be determined using long sequence reads of the plurality of long sequence reads aligned to the VNTR in the reference, as described with reference to block 512.

[0071] Method 500 proceeds from block 512 to block 516, where a computing system receives a plurality of short sequence reads generated from a second sample (or test sample) obtained from a second subject. The short sequence reads are also referred to herein as short reads. The short sequence reads can be, for example, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1250, 1500, 1750, 2000 or more base pairs (bps) in length, respectively. For example, the short sequence reads are each about 100 bp to about 1000 bp in length. The short sequence reads can include paired-end sequence reads. The sequence reads can include single-end sequence reads. The short sequence reads can be generated by targeted sequencing. The short sequence reads can be generated by whole genome sequencing (WGS). The short sequence reads can be generated by whole genome sequencing (WGS). The WGS can be clinical WGS (cWGS). The second sample can include cells, cell-free DNA, cell-free fetal DNA, amniotic fluid, a blood sample, a biopsy sample, or a combination thereof. The second subject can be a human subject. In some embodiments, the plurality of first subjects includes a second subject. In some embodiments, the plurality of first samples includes a second sample.

[0072] The computing system can store the sequence reads in memory. The computing system can load the sequence reads into memory. The sequence reads can be generated by techniques such as sequencing by synthesis, sequencing by ligation, or sequencing by ligation. The sequence reads can be generated using instruments such as MINISEQ, MISEQ, NEXTSEQ, HISEQ, and NOVASEQ sequencing instruments from Illumina, Inc. (San Diego, CA).

[0073] From block 516, the method 500 proceeds to block 520, where the computing system aligns, for each of a plurality of haplotypes of the VNTR, a short sequence read of the plurality of short sequence reads (re)aligned to the VNTR to the haplotype to generate a realignment. In some embodiments, no gaps are allowed in the realignment. In some embodiments, gaps are allowed in the realignment. Computing systems include Burrows-Wheeler Aligner (BWA), iSAAC, BarraCUDA, BFAST, BLASTN, BLAT, Bowtie, CASHX, Cloudburst, CUDA-EC, CUSHAW, CUSHAW2, CUSHAW2-GPU, drFAST, ELAND, ERNE, GNUMAP, GEM, GensearchNGS, GMAP and GSNAP, Geneious Assembler, LAST, MAQ, mrFAST and mrsFAST, MOM, MOSAIK, MPscan, Novoaligh & NovoalignCS, NextGENe, Omixon, PALMapper, Partek, PASS, PerM, PRIMEX, QPalma, RazerS, REAL, cREAL, RMAP, rNA, RT Short sequence reads can be (re)aligned to haplotypes using aligners or alignment methods such as Investigator, Segemehl, SeqMap, Shrec, SHRiMP, SLIDER, SOAP, SOAP2, SOAP3 and SOAP3-dp, SOCS, SSAHA and SSAHA2, Stampy, SToRM, Subread and Subjunc, Taipan, UGENE, VelociMapper, XpressAlign, and ZOOM.

[0074] Method 500 proceeds from block 520 to block 524, where the computing system uses the realignment of the short sequence reads (re)aligned to haplotypes to determine a probability index for each of the plurality of haplotypes of the VNTR for the second subject. The computing system can determine two or more haplotypes of the plurality of haplotypes, each haplotype having a probability index that meets a probability criterion. The probability index for each of the plurality of haplotypes of the VNTR includes a probability of each of the plurality of haplotypes of the VNTR. The probability criterion can include a probability threshold (e.g., 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95%). The probability threshold can be predetermined. In some embodiments, the probability threshold is determined using a large number of samples, such as 100, 200, 300, 400, 500, 1000, 2000, 3000, 4000, 5000, or more or less samples. The probability criterion can include the highest probability (or the highest several probabilities, such as the highest 2, 3, 4, 5, or more probabilities) among each of the multiple haplotypes.

[0075] Alternatively or additionally, the computing system uses realignment of the short sequence reads (re)aligned to haplotypes to determine a probability index of each pair of haplotypes of the plurality of haplotypes of the VNTR for the second subject. The computing system can determine one or more haplotype pairs of the plurality of haplotypes, each pair having a probability index that meets a probability criterion. The probability index of each pair of haplotypes of the plurality of haplotypes of the VNTR can include a probability of each pair of haplotypes of the VNTR. The probability criterion can include a probability threshold (e.g., 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95%). The probability criterion can include the highest probability (or the highest few probabilities, such as the highest 2, 3, 4, 5, or more probabilities) of each haplotype pair of the plurality of haplotypes.

[0076] For example, a score (e.g., a probability index) for each haplotype / sequence read combination can be determined. The VNTR genotyping model used in some embodiments for scoring is as follows: Lead R with L base i For a given haplotype H1, its probability is:

[0077]

number

[0078]

number

[0079] The method 500 proceeds from block 524 to block 528, where the computing system determines a VNTR status of the second subject based on the probability index of each of the plurality of haplotypes. The accuracy of the VNTR status may be at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95%. The VNTR status may include a haplotype status of the VNTR. The haplotype status may include a haplotype, a length of the haplotype, and / or a confidence interval (CI) of the haplotype length. The confidence interval may include a minimum length of the haplotype and a maximum length of the haplotype.

[0080] The VNTR status may include a genotypic status of the VNTR. The genotypic status may include a genotype, a haplotype length of the genotype, and / or a confidence interval of the length of each haplotype of the genotype. The confidence interval may include a minimum length of each haplotype and a maximum length of each haplotype. The computing system may determine the lengths of the determined two or more haplotypes. The minimum haplotype length may be the minimum length of the lengths of the determined two or more haplotypes. The maximum haplotype length may be the maximum length of the lengths of the determined two or more haplotypes.

[0081] In some embodiments, the computing system generates a user interface (UI), such as a graphical user interface, that includes or represents the state of the VNTR. The UI can include, for example, a dashboard. The UI can include one or more UI elements. The UI element can include or represent the state of the VNTR. The UI element can be a window (e.g., a container window, a browser window, a text terminal, a child window, or a message window), a menu (e.g., a menu bar, a context menu, or a menu extra), an icon, or a tab. The UI element can be for an input control (e.g., a check box, a radio button, a drop-down list, a list box, a button, a toggle, a text field, or a date field). The UI element can be for navigation (e.g., a breadcrumb, a slider, a search field, pagination, a slider, a tag, an icon). The UI element can provide information (e.g., a tooltip, an icon, a progress bar, a notification, a message box, or a modal window). The UI element can be a container (e.g., an accordion).

[0082] The method 500 ends at block 532.

[0083] Execution environment FIG. 6 illustrates a general architecture of an exemplary computing device 600 configured to determine a VNTR status, such as genotyping a VNTR. The general architecture of the computing device 600 illustrated in FIG. 6 includes an arrangement of computer hardware and software components. The computing device 600 may include more (or less) elements than those illustrated in FIG. 6. However, not all of these general conventional elements need to be illustrated to provide a useful disclosure. As illustrated, the computing device 600 includes a processing unit 610, a network interface 620, a computer-readable medium drive 630, an input / output device interface 640, a display 650, and an input device 660, all of which may communicate with each other via a communication bus. The network interface 620 may provide connectivity to one or more networks or computing systems. The processing unit 610 may therefore receive information and instructions from other computing systems or services via a network. The processing unit 610 may also communicate with a memory 670 and further provide output information for the optional display 650 via the input / output device interface 640. The input / output device interface 640 may also accept input from any input device 660, such as a keyboard, a mouse, a digital pen, a microphone, a touch screen, a gesture recognition system, a voice recognition system, a game pad, an accelerometer, a gyroscope, or other input device.

[0084] Memory 670 may include computer program instructions (grouped into modules or components in some embodiments) that processing unit 610 executes to implement one or more embodiments. Memory 670 generally includes RAM, ROM, and / or other persistent, secondary, or non-transitory computer-readable media. Memory 670 may store an operating system 672 that provides computer program instructions for use by processing unit 610 in the overall management and operation of computing device 600. Memory 670 may further include computer program instructions and other information for implementing aspects of the present disclosure.

[0085] For example, in one embodiment, memory 670 includes a VNTR status determination module 674 for determining a VNTR status, such as method 500 described with reference to Figure 5. Additionally, memory 670 can include or be in communication with data store 690 and / or one or more other data stores that store one or more inputs, one or more outputs, and / or one or more results (including intermediate results) of determining a VNTR status of the present disclosure, such as determined long reads, determined haplotypes, short reads, and determined VNTR status (e.g., haplotype or genotype of a sample).

[0086] Additional Considerations In at least some of the foregoing embodiments, one or more elements used in one embodiment may be used interchangeably in another embodiment, except where such an exchange is not technically feasible. Those skilled in the art will appreciate that various other omissions, additions, and modifications may be made to the methods and structures described above without departing from the scope of the claimed subject matter. All such modifications and variations are intended to be included within the scope of the subject matter, as defined by the appended claims.

[0087] Those skilled in the art will appreciate that, for this and other processes and methods disclosed herein, the functions performed in the processes and methods may be performed in differing orders. Moreover, the outlined steps and operations are provided only as examples, and some of the steps and operations may be optional, may be combined into fewer steps and operations, or may be expanded to additional steps and operations without departing from the essence of the disclosed embodiments.

[0088] For the use of substantially any plural and / or singular term herein, one of ordinary skill in the art may substitute the plural for the singular and / or the singular for the plural as appropriate to the context and / or application. For clarity, various singular / plural permutations may be expressly set forth herein. As used herein and in the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, phrases such as "a device configured to" are intended to include one or more enumerated devices. Such one or more enumerated devices may also be collectively configured to execute the recited detailed description. For example, "a processor configured to execute detailed description A, B, and C" may include a first processor configured to execute detailed description A and to perform operations in conjunction with a second processor configured to execute detailed description B and C. Any reference to "or" herein is intended to encompass "and / or" unless otherwise indicated.

[0089] Generally, the terms used herein, and particularly the terms used in the appended claims (e.g., the body of the appended claims), are generally intended as "open" terms (e.g., the term "including" should be interpreted as "including but not limited to", the term "having" should be interpreted as "having at least", the term "includes" should be interpreted as "includes but is not limited to", etc.). Those skilled in the art will further understand that where a specific number of introduced claim recitations are intended, such intention will be expressly recited in the claim, and in the absence of such recitation, no such intention exists. For example, to aid in understanding, the following appended claims may include the use of the introductory phrases "at least one" and "one or more" to introduce the claim recitations. However, the use of such phrases should not be construed as limiting any particular claim that includes such an introduced claim statement to embodiments that include only one of such statements, even when the same claim includes "one or more" or "at least one" and an indefinite article such as "a" or "an" (e.g., "a" and / or "an" should be interpreted to mean "at least one" or "one or more"), and the same applies to the use of indefinite articles used to introduce claim statements. Moreover, even when a specific number of introduced claim statements is explicitly recited, one of skill in the art will recognize that such a recitation should be interpreted in the sense of at least the recited number (e.g., the unmodified recitation of "two statements" means, in the absence of other modifications, at least two statements or more than two statements).Furthermore, when phrases similar to "at least one of A, B, and C, etc." are used, generally such structures are intended in the sense that one of ordinary skill in the art would understand the phrase (e.g., "a system having at least one of A, B, and C" includes, but is not limited to, systems having only A, only B, only C, both A and B, both A and C, both B and C, and / or both A, B, and C, etc.). When phrases similar to "at least one of A, B, or C, etc." are used, generally such structures are intended in the sense that one of ordinary skill in the art would understand the phrase (e.g., "a system having at least one of A, B, or C" includes, but is not limited to, systems having only A, only B, only C, both A and B, both A and C, both B and C, and / or both A, B, and C, etc.). Those skilled in the art will further appreciate that virtually any disjunction and / or phrase presenting two or more alternative terms, whether in the specification, claims, or drawings, should be understood to contemplate the possibility of including one of the terms, either of the terms, or both terms. For example, the phrase "A or B" will be understood to include the possibilities of "A" or "B" or "A and B."

[0090] Additionally, where features or aspects of the disclosure are described in terms of a Markush group, one of skill in the art will thereby recognize that the disclosure is also described in terms of any individual members or subgroups of the Markush group members.

[0091] As will be understood by those skilled in the art, for any and all purposes, such as in terms of providing a written description, all ranges disclosed herein also encompass any and all possible subranges and combinations of those subranges. Any recited range can be readily recognized as fully descriptive and allowing the same range to be broken down into at least two equal parts, a third, a quarter, a fifth, a tenth, etc. As a non-limiting example, each range described herein can be readily broken down into a lower third, a middle third, an upper third, etc. As will also be understood by those skilled in the art, all language such as "up to," "at least," "greater than," "less than" and the like refers to a range that includes the recited numbers and can then be broken down into subranges as described above. Finally, as will be understood by those skilled in the art, a range includes each individual component. Thus, for example, a group having 1 to 3 items means a group having 1, 2, or 3 items. Similarly, a group having 1-5 items means a group having 1, 2, 3, 4, or 5 items, etc.

[0092] Various embodiments of the present disclosure have been described herein for illustrative purposes, and it will be understood that various modifications may be made without departing from the scope and spirit of the present disclosure. Accordingly, the various embodiments disclosed herein are not intended to be limiting, the true scope and spirit of which is set forth by the following claims.

[0093] It should be understood that not all objectives or advantages may necessarily be achieved in accordance with any particular embodiment described herein. Thus, for example, those skilled in the art will recognize that a particular embodiment may be configured to operate in a manner that achieves or optimizes an advantage or set of advantages as taught herein, without necessarily achieving other objectives or advantages that may be taught or suggested herein.

[0094] All of the processes described herein may be embodied in, and may be fully automated via, software code modules executed by a computing system including one or more computers or processors. The code modules may be stored on any type of non-transitory computer-readable medium or other computer storage device. Some or all of the methods may be embodied in dedicated computer hardware.

[0095] Many other variations beyond those described herein will be apparent from this disclosure. For example, depending on the embodiment, certain acts, events, or functions of any of the algorithms described herein may be performed in a different order, added, combined, or omitted entirely (e.g., not all described acts or events are necessary to implement the algorithm). Furthermore, in certain embodiments, acts or events may be performed simultaneously rather than sequentially, e.g., via multi-threading, interrupt processing, or multiple processors or processor cores, or on other parallel systems. In addition, different tasks or processes may be performed by different machines and / or computing systems that can function together.

[0096] The various example logic blocks and modules described in connection with the embodiments disclosed herein may be implemented or performed by mechanical devices such as processing units or processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. The processor may be a microprocessor, but alternatively the processor may be a controller, microcontroller, or state machine, combinations thereof, and the like. The processor may include electrical circuitry configured to process computer-executable instructions. In another embodiment, the processor includes an FPGA or other programmable device that performs logical operations without processing computer-executable instructions. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in association with a DSP core, or any other such configuration. Although primarily digital technology is described herein, the processor may also include primarily analog components. For example, some or all of the signal processing algorithms described herein may be implemented in analog circuitry or mixed analog and digital circuitry. The computing environment may include any type of computer system, including, but not limited to, a computer system based on a computing engine within a microprocessor, mainframe computer, digital signal processor, portable computing device, device controller, or appliance, to name a few.

[0097] Any process illustrations, elements or blocks in the flow diagrams described herein and / or shown in the accompanying drawings should be understood as potentially representing modules, segments or portions of code that include one or more executable instructions for implementing a particular logical function or element in the process. Alternate embodiments are included within the scope of the embodiments described herein, and depending on the functionality involved, elements or functions may be omitted or removed from the order from that shown or discussed, including substantially simultaneously or in reverse order, as will be understood by those skilled in the art.

[0098] It should be emphasized that many variations and modifications may be made to the above-described embodiments, and that the elements thereof should be understood to be among the other acceptable examples. All such modifications and variations are intended to be included herein within the scope of this disclosure and protected by the following claims.

Claims

1. A system for determining a variable number tandem repeat (VNTR) status, comprising: A non-transitory memory configured to store executable instructions and a plurality of haplotypes of the VNTR; and A hardware processor in communication with the non-transitory memory, the hardware processor programmed by the executable instructions to: Receive a plurality of short sequence reads generated from a test sample obtained from a test subject; For each of the plurality of haplotypes of the VNTR, realign a short sequence read among the plurality of short sequence reads aligned to the VNTR to the haplotype to generate a realignment; Use the realignment of the short sequence reads realigned to the haplotype to determine a probability for each of the plurality of haplotypes with respect to the test subject; Determine the VNTR status of the test subject; and Execute, a hardware processor A system for determining a variable number tandem repeat (VNTR) status, comprising.

2. The system of claim 1, wherein the plurality of haplotypes of the VNTR are determined using long sequence reads of a plurality of long sequence reads aligned to the VNTR in a reference, and the plurality of long sequence reads are generated from a plurality of reference samples obtained from a plurality of reference subjects.

3. The plurality of haplotypes of the VNTR are: For each of the plurality of samples: Extract a long sequence read among the plurality of long sequence reads of the test sample aligned to the VNTR in the reference; Realign the long sequence reads extracted to the left flanking region and the right flanking region of the VNTR to determine the aligned long sequence reads. determining the haplotypes of the plurality of haplotypes based on the aligned long sequence reads each having an alignment score exceeding an alignment threshold; The system according to claim 2, determined by.

4. the haplotypes of the plurality of haplotypes of the VNTR being aligning and trimming the sequences of the aligned long sequence reads each having an alignment score exceeding the alignment threshold with respect to the left flanking region and the right flanking region to generate trimmed long sequence reads; determining the haplotypes of the plurality of haplotypes based on the trimmed long sequence reads; The method according to claim 3, determined by.

5. The system according to claim 3, wherein the reference sample is homozygous for the VNTR, and the haplotypes of the plurality of haplotypes of the VNTR are determined to include only one of the plurality of haplotypes based on the trimmed long sequence reads.

6. The system according to claim 3, wherein the reference sample is heterozygous for the VNTR, and the haplotypes of the plurality of haplotypes of the VNTR are determined to include two of the plurality of haplotypes based on the trimmed long sequence reads.

7. The system according to claim 3, wherein the haplotypes of the plurality of haplotypes of the VNTR are determined by determining a consensus sequence of the trimmed long sequence reads.

8. The system according to claim 2, wherein the quality of the long sequence reads of the plurality of long sequence reads aligned to the VNTR in the reference and / or the quality of the plurality of haplotypes meet a quality standard.

9. The system according to claim 1, wherein the state of the VNTR includes the haplotype state and / or the genotype state of the VNTR, and optionally, the haplotype state includes a haplotype, the length of the haplotype, and a confidence interval of the length of the haplotype, and optionally, the genotype state includes a genotype, the length of the haplotype of the genotype, and a confidence interval of each of the lengths of the haplotypes of the genotype.

10. The system according to claim 1, wherein each of the probability indicators of the plurality of haplotypes of the VNTR includes the probability of each of the plurality of haplotypes of the VNTR, and the probability standard includes a probability threshold.

11. The system according to claim 1, wherein the accuracy of the haplotype state is at least 60%.

12. The system according to claim 1, wherein the plurality of long sequence reads include sequence reads each having a length of about 10,000 base pairs to about 20,000 base pairs.

13. The system according to claim 1, wherein the plurality of short sequence reads include sequence reads each having a length of about 100 base pairs to about 1000 base pairs.

14. The system according to claim 1, wherein each haplotype of the plurality of haplotypes of the VNTR includes a plurality of copies of a repeat unit.

15. The system according to claim 14, wherein the repeat unit has a length exceeding 6 base pairs.

16. The system according to claim 14, wherein the number of the plurality of copies is at least 3.

17. The system according to claim 14, wherein the arrays of two copies of the plurality of copies of the repeating unit of the haplotype of the plurality of haplotypes differ at one or more differentiation positions.

18. The system according to claim 1, wherein the haplotypes of the plurality of haplotypes of the VNTR are associated with a disease.

19. The system according to claim 1, wherein the hardware processor is programmed by the executable instructions to perform generating a user interface (UI) comprising UI elements representing the state of the VNTR.

20. A method for determining a variable number tandem repeat (VNTR) state, comprising: under the control of a hardware processor: receiving a plurality of long array reads generated from a plurality of first samples obtained from a plurality of first subjects; determining a plurality of haplotypes of the VNTR using the long array reads among the plurality of long array reads aligned to the VNTR in a reference; receiving a plurality of short array reads generated from a second sample obtained from a second subject; for each of the plurality of haplotypes of the VNTR, aligning the short array reads among the plurality of short array reads aligned to the VNTR to the haplotype to generate an alignment; determining a probability metric for each of the plurality of haplotypes of the VNTR for the second subject using the realignment of the short array reads realigned to the haplotype, and determining the state of the VNTR of the second subject based on the probability metric for each of the plurality of haplotypes. A method comprising.