Genotyping by sequencing

JP7899168B2Active Publication Date: 2026-08-03REGENERON PHARMACEUTICALS INC
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
REGENERON PHARMACEUTICALS INC
Filing Date
2021-11-19
Publication Date
2026-08-03

AI Technical Summary

Benefits of technology

【0008】 本開示は、シーケンシングによるジェノタイピングのための核酸プローブを製造する方法であって、a)核酸プローブによって捕捉するための複数の直接観察される遺伝的バリアントを選択すること、b)複数の直接観察される遺伝的バリアントから低信頼度バリアントを排除し、それにより、フィルタリングされた複数の直接観察される遺伝的バリアントを作成すること、c)フィルタリングされた複数の直接観察される遺伝的バリアントをフェージングすること、d)フィルタリングされた複数の直接観察される遺伝的バリアントのうちの各バリアントについて、1つまたは複数のプロキシバリアントの存在または非存在を識別すること、e)フィルタリングされた複数の直接観察される遺伝的バリアントを含むゲノムDNAの複数の候補領域を選択することであって、ゲノムDNAの各候補領域が、約25~約150の塩基を含み、フィルタリングされた複数の直接観察される遺伝的バリアントの中の少なくとも1つのバリアントを含む、選択すること、f)ゲノムDNAの各候補領域について、プローブの捕捉効率及びアラインメント成功を推定するクオリティスコアを算出すること、g)ゲノムDNAの各候補領域について、ゲノムDNAの候補領域によって捕捉されるバリアントの数をクオリティスコアに乗算することにより、プローブスコアを算出することであって、ゲノムDNAの候補領域によって捕捉されるバリアントの数が、ゲノムDNAの候補領域によって捕捉される直接観察されるバリアントの数と、ゲノムDNAの異なる候補領域における対応するプロキシバリアントの数との和である、算出すること、h)ゲノムDNAの領域の最終セットに含めるために、最も高いプローブスコアを有するゲノムDNAの1つまたは複数の候補領域を選択すること、i)ゲノムDNAの領域の最終セットに含めるために、選択されていないゲノムDNAの候補領域に対してステップg)及びh)を繰り返すことであって、選択されていないゲノムDNAの候補領域におけるバリアントの数が、1)選択済みのゲノムDNAの領域内のすべての直接観察されるバリアントを除外した、選択されていないゲノムDNAの候補領域における直接観察されるバリアントの数と、2)選択済みのゲノムDNAの領域内の直接観察されるバリアントに対応するすべてのプロキシバリアントを除外した、ゲノムDNAの異なる候補領域における対応するプロキシバリアントの数との和であり、最大数のゲノムDNAの領域が選択されるまでステップg)及びh)が繰り返される、繰り返すこと、及びj)ゲノムDNAの領域の最終セットの中の各ゲノム領域の核酸配列に相補的な核酸プローブのセットを生成することを含む方法を提供する。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007899168000004
    Figure 0007899168000004
  • Figure 0007899168000005
    Figure 0007899168000005
  • Figure 0007899168000001
    Figure 0007899168000001
Patent Text Reader

Abstract

The present disclosure provides methods for producing nucleic acid probes for genotyping by sequencing, methods for genotyping a DNA sample by sequencing using a set of nucleic acid probes, and systems for performing such methods.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This disclosure, in part, relates to a method for producing nucleic acid probes for sequencing-based genotyping, a method for performing genotyping of a DNA sample by sequencing using a set of nucleic acid probes, and a system for performing such a method. [Background technology]

[0002] Whole-genome sequencing involves sequencing the entire genome of an individual. While the cost of whole-genome sequencing has decreased, it remains substantial. The cost increases with increasing sequencing depth. Because different parts of the genome have different levels of focus or interest, the requirements for deep sequencing vary.

[0003] Instead of sequencing the entire genome at a constant depth as expected, it is possible to pre-select regions of the genome for sequencing (and thus perform most of the sequencing in these regions). Exome sequencing targets the sequencing of gene exons by capturing short DNA strands that overlap with the gene's exons and then sequencing these short DNA strands. Exons are of great interest due to their functional and clinical significance. Directly sequencing exons allows for the observation of genetic variation in specific individual samples without referencing other samples. Exome sequencing targets only about 1% of the genome, but returns unbiased, functional, and actionable genetic variation at a significantly lower cost compared to whole-genome sequencing.

[0004] An alternative to sequencing strategies is the observation of genetic variation using DNA microarray technology, which was developed earlier and on a larger scale than sequencing. DNA microarray technology allows for the assay of, for example, hundreds of thousands of specific variants at once using DNA chips. These genetic variants typically represent genetic variation across the entire genome. Genotyping arrays, which measure genetic variation at hundreds of thousands to millions of variable sites in DNA, are driving modern human genetics. The variable sites measured by each array are typically selected to represent common genetic variation in one or more populations of interest. This strategy provides an inexpensive and effective alternative to direct whole-genome sequencing and is currently used for genotyping millions of DNA samples annually. The data obtained allows consumer-facing genetics companies to estimate an individual's ancestry and match individuals to their DNA relatives. This has also facilitated genome-wide association studies (GWAS), genomic risk scores, and Mendelian randomization analyses, which have yielded many insights into the ecology of diverse complex traits related to human health and behavior, ranging from cardiovascular and metabolic diseases to mental disorders, and from human behavior to age-related disorders and cancer.

[0005] Traditional strategies for array design focus on a set of known common genetic variants and attempt to identify a subset of these variants that are expected to perform well in multiplex genotyping experiments and that adequately represent other known common variants. Typically, each variant is assigned a probe score that measures its expected performance on the array platform. This score summarizes factors such as the presence of other neighboring variants, repetition, the proportion of guanine-cytosine (GC) bases in the probe DNA sequence, and the performance of similar probes in previous genotyping arrays. Each of these factors can influence the performance of a genotyping probe targeting the variant. In addition to this probe score that summarizes the expected performance of the probe, variants are also commonly mapped to a list of other common variants they may represent. Variants that represent mutations in other neighboring common variants are “proxy” or “surrogate” versions of these additional variants. These proxy relationships are commonly observed between neighboring variants in the human genome due to a process known as linkage disequilibrium. Linkage disequilibrium is the result of a genetic variant entering a population through mutation or migration, and then gradually spreading through inheritance, recombination, and gene conversion. Mutation, migration, inheritance, recombination, and gene conversion all often produce neighboring genetic variants in predictable combinations, which usually reflect the ancestral chromosome from which each variant first entered the population.

[0006] Genotyping arrays, such as DNA microarrays, observe only a small subset of variants in individual samples. Selecting a set of directly observable variants to include in a genotyping array ultimately involves selecting a set of directly observable variants with a high "probe score" that can act as a "proxy" for the majority of all known genetic variants. It is possible to indirectly observe (complement) variants from directly observable variants. This process is called complementation. Complementation works because our genetic variations are inherited in such a way that the closer multiple variants are to each other on the same chromosome, the higher the probability that they were inherited from the same ancestor. Complementation has been shown to yield high-quality results for complementing variants that are not directly observable, taking into account inferences about the manner in which DNA segments are inherited. While this strategy yields a list of variants that well represent common genetic variations in humans, it is also inefficient for techniques that measure multiple genetic variants with a single probe. Another problem with DNA microarray assays is that they are entirely separate processes in the laboratory and require replication of many processes, making the experiments inefficient. What is needed is a cost-effective experimental strategy that enables direct sequencing of desired target regions while maintaining the ability to complement variants across the entire genome. [Overview of the project] [Problems that the invention aims to solve]

[0007] Genotyping techniques have remained largely unchanged for nearly 20 years. Arrays produce high-quality data and consistent results at low cost, but they are labor-intensive. Arrays require additional processing and equipment that differs from those used for whole-exome sequencing. Array scalability and customization are limited. Efficient processing of millions of samples is required. [Means for solving the problem]

[0008] This disclosure relates to a method for producing nucleic acid probes for sequencing-based genotyping, comprising: a) selecting a plurality of directly observable genetic variants to be captured by the nucleic acid probe; b) removing low-confidence variants from the plurality of directly observable genetic variants to create a filtered plurality of directly observable genetic variants; c) phasing the filtered plurality of directly observable genetic variants; d) identifying the presence or absence of one or more proxy variants for each variant among the filtered plurality of directly observable genetic variants; e) selecting a plurality of candidate regions of genomic DNA containing the filtered plurality of directly observable genetic variants, wherein each candidate region of genomic DNA contains approximately 25 to approximately 150 bases and contains at least one variant among the filtered plurality of directly observable genetic variants; f) calculating a quality score for each candidate region of genomic DNA to estimate probe capture efficiency and alignment success; g) h) Calculate the probe score for each candidate region of genomic DNA by multiplying the quality score by the number of variants captured by the candidate region of genomic DNA, wherein the number of variants captured by the candidate region of genomic DNA is the sum of the number of directly observed variants captured by the candidate region of genomic DNA and the number of corresponding proxy variants in different candidate regions of genomic DNA, i) Select one or more candidate regions of genomic DNA with the highest probe score to include in the final set of genomic DNA regions, i) Repeat steps g) and h) for unselected candidate regions of genomic DNA to include in the final set of genomic DNA regions, wherein the number of variants in the unselected candidate regions of genomic DNA is 1) the number of directly observed variants in the unselected candidate regions of genomic DNA excluding all directly observed variants in the selected regions of genomic DNA, and 2) the number of proxy variants excluding all directly observed variants in the selected regions of genomic DNA.The present invention provides a method comprising repeating steps g) and h) until the maximum number of genomic DNA regions are selected, and j) generating a set of nucleic acid probes complementary to the nucleic acid sequence of each genomic region in the final set of genomic DNA regions.

[0009] The disclosure also provides a method for genotyping a DNA sample by sequencing, comprising: a) hybridizing a DNA sample with a set of nucleic acid probes manufactured as described above to generate probe-hybridized genomic DNA; b) sequencing the probe-hybridized genomic DNA to create a plurality of sequencing reads; c) mapping the plurality of sequencing reads to a reference genome; d) calling directly observable variants present in the mapped sequencing reads; and e) supplementing unobserved variants from unsequenced regions of the genomic DNA to establish the genotype of the sample DNA.

[0010] The disclosure also provides a method for genotyping a DNA sample by sequencing using a set of nucleic acid probes, comprising: a) selecting multiple regions of genomic DNA from a DNA sample containing multiple directly observable genetic variants; b) identifying a set of nucleic acid probes for hybridization of the selected genomic DNA to multiple regions; c) hybridizing the set of nucleic acid probes to the DNA sample to generate genomic DNA hybridized to the probes; d) sequencing the genomic DNA hybridized to the probes to create multiple sequencing reads; e) mapping the multiple sequencing reads to a reference genome; f) calling directly observable variants present in the mapped sequencing reads; and g) supplementing unobserved variants from unsequenced regions of the genomic DNA, thereby establishing the genotype of the sample DNA.

[0011] This patent or application document includes at least one drawing drawn in color. Copies of this patent or patent application publication, including the color drawing(s), will be provided by the Japan Patent Office upon request and payment of the necessary fees. [Brief explanation of the drawing]

[0012] [Figure 1] The Rsq is shown supplemented by variant bins for two different observations [one being a global screening array (GSA) and the other being a sequencing-based genotyping method (GxS) as described herein] and two in silico versions for comparison [one labeled "Simulated GxS" having all variants in the probe from the observed probe region, and the other labeled "Simulated MEGA" having all variants in the region assayed by a MEGA microarray (containing 1.8M variants)]. [Figure 2] The data shows the mean call rate of 98.9% in sequencing-based genotyping assays performed on 223,266 samples, each evaluated for coverage at a design site, and the percentage of samples with a call rate of 95% or higher (99.3%). The call rate represents the percentage of regions containing actionable genotypes. [Modes for carrying out the invention]

[0013] Provided herein is a schematic strategy that can be used to efficiently design a set of nucleic acid probes in which each probe can target multiple genetic variants for use, for example, in capture-based "sequencing-based genotyping" methods. Such capture-based "sequencing-based genotyping" methods target multiple short segments ("target regions," each typically 10 to several hundred base pairs long) of the genome, each potentially containing multiple known genetic variants. Selecting variants to target individually is inefficient in these experiments. For example, in a worst-case scenario, targeting 100,000 independently selected variants might require 100,000 short target regions. In a more desirable scenario, these 100,000 variants can be clustered together and captured with a significantly smaller number of probes. For example, a more desirable approach would be to capture only 25,000 short target regions (each containing an average of 4 variants) or 50,000 short target regions (each containing an average of 2 variants), while identifying a set of 100,000 variants that could be genotyped. Alternatively, a set of probes could identify 100,000 short target regions capturing 200,000 to 400,000 variants (and thus likely to perform significantly better than selecting 100,000 target regions after independently selecting 100,000 variants).

[0014] The methods described herein identify small sets of genomic regions for sequencing, aiming to approach the comprehensiveness of whole-genome sequencing with significantly reduced cost and effort. These regions are selected so as to be expected to function well in targeted capture experiments. Furthermore, these regions, when considered together, include a set of common genetic variants that accurately summarize the variations within the genome for GWAS, ancestry estimation, identification of genetic relatives, estimation of polygenic risk scores, and other applications that currently rely on genotyping arrays.

[0015] The method described herein provides a sequencing-based alternative to genotyping arrays. The method provides better genomic coverage than standard arrays across multiple ancestors. By selecting a large number of common variants, such as approximately 1.4M, highly accurate complementation between multiple ancestors can be enabled. The method can also cover approximately 4.5M–5.0M common variants per sample with one or more sequencing reads. The reagents described herein have been iteratively refined by application to diverse ancestral samples. Features of the method described herein include, but are not limited to, generating data in parallel with whole-exome sequencing of each sample, selecting a large portion of 1.4M common variants to enable complementation of genome-wide mutations, and targeting peaks, mitochondrial DNA, Y chromosome, and MHC in genome-wide association studies where additional variants are known. The method described herein generates high-fidelity genotypes for approximately 1.4M variants per sample. These 1.4M variants exhibit a call rate of approximately 98.9% and an accuracy of approximately 99.7% compared to high-depth whole-genome sequencing data. These 1.4M variants can be used as an alternative to array genotypes in most applications. The methods described herein are bioinformatically efficient, requiring less than 10 hours of CPU time in addition to typical exome processing procedures. Each sample can be processed and handled independently.

[0016] The sequencing-based methods for genotyping described herein are based on the high-throughput DNA capture techniques described herein. The DNA capture methodologies described herein are highly automated and scaled to process millions of samples annually. High-quality exome data and genotyping can be performed simultaneously, facilitating the integration of results. The methods described herein also have the advantage of evolving over time, improving coverage of regions or variants of interest. The methods described herein achieve varying sequence coverage and accuracy for high-value variants. The methods described herein maximize tagging while minimizing the number of capture targets. The probe sets described herein have been validated and improved by using them on a variety of samples to remove / replace poor targets. The probes are selected and experimentally validated to represent genetic variations across multiple ancestors. The probe sets target approximately 1.5M variant sites per sample, covering approximately 2.6% of the genome.

[0017] The terms used herein are intended to describe only specific embodiments and are not intended to limit them. The methods described herein provide for the selection and manufacture of sets of nucleic acid probes such that each probe can efficiently capture short strands of DNA overlapping with the probe and similarly produce sequencing reads that can be aligned. Furthermore, the methods described herein focus on regions of genomic DNA with genetic variations that enable good complementation of nearby unobserved genetic variations (i.e., complementary variants) and / or direct observation of significant variations.

[0018] This disclosure relates to a method for producing nucleic acid probes for sequencing-based genotyping, comprising: a) selecting a plurality of directly observable genetic variants to be captured by the nucleic acid probe; b) removing low-confidence variants from the plurality of directly observable genetic variants to create a filtered plurality of directly observable genetic variants; c) phasing the filtered plurality of directly observable genetic variants; d) identifying the presence or absence of one or more proxy variants for each variant among the filtered plurality of directly observable genetic variants; e) selecting a plurality of candidate regions of genomic DNA containing the filtered plurality of directly observable genetic variants, wherein each candidate region of genomic DNA contains approximately 25 to approximately 150 bases and contains at least one variant among the filtered plurality of directly observable genetic variants; f) calculating a quality score for each candidate region of genomic DNA to estimate probe capture efficiency and alignment success; g) h) Calculate the probe score for each candidate region of genomic DNA by multiplying the quality score by the number of variants captured by the candidate region of genomic DNA, wherein the number of variants captured by the candidate region of genomic DNA is the sum of the number of directly observed variants captured by the candidate region of genomic DNA and the number of corresponding proxy variants in different candidate regions of genomic DNA, i) Select one or more candidate regions of genomic DNA with the highest probe score to include in the final set of genomic DNA regions, i) Repeat steps g) and h) for unselected candidate regions of genomic DNA to include in the final set of genomic DNA regions, wherein the number of variants in the unselected candidate regions of genomic DNA is 1) the number of directly observed variants in the unselected candidate regions of genomic DNA excluding all directly observed variants in the selected regions of genomic DNA, and 2) the number of proxy variants excluding all directly observed variants in the selected regions of genomic DNA.The present invention provides a method comprising repeating steps g) and h) until the maximum number of genomic DNA regions are selected, and j) generating a set of nucleic acid probes complementary to the nucleic acid sequence of each genomic region in the final set of genomic DNA regions.

[0019] This disclosure relates to a method for designing nucleic acid probes for sequencing-based genotyping, comprising: a) selecting multiple directly observable genetic variants to be captured by the nucleic acid probe; b) eliminating low-confidence variants from the multiple directly observable genetic variants to create a filtered set of directly observable genetic variants; c) phasing the filtered set of directly observable genetic variants; d) identifying the presence or absence of one or more proxy variants for each variant among the filtered set of directly observable genetic variants; e) selecting multiple candidate regions of genomic DNA containing the filtered set of directly observable genetic variants, wherein each candidate region of genomic DNA contains approximately 25 to approximately 150 bases and contains at least one variant among the filtered set of directly observable genetic variants; f) calculating a quality score for each candidate region of genomic DNA to estimate probe capture efficiency and alignment success; g) geno h) calculate the probe score for each candidate region of genomic DNA by multiplying the quality score by the number of variants captured by the candidate region of genomic DNA, wherein the number of variants captured by the candidate region of genomic DNA is the sum of the number of directly observed variants captured by the candidate region of genomic DNA and the number of corresponding proxy variants in different candidate regions of genomic DNA, i) select one or more candidate regions of genomic DNA with the highest probe score to include in the final set of genomic DNA regions, and i) repeat steps g) and h) for unselected candidate regions of genomic DNA to include in the final set of genomic DNA regions, wherein the number of variants in the unselected candidate regions of genomic DNA is 1) the number of directly observed variants in the unselected candidate regions of genomic DNA excluding all directly observed variants in the selected regions of genomic DNA, and 2) the number of proxy variants excluding all directly observed variants in the selected regions of genomic DNA.The method also provides a method that includes repeating steps g) and h), where the sum of the number of corresponding proxy variants in different candidate regions of genomic DNA is used, and steps g) and h) are repeated until the maximum number of genomic DNA regions is selected.

[0020] The present method involves selecting multiple genetic variants to be captured by nucleic acid probes. These selected variants constitute a desired set of “directly observable genetic variants.” A “directly observable genetic variant” is a variant present in genomic DNA that is captured by hybridization of at least one probe and subsequently sequenced. Directly observable variants are distinct from the remaining genetic variants, including complementary variants. Any complementary variants are likely to be present in the same genomic DNA but are not captured by hybridization of at least one probe, and therefore are not subsequently sequenced. The presence of directly observable variants in the genomic DNA and subsequent sequencing enables the complementation of complementary variants.

[0021] Multiple directly observable genetic variants to be captured by nucleic acid probes can include any desired number of known common variants. For example, a set of M known genetic variants could be V1, V2, V3…V M This can be considered as follows: The exponents m and n vary between 1 and M and are used to specify individual variants. Each variant V m This is a known chromosome position P m and allele A m It has a set of variants V n This is a known chromosome position P n and allele A nhas a set. In some embodiments, the plurality of directly observed genetic variants includes every known common variant. In some embodiments, the plurality of directly observed genetic variants is selected from a database of genome-wide associations of genetic variants, a database of pharmacogenetic associations of genetic variants, a database containing genetic variants within all mitochondrial chromosomes, and / or a database of genetic variants within a microarray, or any combination thereof.

[0022] In some embodiments, the plurality of directly observed genetic variants is selected from one or more databases of genome-wide associations of genetic variants. Any of the databases of genome-wide associations of genetic variants can be used to identify one or more directly observed genetic variants for inclusion. In some embodiments, the database of genome-wide associations of genetic variants is a catalog of known genome-wide association hits (see, e.g., the world wide web at "ebi.ac.uk / gwas / "). In some embodiments, the source file was "gwas_catalog_v1.0.2-associations_e96_r2019-07-30.tsv." In some embodiments, not all variants within the database of genome-wide associations of genetic variants are selected. In some embodiments, variants within the database of genome-wide associations of genetic variants are selected to be included in the plurality of directly observed genetic variants if the association of the variant with the trait has a p-value ≤ 10 -9 and is selected to be included in the plurality of directly observed genetic variants. In some embodiments, variants within the database of genome-wide associations of genetic variants are not selected to be included in the plurality of directly observed genetic variants if the association of the variant with the trait has a p-value > 10 -9If a variant has this characteristic, it is excluded from multiple directly observed genetic variants. In some embodiments, this p-value analysis excludes variants present on the Y chromosome and mitochondrial chromosomes. In some embodiments, the number of variants selected from the genome-wide association database(s) of genetic variants is approximately 30,000 to 45,000. In some embodiments, the number of variants selected from the genome-wide association database(s) of genetic variants is approximately 35,000 to 40,000. In some embodiments, the number of variants selected from the genome-wide association database(s) of genetic variants is approximately 38,000. The number of variants selected from the genome-wide association database(s) of genetic variants is expected to change over time.

[0023] In some embodiments, multiple directly observed genetic variants are selected from one or more databases of genetic variant gene pharmacological relevance. Any of the genetic variant gene pharmacological relevance databases can be used to identify one or more directly observed genetic variants to include. In some embodiments, the genetic variant gene pharmacological relevance database is data published by PharmGKB on gene pharmacological relevance. In some embodiments, all sites that are within dbSNPs and observed as single nucleotide polymorphisms (SNPs) that overlap with genes of pharmacogenetic interest are included. In some embodiments, the number of variants selected from the gene pharmacological relevance database(s) is approximately 2,000 to approximately 10,000. In some embodiments, the number of variants selected from the gene pharmacological relevance database(s) is approximately 4,000 to approximately 6,000. In some embodiments, the number of variants selected from the gene pharmacological relevance database(s) is approximately 5,000.

[0024] In some embodiments, multiple directly observable genetic variants are selected from one or more databases containing genetic variants within all mitochondrial chromosomes. Any of the databases containing genetic variants within all mitochondrial chromosomes can be used to identify one or more directly observable genetic variants to include. In some embodiments, all mitochondrial chromosomes are tiled from end to end.

[0025] In some embodiments, multiple directly observable genetic variants are selected from one or more databases of genetic variants in one or more microarrays. Any of the databases of genetic variants in microarrays can be used to identify one or more directly observable genetic variants to include. An exemplary database is the variants on microarrays used by the UK Biobank. In some embodiments, the database of genetic variants in microarrays includes genetic variants in the HLA region of chromosome 6, the Y chromosome, two killer cell immunoglobulin-like receptor (KIR) regions on chromosome 19, and pseudoautosomal regions 1 and 2 (Par1 and Par2) on the X chromosome.

[0026] In some embodiments, the database of genetic variants in the microarray includes genetic variants in the HLA region of chromosome 6. In some embodiments, the database of genetic variants in the microarray includes genetic variants in the HLA region of chromosome 6, defined as Chr6:28011410-33978119. Naturally, equivalent coordinates in alternative human genome assemblies are also included herein.

[0027] In some embodiments, the database of genetic variants in the microarray includes genetic variants on the Y chromosome. In some embodiments, the database of genetic variants in the microarray includes genetic variants in two KIR regions on chromosome 19. Naturally, equivalent coordinates in alternative human genome assemblies are also included herein.

[0028] In some embodiments, the database of genetic variants in the microarray includes genetic variants at Par1 and Par2 on the X chromosome. In some embodiments, the database of genetic variants in the microarray includes genetic variants at Par1 and Par2 on the X chromosome, defined as ChrX:10425-2774669 and ChrX:155704030-156003450. Naturally, equivalent coordinates in alternative human genome assemblies are also included herein. In some embodiments, the number of variants selected from the database(s) of genetic variants in the microarray is approximately 700,000 to approximately 900,000. In some embodiments, the number of variants selected from the database(s) of genetic variants in the microarray is approximately 800,000 to approximately 850,000. In some embodiments, the number of variants selected from the database(s) of genetic variants in the microarray is approximately 830,000.

[0029] In some embodiments, a multiallelic variant is converted into one or more sets of biallelic variants. The conversion involves two steps: one step converts the variant of the abstract, and the other converts the individual genotypes. In some embodiments, the multiallelic genotype of the original multiallelic variant is converted into the biallelic genotype of each of the decomposed genetic variants, enabling estimation of linkage disequilibrium coefficients and proxy relationships between the genetic variants. The methods described herein can address multiallelic variants by decomposing each multiallelic variant into a set of biallelic variants, each of which is assigned the same chromosomal location but different alleles. For example, if a particular multiallelic variant has one reference allele and three surrogate alleles, the multiallelic variant is converted into three sets of biallelic variants (i.e., the reference allele and the first surrogate allele, the reference allele and the second surrogate allele, and the reference allele and the third surrogate allele).

[0030] In some embodiments, we procured the whole-genome sequencing dataset from the 1000 Genomes Project (denoted as 1KG) to calculate metrics for possible completion success. High-coverage (30x) sequencing of 2,504 samples from 26 different populations was published for commercial use by the New York Genome Center in May 2019 (see the World Wide Web at "internationalgenome.org / data-portal / data-collection / 30x-grch38").

[0031] The method described here also includes eliminating low-confidence variants from multiple directly observable genetic variants, thereby creating a filtered set of multiple directly observable genetic variants. Eliminating low-confidence variants from multiple directly observable genetic variants serves as a quality control to limit the selected variants to those with higher confidence. In some embodiments, eliminating low-confidence variants from multiple potential directly observable genetic variants leaves approximately 15 million variants. Eliminating low-confidence variants from multiple directly observable genetic variants may include one or more of the following:

[0032] In some embodiments, excluding low-confidence variants from multiple directly observed genetic variants includes excluding all variants having minor allele frequencies (MAFs) below a desired threshold. For example, the allele frequency range is f min f max It can be considered that the variant in V is f min more than f max It may be restricted to variants having the following minor allele frequencies. For example, f max This can be set to 0.50. Furthermore, f min This can be 1% (0.01) or 5% (0.05). In some embodiments, the desired threshold is 1% (0.01). In some embodiments, this MAF threshold can be lowered to 0.1% (0.001).

[0033] In some embodiments, eliminating low-confidence variants from multiple directly observed genetic variants includes eliminating all variants with missing values ​​exceeding a desired threshold. In some embodiments, the desired threshold is 2%.

[0034] In some embodiments, excluding low-confidence variants from multiple directly observed genetic variants is possible with a P-value < 10 in the Hardy-Weinberg test in any of the sample populations. -8This includes excluding variants that have the same relationship.

[0035] The present method also includes phasing a filtered set of potential directly observable genetic variants. In some embodiments, the method includes phasing all variants observed in a 1000-genome sample or another reference panel. These variant phasings help the method and algorithm select “directly observable variants” and “probes” that work better. Phasing produces the best estimate of the variant sequence in each of the two chromosomes for each sample. Phasing variants in a 1000-genome reference panel (or another panel of reference individuals) improves the handling of missing data and estimates of linkage disequilibrium and proxy relationships between variants. In contrast, genotyping only has information about the count of a particular allele in the combination of both chromosomes. For example, the sequence with allele count 0,1,2,2,1,1 may be phased as two binary sequences 0,1,1,1,1,1 and 0,0,1,1,0,0, representing two sequences on each chromosome. Genotype calling phasing can be performed using commercially available software such as SHAPEIT4 (see the World Wide Web at "odelaneau.github.io / shapeit4 / ") with all standard defaults.

[0036] The method also includes identifying the presence or absence of one or more proxy variants for each of the filtered directly observed genetic variants. Each of the filtered directly observed genetic variants may potentially be a proxy (i.e., a proxy variant) for other variants that are neither probed nor sequenced (i.e., the proxy variant is complemented into the sample DNA genome based on the presence of the directly observed variant). These proxy relationships are commonly observed between neighboring variants in the human genome due to linkage disequilibrium. For example, to describe a proxy relationship between two variants, variant V m and V n Entry R describes the chain disequilibrium relationship between mn A matrix R containing the following can be used. Several suitable measures of linkage disequilibrium between variants exist and can be used in the methods described herein. In some embodiments, when the directly observed genetic variant and proxy variant are within 1 MB of each other, and the linkage disequilibrium between the two variants is r 2 If a squared correlation exceeds a desired threshold (t) using a scale, then the variants in the filtered, directly observed genetic variants have corresponding proxy variants in another region of genomic DNA. The tunable parameter t represents the minimum amount of linkage disequilibrium required before two variants can be considered proxies to each other. In some embodiments, the linkage disequilibrium between the two variants is r of the linkage disequilibrium. 2 The scale has a correlation squared (t) of at least 0.2. In some embodiments, the chain imbalance between the two variants is the r of the chain imbalance. 2 The scale has a correlation squared (t) of at least 0.5. In some embodiments, the chain imbalance between the two variants is the r of the chain imbalance. 2 The scale has a correlation squared (t) of at least 0.8. In some embodiments, the chain imbalance between the two variants is the r of the chain imbalance.2 The scale has a correlation squared (t) of at least 0.9. In some embodiments, the chain imbalance between the two variants is the r of the chain imbalance. 2 The scale has a correlation squared (t) of at least 1.0. In some embodiments, the proxy variant resides in a different candidate region of genomic DNA compared to its corresponding directly observed variant. Therefore, R mn When the value of is greater than t, there are two variants V m and V n They are each other's proxies.

[0037] Typically, a set of known genetic variants V and their linkage disequilibrium relationships R can be estimated by sequencing or genotyping a small set of individuals. The quality of the regions selected for sequencing improves as the number of individuals in this set increases. Furthermore, the individuals in this set should have diverse ancestors, or at least, it is desirable that they match the ancestral composition of the individuals studied using the selected target regions.

[0038] In some embodiments, identifying the presence or absence of one or more proxy variants for each directly observed variant can be done by software relating to chain imbalances. One such example is emeraLD, which uses normal defaults (see the World Wide Web at "github.com / statgen / emeraLD"). Using such software, it is possible to generate a list of variant pairs that are within 1Mb of each other and have a squared correlation exceeding a desired threshold t.

[0039] The method described here also involves selecting multiple candidate regions (i.e., target regions) of genomic DNA to be captured by nucleic acid probes. One goal is to obtain a set of K candidate regions of genomic DNA, T=T1, T2, T3, ...T KThe objective is to identify each candidate region of genomic DNA. The index k varies between 1 and K and can be used to specify individual candidate regions of genomic DNA. k This is the starting position Start(T k ) and the end position End(T k ) and the corresponding probe score (T k The probe score represents the expected performance of a candidate region of genomic DNA in targeted experiments. The candidate region of genomic DNA includes multiple directly observable genetic variants that have been filtered.

[0040] The adjustable parameter L defines the maximum allowable length of each candidate region in the genomic DNA, which is the starting position (Start(T)) of the candidate region in the genomic DNA. k ) and end position End(T kL is the distance between bases. Setting L=1 results in a strategy similar to the pairwise tagging algorithm often used to design standard arrays. In contrast, the present method described herein can use L in the range of 25 to 150. In some embodiments, each candidate region of genomic DNA contains about 25 to about 150 bases and includes at least one variant from a filtered group of directly observable genetic variants. In some embodiments, each candidate region of genomic DNA contains about 35 to about 140 bases and includes at least one variant from a filtered group of directly observable genetic variants. In some embodiments, each candidate region of genomic DNA contains about 45 to about 130 bases and includes at least one variant from a filtered group of directly observable genetic variants. In some embodiments, each candidate region of genomic DNA contains about 55 to about 125 bases and includes at least one variant from a filtered group of directly observable genetic variants. In some embodiments, each candidate region of genomic DNA contains approximately 65 to approximately 125 bases and includes at least one variant from a filtered set of directly observable genetic variants. In some embodiments, each candidate region of genomic DNA contains approximately 75 to approximately 125 bases and includes at least one variant from a filtered set of directly observable genetic variants. In some embodiments, each candidate region of genomic DNA contains approximately 85 to approximately 125 bases and includes at least one variant from a filtered set of directly observable genetic variants. In some embodiments, each candidate region of genomic DNA contains approximately 95 to approximately 125 bases and includes at least one variant from a filtered set of directly observable genetic variants. In some embodiments, each candidate region of genomic DNA contains approximately 105 to approximately 125 bases and includes at least one variant from a filtered set of directly observable genetic variants. In some embodiments, each candidate region of genomic DNA contains approximately 120 to approximately 125 bases.

[0041] In some embodiments, the candidate regions of genomic DNA contain approximately 5 million to 50 million variants. In some embodiments, the candidate regions of genomic DNA contain approximately 10 million to 40 million variants. In some embodiments, the candidate regions of genomic DNA contain approximately 20 million to 30 million variants.

[0042] In some embodiments, the total of multiple candidate regions of genomic DNA contains approximately 1 million to 100 million base pairs. In some embodiments, the total of multiple candidate regions of genomic DNA contains approximately 5 million to 75 million base pairs. In some embodiments, the total of multiple candidate regions of genomic DNA contains approximately 10 million to 50 million base pairs. In some embodiments, the total of multiple candidate regions of genomic DNA contains approximately 20 million to 40 million base pairs.

[0043] In some embodiments, multiple candidate regions of genomic DNA are divided into separate analysis groups. In some embodiments, multiple candidate regions of genomic DNA are divided into separate chromosome analysis groups.

[0044] In some embodiments, multiple candidate regions of genomic DNA include two or more directly observable variants from a filtered set of directly observable genetic variants. For example, a candidate region of genomic DNA containing 120 bases may include four directly observable variants (i.e., V1, V2, V3, and V4). In this scenario, each of the four directly observable variants is present in the region of DNA being probed by the nucleic acid probe set. The 120-base candidate region of genomic DNA may begin at the position of the first variant (i.e., V1…V2…V3…V4…). The 120-base candidate region of genomic DNA may end at the position of the last variant (i.e., …V1…V2…V3…V4). Alternatively, the 120-base candidate region of genomic DNA may begin and end at positions other than these variant positions (i.e., …V1…V2…V3…V4…). Candidate regions of genomic DNA containing 120 base pairs and directly observable variants can be numerous and diverse (i.e., by shifting the start position of the candidate region). Therefore, multiple different candidate regions of genomic DNA containing 120 base pairs can contain the same directly observable variant(s).

[0045] The method also includes calculating a quality score for each candidate region of genomic DNA to estimate the capture efficiency and alignment success of probes hybridizing to it. The quality score can be used to determine which probes (and corresponding candidate regions of genomic DNA) should be avoided. As described above, multiple different candidate regions of genomic DNA containing 120 bases may contain the same directly observed variant(s), and therefore, a quality score is calculated for each of these candidate regions of genomic DNA containing the same directly observed variant(s). Furthermore, a quality score is calculated for each of the other candidate regions of genomic DNA containing different directly observed variant(s). In some embodiments, calculating the quality score includes determining component scores for each of the mapping possibility metrics, insertion-deletion metrics, and classification metrics for the candidate regions of genomic DNA. The quality score aims to enable reverse mapping of probes and subsequent sequencing reads that function well in capturing the appropriate strands of DNA by combining these three types of information, avoiding regions containing insertion-deletion polymorphisms or mutations, and prioritizing regions that function well according to the expected performance of probe hybridization to DNA, which can be estimated as a function of sequence composition and uniqueness. The quality score for each candidate region of genomic DNA is the product of the respective product of the component scores for that candidate region of genomic DNA. The final result is a quality score between 0 and 1 that correlates with the probability of probe success. If any of the component scores are zero, the overall quality score is also zero.

[0046] In some embodiments, the mapping possibility metric (or multi-read mapping possibility metric) is the probability that a randomly selected read of length k in a given region is uniquely mappable. In some embodiments, the mapping possibility metric is the UMAP metric. In some embodiments, the component score of the mapping possibility metric is the multi-read mapping possibility metric (UMAPMRM for position i).i It is an exponential function of 10 times (denoted as). In some embodiments, the component score of the mapping possibility metric is exp(10 × UmapMRM). i -9) and here UmapMRM i This is a multi-read mapping possibility metric for variant position i within a candidate region of genomic DNA. In some embodiments, UMAP mapping metrics, particularly the 100bp multi-read mapping possibility metric, are pre-calculated for the entire genome and compiled into downloadable tables (see the World Wide Web at "bismap.hoffmanlab.org / ").

[0047] In some embodiments, the insertion-deletion metric is a measure of the presence or absence of insertions or deletions of bases (e.g., insertion-deletion polymorphisms or mutations) within candidate regions of genomic DNA. An insertion-deletion is included as if position i were linked to an insertion-deletion mutation, and this position is then downweighted. In some embodiments, the component score for an insertion-deletion mutation is exp(SV score). i ) In some embodiments, if variant position i is not linked to an insertion-deletion mutation, or if it is linked to an insertion-deletion mutation of less than 5 bases, the SV score i The value is 2. In some embodiments, when variant position i is linked to an insertion-deletion mutation of 5 to 10 bases (e.g., a medium-sized insertion-deletion variant), the SV score i The value is 1. In some embodiments, when variant position i is linked to an insertion-deletion mutation of more than 10 bases (e.g., a large-sized insertion-deletion), the SV score i The value is 0. In some embodiments, if the variant location is not near the insertion-deletion variant, the SV score i The value is 2, and if the variant position is near an insertion-deletion variant of ≥5 and <10 bases, the SV score i The SV score is 1, and the variant position is near an insertion-deletion variant of ≥10 bases. iThis value is 0. The tunable parameter can define the maximum length of insertion-deletion polymorphisms contained within candidate regions of genomic DNA. This tunable parameter may depend on the tolerance for mismatch between the probe used for targeting and the sequences present in each sample being investigated.

[0048] In some embodiments, the classification metric for candidate regions of genomic DNA includes a first category (e.g., the worst performing category), a second category (e.g., the poor performing category), a third category (e.g., the insufficient performing category), and a fourth category (e.g., the good performing category). The order from best performance to worst performance is the fourth category, the third category, the second category, and the first category. In some embodiments, the first component score of the classification metric is exp(Region_score). i This is a score based on position, where variant position i for the first category is scored as 0, variant position i for the second category is scored as 1, variant position i for the third category is scored as 1.6, and variant position i for the fourth category is scored as 2. In some embodiments, the second component score, which is the minimum absolute distance score of the classification metric,

[0049]

number

[0050] And here, dist2category1 i This is the minimum absolute distance from the variant position i of the first category to the region. In some embodiments, the third component score of the classification metric is

[0051]

number

[0052] And here, dist2category2i is the minimum absolute distance from variant position i of the second category to the region. These two component scores downweight probes that are very close but not in category 1 or category 2 (i.e., the poor region or the worst region), so that leads produced from the probe may have poor alignment.

[0053] In some embodiments, the trait used to categorize specific candidate regions of genomic DNA may be the %GC content of the corresponding complementary probe / primer. For example, the %GC content of the probe / primer is preferably about 40% to about 55%. Thus, in some embodiments, a first category may have corresponding probes / primers with a %GC content of less than about 40%, a second category may have corresponding probes / primers with a %GC content greater than 55%, a third category may have corresponding probes / primers with a %GC content of about 50% to about 55%, and a fourth category may have corresponding probes / primers with a %GC content of about 40% to about 55%. Additional traits that can be used to categorize specific candidate regions of genomic DNA include, but are not limited to, the melting temperature of the primer / probe, the annealing temperature of the primer / probe, the presence or absence of GC clamping, and the stability of the 3' end. Each of these traits may be divided into four categories based on the user's desired preference tendencies.

[0054] The overall quality score is the product of the multiplication of the five component scores. In some embodiments, the quality score of each candidate region of genomic DNA is the maximum score (exp(5) × 1.2 2 It is scaled to 0-1 by dividing by (or approximately 213.7149), thereby creating a quality score for each candidate region of genomic DNA.

[0055] Regarding the overall quality score, the decision made about which probe to select for a particular candidate region of genomic DNA can be relative. Therefore, regional characteristics that lower the scores of many neighboring probes (such as GC content) do not necessarily mean that the region should be excluded from consideration. Rather, our method attempts to select the best available probe for such a region. Furthermore, the quality score may include metrics that prioritize probes that are evenly distributed across the genome.

[0056] The method described herein also includes calculating a probe score for each candidate region of genomic DNA. In some embodiments, the probe score is calculated by multiplying the quality score by the number of variants captured by the candidate region of genomic DNA. For example, for each candidate region T of genomic DNA k This may overlap with the set of genetic variants, and this is called OverlapSet(T k ) can be called Start(T k ) and End(T k Includes all genetic variants located between ) and each candidate region T of the genomic DNA. k In addition to the variants that directly overlap, OverlapSet(T k This also captures variants that have a proxy in ) and places this set in region T. k This can be called a proxy set, and this is ProxySet(T k ) can be called OverlapSet(T k Not only all variants in ) but also R mn > OverlapSet(T k This includes all other variants m in which a corresponding variant n exists. Thus, in some embodiments, the number of variants captured by a candidate region of genomic DNA is the sum of the number of directly observed variants (i.e., within the candidate region hybridized to the probe) captured by the candidate region of genomic DNA and the number of corresponding proxy variants in different candidate regions of genomic DNA.

[0057] For example, a particular candidate region of genomic DNA contains three directly observed variants (i.e., V1, V2, and V3), where V1 is one of two corresponding proxy variants PV a and PV b V2 has four corresponding proxy variants PV located in different candidate regions of genomic DNA. c , PV d , PV e , and PV f V3 has five corresponding proxy variants PV located in different candidate regions of genomic DNA. g , PV h , PV i , PV j , and PV k Assuming that these variants are present in different candidate regions of genomic DNA, the number of directly observed variants captured by the candidate regions of genomic DNA is 3 (i.e., V1, V2, and V3), and the number of corresponding proxy variants in the different candidate regions of genomic DNA is 11 (i.e., PV a , PV b , PV c , PV d , PV e , PV f , PV g , PV h , PV i , PV j , and PV k Therefore, the sum of the number of directly observed variants captured by the candidate region of genomic DNA and the number of corresponding proxy variants in different candidate regions of genomic DNA is 14. Thus, the probe score of this particular candidate region of genomic DNA is the product of the quality score and 14.

[0058] The method also includes selecting one or more candidate regions of genomic DNA having the highest probe score for inclusion in the final set of genomic DNA regions. In some embodiments, a single candidate region of genomic DNA having the highest probe score is selected for inclusion in the final set of genomic DNA regions. In some embodiments, two or more candidate regions of genomic DNA having the highest probe score are selected for inclusion in the final set of genomic DNA regions. In some embodiments, if there are multiple candidate regions of genomic DNA with the highest probe score, the candidate regions(s) of genomic DNA that are more evenly spaced across the genome are selected.

[0059] When selecting a set of candidate genomic DNA regions to measure experimentally, one goal is to minimize the number of regions within T and achieve an overall probe score (T). k To maximize the overall quality of these regions summarized by ) and the ProxySet(T) of candidate regions of genomic DNA k The goal is to maximize the number of variants captured in the union of the following. If there are multiple sets of candidate genomic DNA regions that function similarly, sets of candidate genomic DNA regions that are evenly spaced across the genome may actually perform better than alternatives, and therefore these evenly spaced sets of candidate genomic DNA regions can be preferred.

[0060] As described herein, one step in the method described herein is to identify a set of candidate regions of genomic DNA to be evaluated. Since the human genome is approximately 3 billion base pairs long, there are 3 × 10⁶ potential candidate regions of genomic DNA of length L. 9There may be about several (when L is small compared to the genome size). The number of potentially selected candidate variants is significantly smaller, typically about 5 to 50 million variants (depending on the allelic frequency range of the variants). In the list of candidate regions of genomic DNA, the recommended candidate regions of genomic DNA for each variant are seeded. The recommended candidate regions of this genomic DNA include this variant and all variants within L base pairs to the right of it. Among all possible candidate regions of genomic DNA that meet this criterion, the focus is on the recommended candidate region of genomic DNA with the highest probe score Score(T k ). Performance can be improved by considering regions that contain only a subset of variants that are L base pairs to the right but have a higher regional probe score. For example, if variant V m as well as three additional variants V m+1 , V m+2 , and V m+3 are all within L base pairs to the right of it. Without loss of generality, the three variants can be sorted from left to right according to their coordinates. V m , V m+1 , V m+2 , and V m+3 can be included to identify the candidate region with the highest possible score. It is also possible to identify the highest-scoring candidate regions that contain only V m , V m+1 , and V m+2 , or only V m and V m+1 . These additional regions have probe scores that are higher than those of V m , V m+1 , V m+2 , and V m+3Only regions with a higher probe score than the highest-scoring region containing the region are added to the list of potential candidate regions of genomic DNA. If these additional regions have low region probe scores, they are never selected and can be safely ignored, because the list of variants they can proxy is always smaller than or equal to the list of regions that higher-scoring regions can proxy. This optional selection step reduces the number of candidate regions of genomic DNA that need to be considered in each iteration from billions to millions, saving a significant amount of computation time.

[0061] In some embodiments, an additional adjustable parameter can be used to define the maximum number of variants allowed for each candidate region of genomic DNA. In some embodiments, if a candidate region of genomic DNA contains more directly observed variants than a desired threshold, the candidate region is removed from the final set of genomic DNA regions. In some embodiments, the desired threshold is five directly observed variants.

[0062] The method also includes repeating steps g) (i.e., calculating a probe score for each candidate region of genomic DNA) and h) (i.e., selecting one or more candidate regions of genomic DNA with the highest probe scores to be included in the final set of genomic DNA regions) for unselected candidate regions of genomic DNA to be included in the final set of genomic DNA regions. Thus, to identify a set of candidate regions of genomic DNA, the method described herein proceeds by repeating a series of steps. In each iteration, one or more candidate regions of genomic DNA are selected to be included in the final set of candidate regions of genomic DNA, and the scores of other candidate regions of genomic DNA are updated. The selection of candidate regions of genomic DNA to be included in the final set of candidate regions of genomic DNA continues until the maximum number of candidate regions of genomic DNA are selected, or until all variants of interest are located within the selected candidate regions of genomic DNA, or have proxies within the selected candidate regions of genomic DNA.

[0063] For example, after the initial selection of one or more candidate regions of genomic DNA as described in the previous step, the remaining unselected candidate regions of genomic DNA become available for recalculating probe scores and for selection to be included in the final set of genomic DNA regions. In iterating through such steps, the number of variants in a particular unselected candidate region of genomic DNA is the sum of 1) the number of directly observable variants in the unselected candidate region of genomic DNA, excluding all directly observable variants in the selected candidate region of genomic DNA, and 2) the number of corresponding proxy variants in different candidate regions of genomic DNA, excluding all proxy variants corresponding to directly observable variants in the selected candidate region of genomic DNA.

[0064] For example, suppose a candidate region 1) from a selected candidate region of genomic DNA (i.e., step h) contains two directly observed variants (i.e., V1 and V2). Also, V1 contains two corresponding proxy variants PV a and PV b V2 has two corresponding proxy variants PV in different candidate regions of genomic DNA. c and PV d Assume that the genomic DNA has different candidate regions. Also, candidate region 2 being considered for selection has two directly observed variants (i.e., V2 and V3), where V2 has two corresponding proxy variants PV c and PV d V3 has two corresponding proxy variants PV in different candidate regions of genomic DNA. e and PV fAssume that the variants are located in different candidate regions of the genomic DNA. If candidate region 2 is considered for selection, the number of directly observed variants in unselected candidate region 2 excludes all directly observed variants in the selected candidate region of genomic DNA (i.e., V2 from candidate region 1), and the number of corresponding proxy variants in different candidate regions of genomic DNA is all proxy variants corresponding to the directly observed variants in the selected candidate region of genomic DNA (i.e., proxy variant PV associated with V2 from candidate region 1). c and PV d ) are excluded. Therefore, in the scenario described herein, candidate region 2 contains two directly observed variants (i.e., V2 and V3), but only one of them (i.e., V3) is counted for the number of directly observed variants to determine the probe score. Furthermore, candidate region 2 contains four proxy variants (i.e., PV c , PV d , PV e , and PV f ) includes, but only two of them (i.e., PV e and PV f ) is counted against the number of corresponding proxy variants to determine the probe score. Therefore, in the current scenario, the probe score for candidate region 2 is not the product of the product of the quality scores of candidate regions 2 and 6 (i.e., the sum of two directly observed variants and four corresponding proxy variants), but rather the product of the product of the quality scores of candidate regions 2 and 3 (i.e., the sum of a single directly observed variant and two corresponding proxy variants that are not yet present in any of the candidate regions of the selected DNA).

[0065] In some embodiments, after steps g) (i.e., calculating a probe score for each candidate region of genomic DNA) and h) (i.e., selecting one or more candidate regions of genomic DNA with the highest probe scores to include in the final set of genomic DNA regions) are repeated, the probe scores of the remaining unselected candidate regions of genomic DNA are updated.

[0066] In some embodiments, the update includes selecting candidate genomic DNA regions to include in the final set of genomic DNA regions, and then recalculating all probe scores for the remaining unselected genomic DNA candidate regions, including proxies of directly observed variants that were present in the selected genomic DNA candidate regions. In some embodiments, the update includes eliminating all unselected genomic DNA candidate regions that contain only directly observed variants and / or corresponding proxy variants that were already selected in the previous round of selection to include in the final set of genomic DNA regions. In some embodiments, the update includes both of the updates described above.

[0067] In some embodiments, steps g) and h) are repeated until the maximum number of genomic DNA regions are selected. In some embodiments, steps g) and h) are repeated until all directly observable variants and proxy variants are included in the final set of genomic DNA regions.

[0068] All potential candidate regions of genomic DNA are repeated in each repeat. k Probe score (T k The increment value as the product of ) and the proxy set that is not in the selected region's proxy set ProxySet(T kThe number of variants within the region is measured. One goal is to identify and select the candidate genomic DNA region with the highest incremental value. If ties exist, the distance between the candidate genomic DNA region with the largest product and all the selected candidate genomic DNA regions and the tie is resolved by selecting the candidate genomic DNA region that is furthest from the selected candidate genomic DNA regions. This tie-resolving strategy promotes even spacing of selected candidate genomic DNA regions across the genome and improves methodological performance when combined with modern haplotyping and complementation methodologies for analyzing the resulting candidate genomic DNA regions and data.

[0069] After selecting the candidate genomic DNA region with the highest increment value and resolving ties as necessary, information about the remaining candidate genomic DNA regions may be updated. For example, two optional updates may be considered. First, the number of variants in the proxy set of each candidate genomic DNA region that are not present in the proxy set of selected candidate genomic DNA regions can be cached. This caching is not mandatory but significantly improves computational efficiency. If caching is enabled, a specific candidate region T of genomic DNA can be cached. k After selecting, the proxy set is ProxySet(T k ) allows access to all overlapping regions, and the cached count of the number of variants in the proxy set that are not within the candidate region of the selected genomic DNA is updated, and some of the variants in the proxy set are within the candidate region T of the selected genomic DNA. k This reflects the fact that it is now being captured by. Secondly, if the probe score of each candidate region of genomic DNA depends on the probe score of other selected candidate regions of genomic DNA (for example, because the targeting technique used does not allow for overlapping regions, or because sequence complementarity between candidate regions of the targeted genomic DNA must be considered), then the probe score of other candidate regions of genomic DNA will depend on the probe score of candidate region T of genomic DNA. k It can be updated to reflect that it has been selected.

[0070] Before starting the next iteration, all candidate genomic DNA regions whose proxy set is empty or which are entirely contained within the union of proxy sets of currently selected candidate genomic DNA regions may be removed from the list of candidate genomic DNA regions to be evaluated. If caching is implemented, these regions will have a cache score of zero. These regions can never be selected as they do not improve the design and can be safely removed from the list of candidate genomic DNA regions to be evaluated to improve computational efficiency and speed up future iterations. Furthermore, candidate genomic DNA regions with a cache score of 1 (i.e., capturing only a single incremental variant) can be safely reserved for evaluation in the final custom iteration if the captured variant is not captured by any other candidate genomic DNA region. This methodology can proceed iteratively, selecting one candidate genomic DNA region at a time, until all variants are contained within the proxy set of one candidate genomic DNA region selected for targeting, or until the maximum number of candidate genomic DNA regions have been targeted.

[0071] The methods described herein can be incorporated into algorithms. Additional information can also be used to improve the computational efficiency of algorithms. For example, a challenging aspect of such algorithms might be the storage of a matrix R. If the number of variants considered M is large, the number of entries in this matrix, proportional to M × M, becomes very large and may exceed the capacity of the random access memory (RAM) of most modern computers. In such situations, a sparse representation of the matrix can be used, using only entries whose values ​​exceed a user-defined threshold t that establishes proxy relationships loaded into RAM. In typical human data, large chain disequilibrium coefficients are limited to a small number of variant pairs, and this sparse representation of the matrix can be easily stored in memory and used for the necessary computations.

[0072] Furthermore, while the algorithm may be efficient enough to apply directly to the entire genome, selecting candidate regions of genomic DNA for targeting can be considered, particularly in situations where it does not affect the probe scores of other distant candidate regions of the genomic DNA under consideration, and several efficiencies may be improved. One of these efficiencies is to partition the genome into a series of regions where candidate regions of genomic DNA can be selected independently. In the simplest case, these regions can be individual chromosomes. In a more refined case, the genome can be partitioned into a series of non-overlapping regions such that R mn is guaranteed to be <t> when indexing variants within regions where m and n are different. This partitioning can be done using standard algorithms to identify connected components within a graph. The partitioning improves computational efficiency and allows the algorithm to consider pairs, triples, or other small tuples of candidate regions of genomic DNA for each iteration rather than just one candidate region of genomic DNA for each iteration.

[0073] Iterative algorithms can provide very high-quality solutions that take known linkage disequilibrium relationships into account, prioritize clustered groups of variants that can be targeted together because they fall within a contiguous window of L base pairs or less, allow probe scores for candidate regions of genomic DNA, and distribute probes evenly across the genome, all of which can be achieved in a computationally efficient manner. When the number of candidate regions of genomic DNA is reasonable (or when an algorithm is used that divides the genome into blocks that can be considered independently), it is possible to exhaustively enumerate and evaluate all possible combinations of candidate regions of genomic DNA. In this case, a global scoring scheme can be used to select the optimal combination of candidate regions of genomic DNA from all the enumerated possibilities. To do this, a global scoring scheme can summarize the number of variants with proxies within the candidate regions of genomic DNA, the overall probe score of the candidate regions of genomic DNA, and the even spacing of the candidate regions of genomic DNA. Given a set T of candidate regions of genomic DNA, many suitable scoring schemes can be devised. Each variant of interest may be assigned a probe score to the candidate genomic DNA region with the highest score among the candidate genomic DNA regions in the selected genomic DNA region that includes the variant in the proxy set. Variants not included in any proxy set may be assigned a score of zero. The overall global score for each configuration may then be a weighted sum of the scores assigned to each of these variants (sum of all variants), a measure of the uniformity of the spacing between candidate genomic DNA regions, such as the kurtosis of the distribution of distances between consecutive selected probes, and a penalty for favoring configurations with fewer targets. This global scoring scheme can also be used in conjunction with pseudo-annealing or another Monte Carlo algorithm to refine the iterative solution recommended by the algorithm. This refinement may be possible even when the set of all possible combinations of candidate genomic DNA regions is too large to enumerate.Like other Monte Carlo schemes, pseudo-annealing requires a proposal scheme to explore solutions in the neighborhood of the current solution and recommend new solutions in the neighborhood of the current solution (e.g., by adding, removing, or substituting candidate regions of genomic DNA in the currently selected set); a scheme to probabilistically accept or reject the proposed updates (e.g., by always accepting solutions that improve the global score and occasionally accepting solutions that decrease the global score to avoid being tied to a local minimum); and a scheme to manage the probabilistic component of the process so that the process becomes progressively stringent and to determine when convergence has been achieved.

[0074] The present method also optionally includes generating a set of nucleic acid probes. Each individual probe in the set of nucleic acid probes is complementary to the nucleic acid sequence of a genomic region in the final set of selected genomic DNA regions. Thus, the entire set of nucleic acid probes is complementary to the entire nucleotide sequence of the final set of selected genomic DNA regions. In some embodiments, the set of nucleic acid probes includes about 200,000 to about 700,000 probes. In some embodiments, the set of nucleic acid probes includes about 200,000 to about 600,000 probes. In some embodiments, the set of nucleic acid probes includes about 200,000 to about 500,000 probes. In some embodiments, the set of nucleic acid probes includes about 200,000 to about 400,000 probes. In some embodiments, the set of nucleic acid probes includes about 500,000 to about 700,000 probes. In some embodiments, the set of nucleic acid probes includes about 600,000 to about 650,000 probes. In some embodiments, each individual probe in a set of nucleic acid probes contains about 25 to about 150 bases and is hybridizable to a specific candidate region of genomic DNA containing at least one directly observable variant. In some embodiments, each individual probe in a set of nucleic acid probes contains about 120 to about 125 bases. In some embodiments, one or more individual probes in a set of nucleic acid probes contain the same number of bases as the corresponding candidate region of genomic DNA to which it is designed to hybridize. In some embodiments, one or more individual probes in a set of nucleic acid probes contain a larger number of bases than the corresponding candidate region of genomic DNA to which it is designed to hybridize.

[0075] The disclosure also provides a method for genotyping a DNA sample by sequencing, comprising: a) hybridizing a DNA sample with a set of nucleic acid probes manufactured as described herein to generate probe-hybridized genomic DNA; b) sequencing the probe-hybridized genomic DNA to create a plurality of sequencing reads; c) mapping the plurality of sequencing reads to a reference genome; d) calling directly observable variants present in the mapped sequencing reads; and e) supplementing unobserved variants from unsequenced regions of the genomic DNA to establish the genotype of the sample DNA.

[0076] The DNA sample can be any DNA sample that serves as the DNA source for genotyping. In some embodiments, the DNA sample is obtained from a subject having a disease or condition. In some embodiments, the DNA sample is obtained from a tumor of the subject.

[0077] The present method involves hybridizing a set of nucleic acid probes, prepared as described herein, to a DNA sample to generate genomic DNA hybridized to the probes. The set of nucleic acid probes is brought into contact with the DNA sample under typical conditions under which hybridization occurs. In some embodiments, if the average probe yields a coverage of X, probes with coverage < 0.33X may be removed. Thus, for example, any probes that yield less than 8X coverage of directly observable variants among multiple sequencing reads (if the average probe has a coverage of 24X) are removed from the set of nucleic acid probes. In some embodiments, any probes that result in inefficient capture of sample DNA are removed from the set of nucleic acid probes. In some embodiments, probes that yield low average coverage but target high-value variants (for mapping to known functional regions of the genome or to act as proxies for many other variants) may be supplemented with additional copies in the capture reagent rather than being truncated. This supplementation may help improve the coverage they provide and facilitate accurate genotyping.

[0078] The method also includes sequencing the genomic DNA hybridized to the probe to create multiple sequencing reads. In some embodiments, the multiple sequencing reads include approximately 30 million sequencing reads. In some embodiments, the multiple sequencing reads include approximately 25 million sequencing reads. In some embodiments, the multiple sequencing reads include approximately 20 million sequencing reads. In some embodiments, the multiple sequencing reads include approximately 15 million sequencing reads. In some embodiments, the multiple sequencing reads include approximately 10 million sequencing reads. In some embodiments, the multiple sequencing reads include approximately 5 million sequencing reads. In some embodiments, the multiple sequencing reads include approximately 1 million sequencing reads.

[0079] This method also involves mapping multiple sequencing reads to a reference genome. The method in question also includes calling directly observed variants present in the mapped sequencing reads. In some embodiments, low-confidence called variants resulting from reads with low coverage are eliminated to create a final set of called directly observed variants. In some embodiments, low-confidence called variants resulting from reads with coverage less than 8X are eliminated. In some embodiments, eliminating low-confidence called variants includes supplementing the same called directly observed variants from the variant reference panel.

[0080] In some embodiments, the method further includes phasing a called directly observed variant into a set of known haplotypes. An example of phasing can be found, for example, in U.S. Patent Application Publication 2019 / 0205502.

[0081] In some embodiments, the software GLIMPSE (see the World Wide Web at "odelaneau.github.io / GLIMPSE / ") or software providing the same functionality can be used to return a refined variant call after incorporating information from neighboring variants. Given neighboring variant calls for each sample, GLIMPSE can significantly reduce the uncertainty of variant calls from low-coverage reads. The second step of GLIMPSE is to obtain these refined variant calls and phase the genotype calls into chromosome-specific variant calls. GLIMPSE can be run using default parameters.

[0082] In some embodiments, the percentage of called variants having coverage greater than 10X is determined. In such embodiments, if the percentage of called variants having coverage greater than 10X is less than about 95%, the set of nucleic acid probes is re-hybridized to the DNA sample. This embodiment serves as an internal control for the hybridization and sequencing steps described herein.

[0083] In some embodiments, if a called directly observable variant is near or within a region of genomic DNA that can hybridize to a probe excluded from the set of nucleic acid probes, such directly observable variant is removed from the final set of called directly observable variants.

[0084] The method described here also includes supplementing unobserved variants from unsequenced regions of genomic DNA, thereby establishing the genotype of the sample DNA. In some embodiments, unobserved variants are supplemented from a variant reference panel based on the presence of directly observed variants that have been called in the DNA sample.

[0085] In some embodiments, the software Minimac3 (see the World Wide Web entry at "genome.sph.umich.edu / wiki / Minimac3") may be used for variant completion from variant calls of each haplotype (for variants that have not been observed or sequenced). Minimac3 can be implemented using default parameters.

[0086] The disclosure also provides a method for genotyping a DNA sample by sequencing using a set of nucleic acid probes, comprising: a) selecting multiple regions of genomic DNA from a DNA sample containing multiple directly observable genetic variants; b) identifying a set of nucleic acid probes for hybridization of the selected genomic DNA to multiple regions; c) hybridizing the set of nucleic acid probes to the DNA sample to generate genomic DNA hybridized to the probes; d) sequencing the genomic DNA hybridized to the probes to create multiple sequencing reads; e) mapping the multiple sequencing reads to a reference genome; f) calling the directly observable variants present in the mapped sequencing reads; and g) supplementing unobserved variants from unsequenced regions of the genomic DNA, thereby establishing the genotype of the sample DNA. Steps a) through g) can be carried out according to the disclosure herein.

[0087] This disclosure also provides a system and computer-readable media for carrying out the methods described herein. In some embodiments, a computer program product is provided, comprising a computer-readable medium containing encoded instructions for performing any of the methods described herein. In some embodiments, the computer program product can cause a computer having a processor to perform any of the methods described herein. In some embodiments, the computer program product is encoded such that the program can receive all the parameters necessary to perform any of the methods described herein when implemented by a suitable computer or system. In some embodiments, a computer system is provided for performing any of the methods described herein, comprising a processor and memory connected to the processor, the memory encoding one or more computer programs that cause the processor to perform any of the methods described herein.

[0088] Computer software products can be written using any suitable programming language known in the art. System components may include any suitable hardware known in the art. Suitable programming languages ​​and suitable hardware system components are described in U.S. Patent No. 7,197,400 (see, e.g., columns 8-9), U.S. Patent No. 6,691,042 (see, e.g., columns 12-25); U.S. Patent No. 8,245,517 (see, e.g., columns 16-17); U.S. Patent No. 7,272,584 (see, e.g., column 4, lines 26-5, line 18); U.S. Patent No. 8,2 This includes what is described in U.S. Patent No. 03,987 (see, for example, columns 19-20); U.S. Patent No. 7,386,523 (see, for example, column 2, lines 26-3, line 3; also see column 8, lines 21-9, line 52); U.S. Patent No. 7,353,116 (see, for example, column 5, lines 50-8, line 5); and U.S. Patent No. 5,985,352 (see, for example, column 31, lines 37-32, line 21).

[0089] In some embodiments, a computer system capable of performing the computer implementation methods described herein comprises a processor, a fixed storage medium (i.e., a hard drive), system memory (e.g., RAM and / or ROM), a keyboard, a display (e.g., a monitor), a data input device (e.g., a device capable of providing the system with raw or converted microarray data), and optionally, a drive capable of reading and / or writing to a computer-readable medium (i.e., a removable storage device, e.g., a CD or DVD drive). The system also optionally comprises a network input / output device and a device enabling connectivity to the Internet.

[0090] In some embodiments, computer-readable instructions (e.g., computer software products) (i.e., software for performing any of the method steps described herein) that enable the system to perform any of the methods described herein are encoded on a fixed storage medium, and enable the system to display the results to a user, or to provide the results to a second set of computer-readable instructions (i.e., a second program), or to transmit the results to a data structure residing on the fixed storage medium, or to another network computer, or to a remote location via the Internet.

[0091] Examples are provided below to allow for a more efficient understanding of the subject matter disclosed herein. These examples are for illustrative purposes only and should not be construed as limiting the claimed subject matter in any way. [Examples]

[0092] Example 1: Pilot Study After selecting directly observable variants, selecting candidate regions of genomic DNA containing the selected directly observable variants, and selecting a probe set as described herein, a pilot study was conducted.

[0093] Forty-eight samples were selected from a 1KG sample set, and samples of these DNAs were accessed from Coriell (see the World Wide Web at "coriell.org / 1 / NHGRI / Collections / 1000-Genomes-Collections / 1000-Genomes-Project"). In this example, the 48 samples were treated as completely new and processed with the sequencing-based genotyping probe set described herein. The results of sequencing-based genotyping of the 48 samples were compared to control results obtained from whole-genome sequencing with 30X coverage (after filtering). The reference panel was considered to be 1KG WGS data excluding the 48 samples.

[0094] The pilot set of samples was selected to be diverse. One sample was excluded because it did not have enough DNA for sequencing, so 47 samples remained for testing. The samples are summarized in Table 1.

[0095] [Table 1]

[0096] The primary objective was to determine how well the probes actually performed (i.e., whether the probe set captured sequences specific to the target locations in the genome). Two reasons were considered for excluding certain probes from the initial probe set: 1) the variant coverage was too low for some DNA samples to produce no signal, and 2) many reads were shown not to readily map to the genome at the sites captured by that probe. The overall goal was to eliminate probes that resulted in inefficient capture and probes that did not provide sufficient signaling for the desired variants. Many probes fell into both categories. As a result, approximately 14,000 probes were identified that achieved too little coverage.

[0097] Computational experiments showed that the excluded probes did not make a significant difference in the overall completion performance, and this data was observed by filtering the WGS experiments to represent what could be observed.

[0098] Another objective was to determine whether the information extracted from the sequencing reads could supplement the directly observed variants and enable the completion of other variants. To evaluate the accuracy of the completion, two processes were performed: 1) Variants that were close to or within the excluded probes were removed from the called variants; and 2) The remaining called variants were processed to return the completed variants (for all of the estimated 15 million variants).

[0099] Data preparation methods - Variant calls for supplementation A new haplotype reference set was used to perform complementation on the pilot sample. The reference was a 1KG WGS dataset with the pilot sample removed. This new reference data was used twice: 1) with the GLIMPSE program to improve variant calling and phasing, and 2) with the Minimac3 program for variant complementation. The complemented variant calls were then compared to variant calls observed directly from whole-genome sequencing.

[0100] Evaluation of Complementary Quality To assess the quality of complementation, we evaluated the squared correlation between the directly observed genotype and the complemented genotype. This metric is commonly referred to as "complementary Rsq" or "r 2Called the "scale" or "r-squared," it is the square of the correlation coefficient between the true genotype and the experimentally derived counterpart, estimated from the complementation. When r² is 1.0, these two are identical. When it is close to 0.0, the experimentally derived counterpart is equivalent to a blind estimate. Specifically, from whole-genome sequencing data, we created genotype vectors of directly observed genotypes, encoded as 0 if the genotype is for two reference alleles, as 1 if the genotype is for one reference and one substitute allele, and as 2 if the genotype is for two reference alleles. For the complemented genotype vectors, this was different because each of the three states has a probability. For example, there could be an 80% probability of being 0, a 20% probability of being 1, and a 0% probability of being 2. For the complemented genotype vectors, 0.8 * 0+0.2 * 1+0 * From 2, a predicted genotype of 0.2 was returned.

[0101] Pearson's correlation coefficient was used with two vectors. It was noted that there were only 47 samples per genotype. To improve the measurement across variants, variants were pooled together by allele frequency (to ensure all had the same expected genotype), and vector correlations were performed between samples and variants. This complementary Rsq process followed standard methods.

[0102] Figure 1 shows the Rsq of the difference frequency bins complemented by complementation from different observational data. The highest correlation (and best complementation) occurred when whole-genome sequencing was filtered to observe only variants within the selected probe region. The line thus formed represented the desired best performance. The blue line represents the global screening array directly assayed with these samples (performed in-house under standard protocols). Complementation from the pilot study was desired to be at least as good as that of the global screening array. The green line represents the complementation quality of the directly observed sequencing-based genotyping design after the processing described herein. The sequencing-based genotyping design performed significantly better than the global screening array and, given the selected probes, came close to the desired best performance. This pilot study demonstrates that sequencing-based genotyping design can perform better than the global screening array at a reasonable cost. The pilot study was not merely a simulation study, but a direct comparison of the performance of two assays from DNA samples to complementation comparison. Finally, we compared sequencing-based genotyping designs to a very large array called the MEGA array (Multi-Ethnic Genotyping Array), which has three times the number of variants as the global screening array. When simulating the array by fully observing all the variants that the array would assay from the whole-genome sequencing version of the pilot data, the sequencing-based genotyping designs performed similarly to the best possible performance of the MEGA array. In fact, the MEGA array performed lower. The sequencing-based genotyping designs performed similarly to the MEGA array at a cost comparable to the global screening array (one-third the cost of the MEGA array). Therefore, sequencing-based genotyping designs provide a very cost-effective strategy for assaying genetic information and functioned well to provide high-quality complementarity.

[0103] Example 2: Genotyping by Sequencing For 223,266 samples, each evaluated for coverage at a design site, sequencing-based genotyping assays were successfully performed. Call rate is the percentage of regions containing actionable genotypes. Figure 2 shows the average call rate of 98.9% and 99.3% of the samples with a call rate of 95% or higher.

[0104] In addition to those described herein, various modifications of the subject matter described herein will be apparent to those skilled in the art from the foregoing description. Such modifications are also intended to be included in the appended claims. Each reference cited herein (including, but not limited to, academic journal articles, U.S. and non-U.S. patents, patent application publications, international patent application publications, gene bank acceptance numbers, etc.) is incorporated herein by reference in its entirety.

Claims

1. A computer implementation method for generating information about nucleic acid probes for genotyping by sequencing, a) Selecting multiple directly observable genetic variants to be captured by the nucleic acid probe using a computer, b) Using the computer to remove low-confidence variants from the plurality of directly observable genetic variants, thereby creating a filtered plurality of directly observable genetic variants. c) The computer phases the filtered plurality of directly observable genetic variants, d) The computer identifies, for each of the filtered direct-observable genetic variants, the presence or absence of one or more proxy variants. e) Selecting a plurality of candidate regions of genomic DNA containing the filtered plurality of directly observable genetic variants by the computer, wherein each candidate region of genomic DNA contains 25 to 150 bases and contains at least one variant from the filtered plurality of directly observable genetic variants. f) The computer calculates a quality score for each candidate region of genomic DNA to estimate the probe capture efficiency and alignment success. g) The computer calculates a probe score for each candidate region of genomic DNA by multiplying the quality score by the number of variants captured by the candidate region of genomic DNA, wherein the number of variants captured by the candidate region of genomic DNA is the sum of the number of directly observed variants captured by the candidate region of genomic DNA and the number of corresponding proxy variants in different candidate regions of genomic DNA. h) The computer selects one or more candidate regions of genomic DNA having the highest probe score to be included in the final set of genomic DNA regions. i) Repeating steps g) and h) for unselected candidate regions of genomic DNA to be included in the final set of regions of genomic DNA by the computer, wherein the number of variants in the unselected candidate regions of genomic DNA is the sum of 1) the number of directly observable variants in the unselected candidate regions of genomic DNA, excluding all directly observable variants in the selected regions of genomic DNA, and 2) the number of corresponding proxy variants in different candidate regions of genomic DNA, excluding all proxy variants corresponding to directly observable variants in the selected regions of genomic DNA, and repeating steps g) and h) until the maximum number of regions of genomic DNA are selected, and j) The computer generates a set of nucleic acid probes complementary to the nucleic acid sequence of each genomic region in the final set of genomic DNA regions. Computer implementation methods, including those mentioned above.

2. The method according to claim 1, wherein the plurality of directly observable genetic variants are selected from a genome-wide association database of genetic variants, a database of genetic pharmacological associations of genetic variants, a database containing genetic variants within all mitochondrial chromosomes, and / or a database of genetic variants within microarrays, or any combination thereof.

3. The method according to claim 2, wherein, with respect to a variant in the genome-wide association database of the genetic variant, if the square of the association between the variant and the trait has a p-value ≤ 10⁻⁹, the variant is retained in the plurality of directly observed genetic variants, and if the square of the association between the variant and the trait has a p-value > 10⁻⁹, the variant is excluded from the plurality of directly observed genetic variants.

4. The method according to claim 2, wherein the database of genetic variants in the microarray includes genetic variants in the HLA region of chromosome 6, the Y chromosome, two KIR regions on chromosome 19, and pseudoautosomal regions 1 and 2 (Par1 and Par2) on the X chromosome.

5. The method according to any one of claims 1 to 4, wherein a multi-allergen variant is converted into one or more sets of biallergen variants.

6. The method according to any one of claims 1 to 5, wherein excluding low-confidence variants from the plurality of directly observed genetic variants includes excluding all variants having minor allele frequencies (MAFs) below a desired threshold.

7. The method according to claim 6, wherein the threshold for the minor allele frequency is 1%.

8. The method according to any one of claims 1 to 5, wherein excluding low-confidence variants from the plurality of directly observed genetic variants includes excluding all variants with missing values ​​exceeding a desired threshold.

9. The method according to claim 8, wherein the threshold for missing data is 2%.

10. The method according to any one of claims 1 to 9, wherein when the directly observable genetic variant and the proxy variant are within 1 MB of each other, and the linkage disequilibrium between the two variants has a squared correlation of at least 0.2, at least 0.5, at least 0.8, at least 0.9, or at least 1.0 using the r² scale of the linkage disequilibrium, one of the filtered directly observable genetic variants has a corresponding proxy variant in another candidate region of genomic DNA.

11. The method according to any one of claims 1 to 10, wherein multiple candidate regions of the genomic DNA are divided into separate analysis groups, so that each chromosome is a separate analysis group.

12. The method according to any one of claims 1 to 11, wherein each candidate region of genomic DNA contains 120 to 125 bases.

13. The method according to any one of claims 1 to 12, wherein the plurality of candidate regions of the genomic DNA include 5 million to 50 million variants.

14. The method according to any one of claims 1 to 13, wherein the entirety of the multiple candidate regions of the genomic DNA comprises 1 million to 100 million base pairs, 5 million to 75 million base pairs, 10 million to 50 million base pairs, or 20 million to 40 million base pairs.

15. The method according to any one of claims 1 to 14, wherein calculating the quality score involves determining component scores for each of the mapping feasibility metric, insertion-deletion mutation metric, and classification metric for the candidate region of the genomic DNA, and the quality score is the product of the multiplication of each of the component scores.

16. The component score of the mapping possibility metric is exp(10 × UmapMRM) i -9) and here, UmapMRM i The method according to claim 15, wherein is a multi-read mapping possibility metric for variant position i within the candidate region of the genomic DNA.

17. The insertion-deletion mutation metric is a measure of the presence or absence of base insertions or deletions within a candidate region of the genomic DNA, and the component score of the insertion-deletion mutation is exp(SV score i ) and here, if the variant position i is not linked to an insertion-deletion mutation, or if it is linked to an insertion-deletion mutation of less than 5 bases, the SV score i If the variant position i is 2, and the variant position i is linked to an insertion-deletion mutation of 5 to 10 bases, then the SV score i If is 1, and the variant position i is linked to an insertion-deletion mutation of more than 10 bases, then the SV score i The method according to claim 16, wherein is 0.

18. The classification metric of the candidate region of the genomic DNA includes a first category, a second category, a third category, and a fourth category, and the first component score of the classification metric is exp(Region_score i ), whereby the variant position i of the first category is scored as 0, the variant position i of the second category is scored as 1, the variant position i of the third category is scored as 1.6, the variant position i of the fourth category is scored as 2, and the second component score of the classification metric is (1 + 1.2(min(dist2category1 i , 60) / 60)), where dist2category1 i is the minimum absolute distance from the variant position i of the first category to the region, and the third component score of the classification metric is (1 + 1.2(min(distcategory2 i , 60) / 60)), where distcategory2 i is the minimum absolute distance from the variant position i of the second category to the region. The method according to claim 15.

19. Selecting one or more candidate regions of the genomic DNA having the highest probe score, The method according to any one of claims 1 to 18, comprising selecting a candidate region having the highest probe score from among candidate regions having up to five variants.