Method for determining pharmacogenetic star alleles and rare genetic variants from high-throughput sequencing data

A standardized method for pharmacogenetic analysis standardizes results from multiple tools by right-shortening allele designations and determining diplotype frequencies, addressing inconsistency and improving accuracy for pharmacogenetic star alleles and rare variants, enhancing personalized medicine.

WO2025252520A1PCT designated stage Publication Date: 2025-12-11ROBERT BOSCH FUR MEDIZINISCHE FORSCHUNG MBH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/064517
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-03
Filing Date
2025-05-26
Publication Date
2025-12-11

AI Technical Summary

Technical Problem

Current pharmacogenetic analysis tools yield inconsistent and non-comparable results due to differing technical approaches, gene and allele implementations, and accuracy issues, hindering reliable determination of pharmacogenetic star alleles and rare genetic variants from high-throughput sequencing data.

Method used

A computer-implemented method standardizes result files from multiple tools by right-shortening allele designations, identifying diplotypes, determining diplotype frequencies, and applying confidence criteria to achieve consistent results, while also filtering rare variants for functional prediction using bioinformatics tools.

Benefits of technology

Enhances data comparability and accuracy by standardizing results across different tools, enabling reliable determination of pharmacogenetic star alleles and providing functional predictions for rare variants, thus improving personalized medicine decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000013_0001
    Figure IMGF000013_0001
Patent Text Reader

Abstract

The invention relates to a computer-implemented method for determining pharmacogenetic star alleles from high-throughput sequencing data, the method comprising the following steps: providing a plurality of result files which have been output by a plurality of different computer programs for genotyping and which each have a plurality of result elements, wherein each result element has at least one gene designation; shortening the allele designation as far as possible on the right for each result element of the plurality of result elements from each result file of the plurality of result files; and outputting the gene designation and a specific result diplotype. The invention also relates to a computer-implemented method for determining rare genetic variants. Finally, the invention also relates to a device having a processor and a memory in order to carry out such methods.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Description

[0002] title

[0003] Methods for determining pharmacogenetic star alleles and rare genetic variants from high-throughput sequencing data

[0004] State of the art

[0005] Clinical pharmacogenomics investigates individual genetic variants that influence the efficacy and safety of drugs. By considering these genetic characteristics, personalized treatment approaches can be developed, enabling tailored drug therapy.

[0006] Genetic variants of enzymes and transporters primarily affect the transport and metabolism of drugs in the liver, and thus also drug levels in the blood. Deviations in concentrations outside the therapeutic window can reduce the optimal effect of a drug and cause undesirable side effects. This interindividual variability can be predicted by genetic information. A variant is present when there is a deviation from the current genome or a specifically selected reference genome.

[0007] Evidence-based pharmacogenetic guidelines, such as those from CPIC (Clinical Pharmacogenetics Implementation Consortium) and DPWG (Dutch Pharmacogenetics Working Group), allow for the prior provision of pharmacogenomic information to help minimize adverse drug reactions and improve drug efficacy through dose adjustments or drug selection. Databases and websites (e.g., PharmVar or PharmGKB) collect information on these known variants and their influence on medications. For example, PharmVar lists 2,603 ​​published pharmacogenomic variants (as of February 28, 2024). In the context of personalized cancer therapies, for instance, genetic information is now being analyzed using high-throughput methods (e.g., next-generation sequencing, NGS) to specifically search for individual somatic and germline variants for drug treatment. Bioinformatics pipelines have been developed and established for this purpose (e.g.,within the framework of the NCT MASTER program), which process the NGS raw data up to the reports of the individual variants (Single-Nucleotide Variants, SNV; small insertions and deletions, Indels; and copy number variations, CNVs).

[0008] Predictions regarding drug sensitivity, resistance, adverse drug reactions (ADRs), or toxicity are currently limited. The field of pharmacogenetics investigates germline variants and their association with ADRs and toxicity. Variants of germline genes encoding hepatic enzymes and transporters, for example, influence drug metabolism and transport, which can prevent the therapeutic optimum of a drug from being achieved and, in the worst case, lead to treatment failure.

[0009] For more than 200 medications (as of April 10, 2024), pharmacogenetic guidelines are available that provide recommendations for medication and dosage based on genetic information. These recommendations are based on the expressions or phenotypes of the most important pharmacogenes (e.g., CYP2C9, CYP2C19, CYP2D6, CYP3A5, DPYD, SLCO1B1, TPMT, UGT1A1). These must be determined via haplotypes, i.e., condensed information from several single nucleotide polymorphisms, and translated into drug-specific categories such as slow ("poor"), normal ("normal / extensive"), or rapid ("rapid / ultrarapid") metabolizers. This status determines the dosage and / or drug recommendation according to the guideline.

[0010] Unlike a specific molecular genetic analysis of individual SNVs, the NGS method identifies all SNVs of a gene or DNA segment. Because the DNA structure of pharmacogenes is complex (due to homologies) in addition to known copy number variations (CNVs), special computer programs or software solutions are required for processing NGS data. These are partly freely available, based on different algorithms, and cover different variants and genes. Due to the complex data that must be processed when determining pharmacogenetic star alleles and rare genetic variants, evaluation can only be performed with the help of specialized computer programs and not by a human. Well-known computer programs include Aldy (https: / / github.com / OxTCG / aldy), Stargazer (https: / / stargazer.gs.washington.edu / stargazerweb / ), Cyrius (https: / / github.com / lllumina / Cyrius), and PyPGX (https: / / github.com / sbslee / pypgx). These computer programs were essentially developed independently of each other and follow different approaches to determining pharmacogene alleles.

[0011] Disclosure of the invention

[0012] The inventors based their invention on the observation that medical facilities and institutes typically choose one of the computer programs and then stick with it. This allows for data comparability and the creation of historical records. Furthermore, since each computer program is continuously developed and its operation is constantly evolving, at least slightly, the effort required for adapting to subsequent processes and, if necessary, converting earlier results into a more current presentation format remains manageable.

[0013] However, during the development of the invention, the inventors noticed that each of the computer programs offers specific advantages, but also specific disadvantages. For example, a first computer program might deliver better and more reliable results than a second computer program in a first analysis scenario, whereas the second computer program might provide better and more reliable results in a second analysis scenario.

[0014] Specifically, the inventors recognized that different computer programs have implemented different genes and alleles. This means that the list of variants extracted from the NGS data can differ.

[0015] Furthermore, such differences can change again with each update cycle. The inventors also found that computer programs differ in whether and how CNV determination is implemented for each gene. The computer programs can differ in "phasing," that is, in determining parental lineage and the arrangement of alleles in a chromosome pair.

[0016] It was eventually recognized that the computer programs differed in their accuracy. Further work revealed that, due to these problems, the additional use of another computer program was currently not possible, as the programs, due to differing technical approaches, did not always achieve the same results and even presented the same results differently from a medical perspective.

[0017] A computer-implemented method for determining pharmacogenetic star alleles from high-throughput sequencing data is disclosed, the method comprising the following steps: a) providing a plurality of result files output by a plurality of different computer programs for genotyping, each containing multiple result elements, each result element having at least one gene designation (e.g., CYP2C9, CYP2C19, CYP2D6, etc.); b) performing, for each result element of the multiple result elements from each result file of the plurality of result files, the following step: ba) if the result element has an allele designation without double notation, right-shortening the allele designation as far as possible without ascribing a different function to the allele in the shortened description than to the allele in the unabridged description;and c) Perform, for each gene name from at least a set of selected gene names from all gene names contained in the plurality of result files, the following steps: ca) Identify all result elements in all result files that contain the gene name; cb) Identify, from the identified result elements, the diplotypes associated with the gene name; cc) Identify the respective frequencies of identified diplotypes; cd) Identify a result diplotype based on the number of identified result elements and the respective frequencies of different identified diplotypes; and ce) Output the gene name and the identified result diplotype.

[0018] The majority of result files provided are based on a common sample analyzed by a sequencer. The data format from which the result files were derived is irrelevant. The most common data formats are FASTQ, BAM (Binary Alignment Map), and VCF (Variant Call Format). Regardless of the processing steps the result files underwent before their creation, all result files originate from the same sample.

[0019] Result elements in the result files can, for example, have formats such as CYP2D6*1, CYP2D6*2.001, CYP2D6*4.ALDY or CYP2D6*1B, which, for a gene designation, here CYP2D6, each display the allele of one parent individually, or formats such as CYP2D6*2.001 / CYP2D6*2.001 or CYP2D6*1B / *1B, which, for a gene designation, here again CYP2D6, display the alleles of the parents in combination.

[0020] In the examples mentioned, shortening the allele designation as much as possible to the right leads to the following changes:

[0021] CYP2D6*1 => CYP2D6*1

[0022] CYP2D6*2.001 => CYP2D6*2

[0023] CYP2D6*4.ALDY => CYP2D6*4

[0024] CYP2D6*1B => CYP2D6*1

[0025] CYP2D6*2.001 / CYP2D6*2.001 => CYP2D6*2 / CYP2D6*2

[0026] CYP2D6*1B / *1B => CYP2D6*1 / *1.

[0027] Double notation, e.g., "*1xN", is a type of copy number variation (CNV). The result files can contain various types of result elements. However, only those result elements that include at least one gene name are considered. Step c) is executed for each gene name from at least one set of selected gene names. The selection of this set can be automated according to predefined criteria or manual.

[0028] The data obtained in this way are then used to determine the frequency of occurrence of a diplotype for the resulting diplotype in conjunction with the respective gene designation. Since the result elements have now been made comparable, the results of the different computer programs can be compared, thus yielding more reliable results.

[0029] In an advantageous embodiment, the determination of the result diplotype is carried out such that the diplotype identified is the result diplotype whose frequency is greater than or equal to half the number of identified result elements. This embodiment provides a result diplotype with sufficient confidence, even if not all identified diplotypes agree. If higher confidence is desired, the requirement "greater than or equal to" can be replaced by "greater".

[0030] In an advantageous embodiment, the determination of the result diplotype is carried out such that the determined diplotype is designated as an empty marker if two different determined diplotypes have the same highest frequency among all determined diplotypes. This embodiment takes into account that if two diplotypes have the same highest frequency, e.g., each has the same frequency of two, such as 2 x *1 / *1 and 2 x *1 / *2, this should not be considered sufficient confidence. The empty marker is, in particular, a marker labeled "No" or "None".

[0031] In an advantageous embodiment, the determination of the diplotypes assigned to the gene designation involves replacing the determined diplotype with the blank marker if one of the two elements or alleles of the diplotype does not have an allele designation. If only the allele of one parent could be determined, i.e., if only one of the two elements of the diplotype has an allele designation, the determined diplotype is replaced with the blank marker.

[0032] In an advantageous embodiment, the execution for each of the multiple result elements further comprises the following step: marking the result diplotype with an empty marker or deleting the result element if the result element contains information about a copy number variation (e.g., double single notation caused by duplication or the inclusion of the deletion marker caused by deletion) and a copy number variation indication is supported by only one computer program out of a majority of different computer programs. This step is particularly advantageous if the confidence of a computer program in determining the CNV is insufficient. The result of a CNV analysis depends on various factors, such as the sequencing efficiency or the control region, which itself may be affected by a CNV.Using different computer programs with different CNV detection algorithms increases the accuracy of the result.

[0033] In an advantageous embodiment, the execution for each of the multiple result elements further includes the following step: If the result element contains a description of an allele with double notation, the description is replaced by a deletion marker. The underlying principle here is that a CNV can mean either a deletion or a duplication. This can also be understood as meaning that a deletion or duplication essentially determines the number of copies. Normally, two copies should be available, namely those of the two parents. If one or both copies are missing, a deletion occurs. If three or more copies are present, a duplication occurs.

[0034] In an advantageous embodiment, if the allele description contains the deletion marker, the resulting diplotype is marked with an empty marker or the result element is deleted. In this case, it is indicated that no specific diplotype could be determined, even if the allele of a parent could be identified.

[0035] In an advantageous embodiment, the method further includes the step of alphanumerically sorting the elements of the diplotype. This is based on the premise that a diplotype of type M / N and a diplotype of type N / M, e.g., *1 / *4 and *4 / *1, are sufficiently similar that they should not be recorded as different result diplotypes. The alphanumeric sorting means that, referring to the example given, the diplotype *4 / *1 is sorted into *1 / *4, and thus the diplotype *1 / *4 would be determined twice.

[0036] In an advantageous embodiment, the abbreviated description is searched for in a lookup table. If the abbreviated description is found in the lookup table, the procedure continues. If the abbreviated description is not found in the lookup table, the result element corresponding to the abbreviated description is set aside for manual analysis. This embodiment allows for the identification of missing elements in the lookup table. The lookup table is, in particular, a table or database that maps the potentially different names for various variants at different positions of different genes to a standardized allele name. If the abbreviated description is not found in the lookup table, this indicates that...that an entry is missing in the lookup table. Such an entry is therefore excluded as a result element for manual analysis. Furthermore, a computer-implemented method for determining rare genetic variants from high-throughput sequencing data is disclosed, the method comprising the following steps: a) providing at least one source file obtained from the processing of an output file from a sequencer, the source file containing a plurality of information elements, each information element containing information about a genetic variant; b) performing, for each information element of the plurality of information elements, the following steps: ba) determining a test result as to whether a functional prediction should be performed for the genetic variant described in the information element; bb) passing, if the test result is positive,the genetic variant to a function prediction device, and bc) if a function prediction result could be obtained, supplement the information element with the function prediction result, and c) output the information elements that contain a function prediction result.

[0037] As previously explained, the focus in determining pharmacogenetic star alleles is on the 2,603 ​​known pharmacogenomic variants, as these often provide valuable clues for treatment in practice. However, the inventors have recognized that variants that are rare and therefore not yet well understood with regard to their function or effect can also be of interest.

[0038] The technical solution starts with the source file containing multiple information elements, where each information element represents a genetic variant. Specifically, this involves SNPs (Single Nucleotide Polymorphisms) or single base substitutions, each identified within a single information element.

[0039] Due to the large number of such rare variants, a check is performed to determine whether a functional prediction should be carried out for the genetic variant described in the information element. This reduces the number of variants for which a complex functional prediction needs to be performed. If such a functional prediction is to be carried out, the genetic variant is submitted to the relevant service, and the availability of the functional prediction result is awaited. Alphamissense, for example, can be used for this purpose. This is a large collection of functional prediction datasets that are compared to or against the rare variants. If a functional prediction result is obtained, the information element is updated with the result. A prediction is now available for this rare genetic variant, e.g., regarding a functionality such as "damaging" or "benign."

[0040] For the extraction of rare variants, including new variants and variants with unknown function (additional variants), from NGS data, a VCF file or the result file of a variant caller from germline NGS data is used. This VCF file can be generated with any variant caller, such as the GATK HaplotypeCaller. The term "variant caller" refers to a specialized computer program or algorithm used to identify and interpret genetic variants from sequencing data.

[0041] In an advantageous embodiment of the procedure, the test result is negative if the genetic variant described by the information element is included in a negative list, where the negative list specifically includes all known variants. This is an easily implemented approach that can be carried out, for example, using the bioinformatics computer program bcftools. This effectively prevents a known variant, whose functionality is already sufficiently understood, from being subjected to a functional prediction.

[0042] In an advantageous embodiment of the procedure, the test result is positive if the genetic variant described in the information element is included in a positive list containing predefined variants of interest. With this embodiment, the number of rare variants to be subjected to a functional prediction can be actively controlled. It is possible to name the variants specifically or to specify at least their approximate position or region within a gene.

[0043] In an advantageous implementation of the procedure, the step of determining the test result involves annotating the information element with a genome annotation, and the test result is positive if the genome annotation matches one or more search criteria. This implementation utilizes an annotation software program such as ANNOVAR or Ensembl VEP (Variant Effect Predictor), which annotates the variants with additional information. This allows the rare variants to be subjected to function prediction to be filtered according to specific characteristics. For example, the genome annotation can be used to filter for exonic missense variants, splice variants, and / or cis variants.

[0044] In an advantageous embodiment of the method, the genome annotation includes at least one element from the group consisting of genomic location, consequence to a translated amino acid, and allele frequency. It is assumed that this information provides a good basis for keeping the number of rare variants subjected to functional prediction low.

[0045] Finally, a device comprising a processor and a memory is disclosed, wherein the memory contains instructions that, upon execution of the instructions, cause the processor to execute a method according to any of the preceding claims. The device may be a desktop PC, a laptop, a tablet, a server, a mainframe computer, or a supercomputer. The device may also be a specialized processing unit, such as a graphics processing unit (GPU), a digital signal processor (DSP), a field-programmable gate array (FPGA), or an application-specific integrated circuit (ASIC). The memory may be dynamic RAM (DRAM), static RAM (SRAM), or a cache memory.The storage medium can also be an HDD (Hard Disk Drive), an SSD (Solid State Drive), a USB stick, an SD card, a CD, a DVD, a Blu-ray disc, a ROM (Read-Only Memory), a PROM (Programmable ROM), an EPROM (Erasable Programmable ROM) or an EEPROM (Electrically Erasable Programmable ROM).

[0046] embodiment

[0047] The following section will explain a specific example of how, in a computer-implemented method for determining pharmacogenetic star alleles from high-throughput sequencing data, the resulting diplotype, also called the consensus genotype, can be determined based on the number of identified result elements and the respective frequencies of different identified diplotypes. The term "a gene is supported by a certain number of computer programs for genotyping," hereinafter referred to as "tools," means that these tools are at least fundamentally capable of determining the genotype, specifically the diplotype, of a given gene. This does not mean that such a tool will always determine the genotype, and especially not correctly, but rather that it is capable of determining the genotype.

[0048] If the number is one, this means that of the multiple tools, only one is trained to determine the genotype of a specific gene. If the number is two, this means that of the multiple tools, only two are trained to determine the genotype of a specific gene, including the possibility that exactly two tools are used and both tools can determine the genotype of the specific gene. If the number is three, this means that of the multiple tools, only three are trained to determine the genotype of a specific gene, including the possibility that exactly three tools are used and each of the three tools can determine the genotype of the specific gene.If the number is four, this means that of the multiple tools, only four are trained to determine the genotype for a specific gene, although this includes the possibility that exactly four tools are used and each of the four tools can determine the genotype for the specific gene. This logic applies accordingly to a number greater than four.

[0049] This specific example refers to a number of four tools, where for each gene of interest, one, two, three, or four tools can determine the genotype. The trivial case where none of the tools can determine the genotype is disregarded because in this case no diplotype could be generated and no result is available.

[0050] If a gene is supported by exactly one tool, the following initial set of rules applies:

[0051] Result available: For genes supported by only one tool, such as COMT in Adly, the result from that tool is accepted as the result diplotype. No result available: The result diplotype is "None," meaning the tool could not provide a genotype.

[0052] If a gene is supported by exactly two tools, the following procedure is followed according to a second set of rules:

[0053] Same results: The result is accepted as a result diplotype, e.g. *1 / *1 and *1 / *1 result in *1 / *1.

[0054] Different results: A "No" result is assigned to the result diplotype, e.g. *1 / *1 and *1 / *2 result in "No".

[0055] One result and one "No" result: The tool's result with a valid outcome is used as the result diplotype. *1 / *1 and "No" result in *1 / *1.

[0056] If a gene is supported by exactly three tools, the following procedure is followed according to a third set of rules:

[0057] Same results: The result is accepted as a result diplotype, e.g. *1 / *1 , *1 / *1 and *1 / *1 result in *1 / *1.

[0058] The result of one tool is different: The result of the two tools with similar results is accepted as a result diplotype, e.g. *1 / *1 , *1 / *1 and *1 / *2 result in *1 / *1.

[0059] Different results: The "No" result is assigned to the result diplotype, e.g. *1 / *1 , *1 / *2 and *1 / *3 result in "No".

[0060] Two results and one "No" result: The "No" result is removed and the result is generated according to the second set of rules.

[0061] One result and two "No" results: The "No" results are removed and the result is generated according to the first set of rules.

[0062] If a gene is supported by exactly four tools, the following procedure is followed according to a fourth set of rules:

[0063] Same results: The result is accepted as a result diplotype, e.g., *1 / *1, *1 / *T, *1 / *1 and *1 / *1 yield the same result.

[0064] The result of a tool is different: The result of three tools with similar results is accepted as a result diplotype, e.g., *1 / *1, *1 / *1, *1 / *1, and *1 / *2 result in *1 / *1. The results of two tools are different but the same: The "No" result is assigned to the result diplotype, e.g., *1 / *1, *1 / *1, *1 / *2, and *1 / *2 result in "No".

[0065] Different results from two tools: The results of the two tools with similar results are accepted as a result diplotype, e.g. *1 / *1 , *1 / *1 , *1 / *2 and *1 / *3 result in *1 / *1.

[0066] Different results: The "No" result is assigned to the result diplotype, e.g. *1 / *1, *1 / *2 and *1 / *3, *1 / *4 result in "No".

[0067] Three results and one "No" result: The "No" result is removed and the result is generated according to the third set of rules.

[0068] Two results and two "No" results: "No" results are removed and the result is generated according to the second set of rules.

[0069] One result and three "No" results: "No" results are removed and the result is generated according to the first set of rules.

[0070] The embodiment can be adapted to different objectives by adding, modifying, or deleting individual rules. For example, in some preferred embodiments, a "No" result can be output directly if, in addition to exactly one result, only one or more "No" results are found. In other preferred embodiments, a "No" result can be derived even if different results are obtained.

[0071] This logic applies to a number of four. With a number of three, the fourth set of rules can be omitted; with a number of two, the third set of rules can also be omitted. For a number greater than four, further sets of rules can be established with a similar extension of the logic.

Claims

Claims 1. A computer-implemented method for determining pharmacogenetic star alleles from high-throughput sequencing data, comprising the following steps: a) providing a plurality of result files output by a plurality of different computer programs for genotyping, each containing multiple result elements, each result element having at least one gene designation; b) performing, for each result element of the multiple result elements from each result file of the plurality of result files, the following step: ba) if the result element has an allele designation without double notation, right-shortening the allele designation as far as possible without ascribing a different function to the allele in the shortened description than to the allele in the unabridged description;and c) Perform, for each gene name from at least a set of selected gene names from all gene names contained in the plurality of result files, the following steps: ca) Identify all result elements in all result files that contain the gene name; cb) Identify, from the identified result elements, the diplotypes associated with the gene name; cc) Identify the respective frequencies of identified diplotypes; cd) Identify a result diplotype based on the number of identified result elements and the respective frequencies of different identified diplotypes; and ce) Output the gene name and the identified result diplotype.

2. Computer-implemented method according to claim 1, wherein the determination of the result diplotype is carried out such that the determined diplotype is defined as the result diplotype whose frequency is greater than or equal to half the number of determined result elements.

3. Computer-implemented method according to one of the preceding claims, wherein the determination of the result diplotype is carried out such that the determined A diplotype is determined as an empty marker if two different identified diplotypes have the same highest frequency of all identified diplotypes.

4. Computer-implemented method according to one of the preceding claims, wherein the determination of the diplotypes assigned to the gene designation comprises that the determined diplotype is replaced by an empty marker if one of the two elements of the diplotype does not have an allele designation.

5. Computer-implemented method according to one of the preceding claims, wherein the execution, for each result element of the multiple result elements, further comprises the following step: Mark the result diplotype with an empty marker or delete the result element if the result element contains information about a copy number variation and a copy number variation specification is supported by only one computer program out of the majority of different computer programs.

6. Computer-implemented method according to one of the preceding claims, wherein the execution, for each result element of the multiple result elements, further comprises the following step: Replace the description with a deletion marker if the result element contains a description of an allele with double notation.

7. Computer-implemented method according to the preceding claim, wherein, if the description of the allele includes the deletion marker, the result diplotype is marked with an empty marker or the result element is deleted.

8. Computer-implemented method according to one of the preceding claims, further comprising the step of alphanumeric sorting of the elements of the diplotype.

9. Computer-implemented method according to one of the preceding claims, wherein the abbreviated description is searched for in a lookup table and, if the abbreviated description is found in the lookup table, the method is continued and, if the abbreviated description is not found in the lookup table, the result element corresponding to the abbreviated description is set aside for manual analysis.

10. A computer-implemented method for determining rare genetic variants from high-throughput sequencing data, comprising the following steps: a) providing at least one source file obtained from the processing of an output file from a sequencer, wherein the source file contains a plurality of information elements, each information element containing information about a genetic variant; b) performing, for each information element of the plurality of information elements, the following steps: ba) determining a test result to ascertain whether a function prediction should be performed for the genetic variant described in the information element; bb) if the test result is positive, transferring the genetic variant to a function prediction device; and bc) if a function prediction result could be obtained, supplementing the information element with the function prediction result.and c) Outputting the information elements that exhibit a function prediction result.

11. Computer-implemented method according to claim 10, wherein the test result is negative if the genetic variant described by the information element is included in a negative list, wherein the negative list in particular includes all known variants.

12. Computer-implemented method according to claim 10 or 11, wherein the test result is positive if the genetic variant described by the information element is included in a positive list that includes pre-defined variants of interest.

13. Computer-implemented method according to one of claims 10 to 12, wherein the step of determining the test result comprises annotating the information element with a genome annotation and wherein the test result is positive if the genome annotation corresponds to one or more search criteria.

14. Computer-implemented method according to claim 13, wherein the genome annotation comprises at least one element from the group consisting of genomic localization, consequence to a translated amino acid and allele frequency.

15. Device comprising a processor and a memory, wherein the memory comprises instructions which, upon execution of the instructions, cause the processor to execute a method according to any of the preceding claims.

Citation Information

Patent Citations

  • Methods and Systems for Identification of Causal Genomic Variants

    US20140359422A1

  • Systems and Methods for Genomic Annotation and Distributed Variant Interpretation

    US20150154354A1