Methods and related apparatus for screening for potential pathogenic variants and genes
By annotating and screening gene sequencing variant files and utilizing multiple databases to retain valuable variant site information, the problem of a large number of variant sites with little significance in gene sequencing technology has been solved, enabling efficient screening of potential pathogenic variants and clinical auxiliary diagnosis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-17
- Publication Date
- 2026-03-24
AI Technical Summary
Current gene sequencing technologies obtain a large number of variant sites, but these sites have limited biological significance, which makes subsequent scientific research and clinical diagnosis inconvenient and increases the workload of interpreters.
We used multiple annotation databases to annotate and screen gene sequencing variant files, set filtering conditions, retained valuable variant site information such as exon regions and splice region variants, and linked variants with diseases through the HPO database to filter out variant sites with less biological significance.
It improves screening efficiency, reduces the workload of interpreters, provides a better foundation for downstream analysis, and facilitates scientific research and clinical diagnosis.
Smart Images

Figure CN115101123B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of gene sequencing analysis, and in particular to a method for screening potential pathogenic variants and genes and related equipment. BACKGROUND
[0002] Gene sequencing technology refers to a technical means for detecting gene fragments to analyze specific base sequences. After detecting and analyzing human DNA sequences through gene sequencing technology, we can obtain a large number of variant site information contained in the sample. However, the number of obtained variant sites is too large to be directly applied to scientific research analysis or auxiliary clinical diagnosis. Many of the variant site information have little biological significance, which leads to many inconveniences in subsequent downstream analysis, increases the workload of interpreters, and is time-consuming and laborious. SUMMARY
[0003] Therefore, the purpose of the present application is to provide a method for screening potential pathogenic variants and genes and related equipment to solve the above technical problems.
[0004] In a first aspect, the present application provides a method for screening potential pathogenic variants and genes, comprising: obtaining a gene sequencing variant file of a subject, wherein the gene sequencing variant file contains variant site information; selecting a first annotation database to annotate the variant site information in the gene sequencing variant file to obtain a variant annotation file; screening the variant site information in the variant annotation file according to a predetermined variant screening condition to obtain a variant screening file; and selecting a second annotation database to supplement the annotation of the variant site information in the variant screening file to obtain a potential pathogenic variant file.
[0005] Further, the first annotation database includes a RefSeq database, a dbSNP database, a genomicSuperDups database, a gnomAD database, and a dbNSFP database; and the second annotation database includes a HPO database.
[0006] Further, the variant screening condition includes retaining the variant site information annotated by the RefSeq database as an exon region variant or a splice region variant.
[0007] Further, the variant screening condition further includes filtering out the variant site information annotated by the RefSeq database as a synonymous mutation.
[0008] Further, the variant screening condition further includes filtering out the variant site information with a variant site depth lower than a first threshold.
[0009] Further, the variant screening condition further comprises: filtering out the variant site information annotated by the rs ID in the dbSNP database.
[0010] Further, the variant screening condition further comprises: filtering out the variant site information annotated by the genomicSuperDups database.
[0011] Further, the variant screening condition further comprises: filtering out the variant site information annotated by the gnomAD database with a mutation frequency higher than a second threshold.
[0012] Further, the dbNSFP database comprises a plurality of prediction algorithms for predicting whether the variant site information is a deleterious variant, and the variant screening condition further comprises: filtering out the variant site information annotated by the dbNSFP database and predicted as the deleterious variant by only one of the prediction algorithms.
[0013] In a second aspect, the present application provides a device for screening potential pathogenic variants and genes, comprising: an acquisition module configured to acquire a gene sequencing variant file of a subject, the gene sequencing variant file containing variant site information; an annotation module configured to select a first annotation database to annotate the variant site information in the gene sequencing variant file to obtain a variant annotation file; a screening module configured to screen the variant site information in the variant annotation file according to a predetermined variant screening condition to obtain a variant screening file; and a result module configured to select a second annotation database to supplementally annotate the variant site information in the variant screening file to obtain a potential pathogenic variant file.
[0014] In a third aspect, the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method for screening potential pathogenic variants and genes according to the first aspect when executing the computer program.
[0015] In a fourth aspect, the present application provides a non-transitory computer readable storage medium storing computer instructions, wherein the computer instructions are used to make the computer execute the method for screening potential pathogenic variants and genes according to the first aspect.
[0016] From the above, it can be seen that the present application provides a method and related device for screening potential pathogenic variants and genes. By obtaining the gene sequencing variant file of the subject, the variant site information of the subject can be obtained, providing a basis for subsequent variant screening. The first annotation database is selected to annotate the variant site information, which can enrich the variant site information. The variant screening can be set according to the enriched variant site information, further providing a basis for subsequent variant screening, and facilitating subsequent scientific research analysis. According to the variant screening condition, the variant site information in the variant annotation file is screened, a large number of variant sites with less biological significance can be filtered out, the screening speed and efficiency are high, the more valuable variant site information is reserved, the workload of the interpreter is reduced, and time and effort are saved. The second annotation database is selected to supplement the annotation of the variant site information, associate the variant with the disease, and facilitate the interpreter to analyze the potential pathogenic variants and genes. The method and related device for screening potential pathogenic variants and genes have high screening efficiency, can filter out a large number of variant sites with less biological significance, retain more valuable variant site information, provide a good basis for downstream analysis, reduce the workload of the interpreter, and save time and effort. BRIEF DESCRIPTION OF DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present application or related art, the following will briefly introduce the drawings needed to be used in the embodiments or related art descriptions. Obviously, the drawings in the following description are only embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.
[0018] Fig. 1 A flowchart of a method for screening potential pathogenic variants and genes according to an embodiment of the present application;
[0019] Fig. 2 A structural diagram of a device for screening potential pathogenic variants and genes according to an embodiment of the present application;
[0020] Fig. 3 A structural diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0021] In order to make the purpose, technical solutions and advantages of the present application more clear, the following will further describe the present application in detail with specific embodiments and with reference to the drawings.
[0022] For ease of description, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present application shall have the meanings commonly understood by one of ordinary skill in the art to which the present application belongs. The terms "first", "second", and similar terms used in the embodiments of the present application do not denote any order, quantity, or importance, but are only used to distinguish different components. The terms "include", "contain", and similar terms mean that the elements or objects before the terms encompass the elements or objects listed after the terms and their equivalents, and do not exclude other elements or objects. The terms "connect" or "connected" and similar terms are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect.
[0023] Gene sequencing technology refers to a technical means for analyzing specific base sequences by detecting gene fragments. At present, the most widely used method for gene sequencing is the second-generation sequencing technology. Second-generation sequencing, also known as high-throughput sequencing or next-generation sequencing, is based on the principle of complementary base pairing. The core idea is to synthesize and sequence at the same time. The DNA sequence of the sample to be tested is used as a template, and bases are continuously added for synthesis. Finally, the bases are read to obtain the complete sequence information of the sample to be tested.
[0024] Compared with the first-generation sequencing (Sanger sequencing), the second-generation sequencing has the advantages of high throughput, low cost, and short time, and is widely used in basic biological research and many application fields, such as diagnosis, biotechnology, forensic biology, and biological systematics. It has become an essential gene detection technology.
[0025] Through the detection and analysis of human DNA sequences by the second-generation sequencing technology, we can obtain a large number of variant site information contained in the sample. However, the number of variant sites obtained is too large to be directly applied to scientific research analysis or auxiliary clinical diagnosis. Many of the variant site information has little biological significance, which makes the subsequent downstream analysis inconvenient, increases the workload of the interpreters, and is time-consuming and laborious.
[0026] In the process of implementing the present application, it is found that various annotation databases can be used to annotate the files of the second-generation sequencing, and filtering conditions can be set to screen the annotated files, thereby filtering out variant sites with little biological significance and providing a good foundation for downstream analysis.
[0027] In the following, the technical solutions of the present application will be described in detail through specific embodiments and in combination with Figs. 1 to 3 .
[0028] Some embodiments of the present application provide a method for screening potential pathogenic variants and genes, as shown in Fig. 1 , comprising the following steps:
[0029] S1, obtaining a gene sequencing variation file of a subject, the gene sequencing variation file comprising variation site information.
[0030] The gene sequencing variation file is a second-generation sequencing VCF (Variant Call Format) format file. The VCF format is a commonly used format in gene detection, mainly including some header information and detected variation site information. By obtaining the gene sequencing variation file of the subject, the variation site information of the subject can be obtained, thereby providing a basis for subsequent variation screening.
[0031] S2, selecting a first annotation database to annotate the variation site information in the gene sequencing variation file, to obtain a variation annotation file.
[0032] The first annotation database is, for example, a RefSeq (Reference Sequence Database) database, a dbSNP (The Single Nucleotide Polymorphism Database) database, a genomicSuperDups database, a gnomAD (Genome Aggregation Database) database, or a dbNSFP (database for nonsynonymous SNPs' functional predictions) database, without limitation. The annotation of the variation site information can be more comprehensive and accurate.
[0033] The variation site information of the VCF format file only contains basic information of the variation site, and the amount of information is insufficient. By selecting the first annotation database to annotate the variation site information, the variation site information can be enriched, which facilitates subsequent scientific research analysis and further provides a basis for subsequent variation screening. The enriched variation site information includes, for example, exon region variation, splice region variation, or synonymous mutation, without limitation. The variation screening can set screening conditions according to the enriched variation site information.
[0034] S3, screening the variation site information in the variation annotation file according to predetermined variation screening conditions, to obtain a variation screening file.
[0035] The variation screening conditions are, for example, to retain exon region variation and splice region variation, and to filter out variations in the remaining regions, without limitation.
[0036] According to the variation screening condition, the variation site information in the variation annotation file is screened, a large number of variation sites with less biological significance can be filtered out, the screening speed and efficiency are high, more valuable variation site information is reserved, scientific research analysis or auxiliary clinical diagnosis of an interpreter is facilitated, the workload of the interpreter is reduced, and time and labor are saved.
[0037] S4, selecting a second annotation database to supplement annotation of the variation site information in the variation screening file to obtain a potential pathogenic variation file.
[0038] The second annotation database is, for example, an HPO (Human Phenotype Ontology) database, and is not specifically limited. The second annotation database can supplement annotation of the variation site information, associate the variation with a disease, and facilitate an interpreter to analyze potential pathogenic variations and genes.
[0039] The method for screening potential pathogenic variations and genes has high screening efficiency, can filter out a large number of variation sites with less biological significance, retains more valuable variation site information, provides a good foundation for downstream analysis, reduces the workload of an interpreter, and saves time and labor.
[0040] In some embodiments, the VCF format file includes table header information and detected variation site information.
[0041] The table header information includes a version of the VCF format file, code information, and length information of a chromosome, and also includes information interpretation of columns such as a FILTER column, a FORMAT column, and an INFO column in the variation site information.
[0042] The variation site information includes 10 columns of information, wherein the first column is a CHROM column, recording the number of a chromosome where the variation site is located;
[0043] The second column is a POS column, recording a position of the variation site on the chromosome;
[0044] The third column is an ID column, recording an rs ID (Reference SNP ID) of the variation, the rs ID being an ID set by a dbSNP database for an included variation, and if the variation is not included in the dbSNP database, “.” is used to represent;
[0045] The fourth column is a REF column, recording a normal base sequence of the variation site on a reference genome;
[0046] The fifth column is an ALT column, recording a base sequence of the variation site;
[0047] The sixth column is the QUAL column, recording the quality of the base sequence of the detected variation site. The quality is -10log10(p), where p represents the probability of detection error. The larger the quality value, the smaller the probability of detection error.
[0048] The seventh column is the FILTER column, recording whether the variation site is filtered. If it is "PASS", it means that the filtering is performed, otherwise it is marked with ".".
[0049] The eighth column is the INFO column, recording some additional information, which can be information showing the sequencing platform or showing the variation information.
[0050] The ninth column is the FORMAT column, which is an explanation of the tenth column information, mainly including GT (Genotype), DP (Depth) and AD (Allele Depth), etc. DP is the overall depth of the variation site. AD is the depth of each allele of the variation site.
[0051] The tenth column is the specific information, which corresponds to the ninth column one by one. For example, the ninth column is GT, and the tenth column is 0 / 0, 0 / 1, 1 / 1 or 1 / 2. Among them, 0 / 0 represents that the variation site is a homozygous site, and the base is consistent with the base of REF. 0 / 1 represents that the variation site is a heterozygous mutation, and there are two genotypes of REF and ALT. 1 / 1 represents that the variation site is a homozygous mutation, and there are two genotypes of ALT. 1 / 2 represents that the variation site is a heterozygous mutation, and there are two genotypes of ALT1 and ALT2.
[0052] In some embodiments, step S1 further comprises:
[0053] S101, filtering out the variation site information with a variation site depth lower than a first threshold.
[0054] The first threshold is, for example, 10X, and is not specifically limited. The variation site information with a variation site depth lower than the first threshold can be filtered out according to the ninth and tenth columns of the VCF format file. Compared with low site depth, higher site depth has higher detection accuracy, so the variation with low site depth is discarded. Through step S101, the workload of step S2 annotation can be reduced, the screening efficiency of step S3 can be improved, and the efficiency of the method for screening potential pathogenic variations and genes in the present solution can be further improved, saving time and effort, and facilitating scientific research analysis or auxiliary clinical diagnosis of the interpreter.
[0055] In some embodiments, the annotation tools of the first and second annotation databases employ ANNOVAR software, which is a commonly used database annotation software that annotates the annotation information in the database to the corresponding variant site according to the chromosome number of the variant, the start and end positions of the variant, the REF base and the ALT base.
[0056] In some embodiments, the first annotation database includes, for example, a RefSeq database, a dbSNP database, a genomicSuperDups database, a gnomAD database and a dbNSFP database.
[0057] The RefSeq database is a multi-species public database established by the NCBI (National Center for Biotechnology Information), which covers DNA, RNA and protein sequence, gene, expression and function information.
[0058] The dbSNP database is a small nucleotide variation database established by the NCBI and the NHGRI (National Human Genome Research Institute) in cooperation, which includes human single nucleotide variation, microsatellite, short fragment insertion and deletion, population frequency, genomic information and related information of the variation in the RefSeq database.
[0059] The genomicSuperDups database is a public database for annotating repetitive sequence regions on the genome recorded at the UCSC (University of California, Santa Cruz).
[0060] The gnomAD database is a public database established by an international researcher alliance, which collects exome and genomic sequencing data of various large-scale sequencing projects.
[0061] dbNSFP is a public database for functional prediction and annotation of all potential non-synonymous single nucleotide variations (nsSNVs) in human genome, which integrates multiple prediction algorithms to predict whether the variation is a deleterious variation, including REVEL algorithm, Polyphen2-HDIV algorithm, SIFT (Scale-invariant feature transform) algorithm, Polyphen2-HVAR algorithm, MutationTaster2 algorithm, MutationAssessor algorithm, PROVEAN algorithm and CADD (Computer Aided Drug Design) algorithm, etc., and also includes information of multiple conservation scoring software and multiple public databases, including GERP++ and SiPhy, etc., and 1000 Genomes database, ExAC (the Exome Aggregation Consortium) database and ESP6500 (NHLBI GO Exome Sequencing Project) database, etc.
[0062] In some embodiments, the second annotation database comprises an HPO database.
[0063] The HPO database is a product of the Monarch Initiative, aiming to provide a standardized vocabulary of phenotypic abnormalities encountered in human diseases, and currently includes more than 13,000 terms and more than 156,000 descriptions and annotations of genetic diseases.
[0064] The variation screening file obtained after filtering in step S3 is annotated with protein information, phenotype name and definition and disease information corresponding to the phenotype in the HPO database according to the gene name of the detected variation, so as to obtain a final potential pathogenic variation file annotated with potential pathogenic variations and genes, facilitating analysts to analyze and obtain more scientific and accurate analysis results.
[0065] In some embodiments, the variation screening condition comprises:
[0066] S301, retaining the variation site information annotated with exon region variation or splicing region variation by the RefSeq database.
[0067] Exon region variation refers to a variation occurring in an exon region. An exon is a part of a gene, which is retained on mRNA after RNA splicing and is finally expressed as a protein in the process of protein biosynthesis. Since an exon is involved in the process of protein biosynthesis, if a variation occurs in the position, it may affect the synthesis of the protein, and therefore the exon region variation needs to be retained.
[0068] Splice region variation refers to a variation occurring in a region 1-3 bp away from an exon region or 3-8 bp away from an intron region. RNA splicing refers to a process of removing an intron from a DNA template chain transcription product and linking exons into a continuous RNA sequence. Since mRNA can be formed only after RNA splicing, if a variation occurs in the splice region, it may affect RNA splicing and further affect the formation of mRNA, and therefore the splice region variation needs to be retained.
[0069] Step S301 filters out other region variations except exon region variations and splice region variations, so that most of the variation site information can be filtered out, and the subsequent filtering and screening speed can be improved.
[0070] S302, filter out the variation site information annotated with synonymous mutations by the RefSeq database.
[0071] A synonymous mutation refers to a mutation of a base pair in a DNA sequence, which does not affect the encoded amino acid. For example, after GCC is mutated to GCA, the encoded amino acid is still alanine. Therefore, if the variation is a synonymous mutation, it will not affect the synthesis of the protein, and therefore the variation is discarded.
[0072] S303, filter out the variation site information with a variation site depth lower than a first threshold value.
[0073] The first threshold value is, for example, 10X, and is not specifically limited. The variation site information with a variation site depth lower than the first threshold value can be filtered out according to the ninth column and the tenth column information in the VCF format file. Compared with a low site depth, a higher site depth has a higher detection accuracy, and therefore the variation with a lower site depth is discarded.
[0074] S304, filter out the variation site information annotated with rs ID by the dbSNP database.
[0075] The variation annotated with rs ID is a variation already included in the dbSNP database, and the sequence, position, and distribution frequency of the variation and other related information are recorded. The variation does not belong to a rare variation, and has a lower biological significance, and therefore the variation is discarded.
[0076] S305, filter out the variation site information annotated by the genomicSuperDups database.
[0077] The genomicSuperDups database can detect whether the variant site is located in a repeated fragment. The genetic variation detected in the repeated fragment is mostly caused by sequence alignment errors, so there is a high probability of false positive variation, and therefore the variation is discarded.
[0078] S306, filtering out the variant site information annotated by the gnomAD database with a mutation frequency higher than a second threshold.
[0079] The second threshold is, for example, 0.5% or 1%, and is not particularly limited. Generally, the variation with a high mutation frequency in the population often has no pathogenicity, and therefore this type of variation can be filtered out, and the variation with a low mutation frequency in the population is retained.
[0080] S307, filtering out the variant site information annotated by the dbNSFP database and predicted as a deleterious variation by only one prediction algorithm.
[0081] The dbNSFP database includes multiple prediction algorithms for predicting whether the variant site information is a deleterious variation, including SIFT algorithm, REVEL algorithm, Mutation Taster algorithm and PROVEAN algorithm.
[0082] The SIFT algorithm is based on the amino acid conservation of each site of the homologous protein, and predicts whether the variation affects the protein function through the evolutionary conservation and position-specific scoring matrix, and then performs deleterious variation prediction.
[0083] The REVEL algorithm is an integrated method for predicting missense mutations, which combines the prediction results of multiple tools, including SIFT algorithm, PROVEAN algorithm, FATHMM algorithm, VEST algorithm, PolyPhen algorithm, Mutation Assessor algorithm and Mutation Taster algorithm, uses random forest algorithm, uses pathogenic and rare nonsense mutations for training, and then performs deleterious variation prediction.
[0084] The MutationTaster algorithm is based on Grantham matrix score, evolutionary conservation and Bayesian classifier to perform deleterious variation prediction.
[0085] The PROVEAN algorithm uses BLOSUM62 model, neural network model and evolutionary conservation to predict whether the variation of the protein will affect the function of the protein, thereby performing deleterious variation prediction.
[0086] Step S307 filters out the variant site information predicted as a deleterious mutation by only one of the SIFT algorithm, the REVEL algorithm, the Mutation Taster algorithm, and the PROVEAN algorithm, and retains the variant site information predicted as a deleterious mutation by two or more prediction algorithms, which has higher prediction accuracy and can improve the value of the retained variant site information.
[0087] Through the variant screening conditions of steps S301-S307, a large number of variant sites with less biological significance can be filtered out, the screening speed and efficiency are high, the overall screening time is less than 5 min, more valuable variant site information is retained, scientific research analysis or auxiliary clinical diagnosis is facilitated for an interpreter, the workload of the interpreter is reduced, and time and effort are saved.
[0088] It should be noted that some embodiments of the application have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than the order described above and still achieve desirable results. In addition, the processes depicted in the figures do not necessarily require the particular order shown or sequential order to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous or possible.
[0089] Based on the same inventive concept, the application also provides a device for screening potential pathogenic variants and genes corresponding to any of the above-mentioned embodiments, which refers to Fig. 2 , comprising:
[0090] The acquisition module 21 is configured to acquire a gene sequencing variant file of a subject, wherein the gene sequencing variant file contains variant site information.
[0091] The annotation module 22 is configured to select a first annotation database to annotate the variant site information in the gene sequencing variant file to obtain a variant annotation file.
[0092] The screening module 23 is configured to screen the variant site information in the variant annotation file according to a predetermined variant screening condition to obtain a variant screening file.
[0093] The result module 24 is configured to select a second annotation database to supplement the annotation of the variant site information in the variant screening file to obtain a potential pathogenic variant file.
[0094] For the convenience of description, the above device is described in various modules according to functions. Of course, the functions of each module can be implemented in the same or multiple software and / or hardware when implementing the application.
[0095] The device of the above embodiments is used to implement the method of screening potential pathogenic variants and genes in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which are not repeated here.
[0096] Based on the same inventive concept, the present application also provides an electronic device corresponding to the method of any of the above embodiments, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the method of screening potential pathogenic variants and genes according to any of the above embodiments when executing the program.
[0097] Fig. 3 A more specific hardware structure schematic diagram of an electronic device provided by the present embodiment is shown, which can include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040 and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030 and the communication interface 1040 are connected to each other through the bus 1050 for communication within the device.
[0098] The processor 1010 can be implemented by a general CPU (Central Processing Unit), a microprocessor, an ASIC (Application Specific Integrated Circuit) or one or more integrated circuits, etc., for executing related programs to implement the technical solutions provided by the present embodiment.
[0099] The memory 1020 can be implemented by a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 can store an operating system and other application programs, and when the technical solutions provided by the present embodiment are implemented by software or firmware, the related program codes are stored in the memory 1020 and executed by the processor 1010.
[0100] The input / output interface 1030 is used to connect input / output modules to realize information input and output. The input / output modules can be configured as components in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. The input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, an indicator light, etc.
[0101] The communication interface 1040 is configured to connect a communication module (not shown in the figure) to realize the communication interaction between the device and other devices. The communication module can realize communication through a wired manner (for example, a USB, a network cable, etc.) or a wireless manner (for example, a mobile network, WIFI, Bluetooth, etc.).
[0102] The bus 1050 includes a path for transmitting information between various components (for example, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040) of the device.
[0103] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, the device can also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device can also only contain the components necessary for the implementation of the embodiments of the present application, and does not necessarily contain all the components shown in the figure.
[0104] The electronic device of the above embodiment is used to realize the method of screening potential pathogenic variants and genes in any of the preceding embodiments, and has the beneficial effects of the corresponding method embodiments, which are not described here.
[0105] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present application also provides a non-transitory computer readable storage medium storing computer instructions for causing the computer to execute the method of screening potential pathogenic variants and genes according to any of the above embodiments.
[0106] The computer readable medium of the present embodiment includes permanent and non-permanent, removable and non-removable media, which can be realized by any method or technology to store information. The information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0107] The computer instructions stored on the storage medium of the above embodiments are used to make the computer execute the method for screening potential pathogenic variants and genes as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which are not described here again.
[0108] It should be understood by those of ordinary skill in the art that the discussion of any of the above embodiments is merely exemplary and is not intended to suggest the scope of the present application (including the claims) is limited to these examples; the above embodiments or technical features among different embodiments can also be combined, the steps can be implemented in any order, and there are many other changes of the different aspects of the embodiments of the present application as described above, which are not provided in details for the sake of brevity. It should be understood that the above description is merely illustrative and is not intended to suggest the scope of the present application (including the claims) is limited to these examples.
[0109] In addition, in order to simplify the description and discussion, and so as not to make the embodiments of the present application difficult to understand, the apparatuses can be shown in the form of block diagrams, so as to avoid making the embodiments of the present application difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram apparatuses are highly dependent on the platform to be implemented by the embodiments of the present application (i.e. these details should be fully within the understanding of those skilled in the art). In the case of setting forth specific details to describe the exemplary embodiments of the present application, it is obvious to those skilled in the art that the embodiments of the present application can be implemented without these specific details or with changes to these specific details. Therefore, these descriptions should be considered as illustrative rather than limiting.
[0110] Although the present application has been described in conjunction with the specific embodiments thereof, it is evident that many alternatives, modifications and variations will be apparent to those skilled in the art in light of the foregoing description.
[0111] The embodiments of the present application are intended to cover all such alternatives, modifications and variations as falling within the broad scope of the appended claims. Accordingly, any omission, modification, equivalent replacement, improvement, etc. made within the spirit and principle of the embodiments of the present application should be included in the protection scope of the present application.
Claims
1. A method for screening potential pathogenic variants and genes, characterized in that, include: Obtain the gene sequencing variant file of the subject, the gene sequencing variant file contains variant site information, and filter out the variant site information whose variant site depth is lower than a first threshold. The mutation site information in the gene sequencing mutation file is annotated using a first annotation database to obtain a mutation annotation file. The first annotation database includes the RefSeq database, dbSNP database, genomicSuperDups database, gnomAD database, and dbNSFP database. The dbNSFP database includes various prediction algorithms to predict whether the mutation site information is a harmful mutation. The variant site information in the variant annotation file is filtered according to predetermined variant screening conditions to obtain a variant screening file. The variant screening conditions include: retaining variant site information annotated with exon region variants or splice region variants in the RefSeq database; filtering out variant site information annotated with synonymous mutations in the RefSeq database; filtering out variant site information annotated with rs IDs in the dbSNP database; filtering out variant site information annotated with the genomicSuperDups database; filtering out variant site information annotated with mutations in the gnomAD database whose mutation frequency is higher than a second threshold; and filtering out variant site information annotated with the dbNSFP database that is predicted as a harmful variant by only one of the prediction algorithms. A second annotation database is selected to supplement the mutation site information in the mutation screening file to obtain a potential pathogenic mutation file, wherein the second annotation database includes the HPO database.
2. A device for screening potential pathogenic variants and genes, characterized in that, include: The acquisition module is configured to acquire the gene sequencing variant file of the subject, the gene sequencing variant file containing variant site information, and filter out the variant site information whose variant site depth is lower than a first threshold. The annotation module is configured to select a first annotation database to annotate the variant site information in the gene sequencing variant file to obtain a variant annotation file. The first annotation database includes the RefSeq database, dbSNP database, genomicSuperDups database, gnomAD database and dbNSFP database. The dbNSFP database includes a variety of prediction algorithms to predict whether the variant site information is a harmful variant. A filtering module is configured to filter the variant site information in the variant annotation file according to predetermined variant screening conditions to obtain a variant screening file. The variant screening conditions include: retaining variant site information annotated with exon region variants or splice region variants in the RefSeq database; filtering out variant site information annotated with synonymous mutations in the RefSeq database; filtering out variant site information annotated with rs IDs in the dbSNP database; filtering out variant site information annotated with the genomicSuperDups database; filtering out variant site information annotated with mutations in the gnomAD database whose mutation frequency is higher than a second threshold; and filtering out variant site information annotated with the dbNSFP database that is predicted as a harmful variant by only one of the prediction algorithms. The results module is configured to select a second annotation database to supplement the variant site information in the variant screening file to obtain a potential pathogenic variant file, wherein the second annotation database includes the HPO database.
Citation Information
Patent Citations
Method and device for automatically generating gene testing report, and electronic equipment
CN109754856A
Method and system for screening pathogenic variation
CN111863132A