Clinical annotation method and device for pathogenicity of copy number variation
The method automates CNV annotation by scoring against multiple databases, addressing complexity and performance issues in existing software, enhancing user-friendliness and accuracy.
Patent Information
- Application Number
- CN202111566490.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-20
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-12-20
AI Technical Summary
The existing CNV intelligent interpretation software has complex operation and poor detection performance, and cannot automatically implement the annotation of the pathogenicity of copy number variation, making it difficult to meet the needs of clinicians.
A clinical annotation method and device based on the copy number variation interpretation guide is provided. By obtaining the location information and variation types of copy number variation, it uses databases such as ensembl, ClinGen, dbVar, DECIPHER, DGV, ISCA, OMIM to compare and score, realizes automatic annotation of copy number variation, combines user experience to comprehensively score, and finally performs clinical significance grading.
It realizes the convenient operation and efficient automation of CNV intelligent interpretation software, reduces the workload of manual interpretation, improves the efficiency of genetic variation analysis and clinical interpretation, and ensures the reliability and accuracy of annotation results.
Smart Images

Figure CN114496300B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of gene variation, and particularly to a clinical annotation method and device for the pathogenicity of copy number variation. Background Art
[0002] Copy number variation (CNV) is caused by genomic rearrangement, generally referring to the increase or decrease in the copy number of large genomic fragments with a length of more than 1 kb. CNV is an important part of genomic structural variation (SV), mainly manifested as submicroscopic deletions and duplications. CNV is an important part of human genetic variation and plays a role in human diseases. Accurate identification and clinical annotation are crucial when evaluating patients with neurodevelopmental disorders and congenital abnormalities.
[0003] The American College of Medical Genetics and Genomics (ACMG) promulgated new CNV interpretation guidelines in 2019. The new guidelines established a scoring system to facilitate users to evaluate pathogenicity evidence in a more quantitative way and also provided convenient evidence for the automation of CNV interpretation.
[0004] However, the existing CNV intelligent interpretation software developed based on the new guidelines is complex to operate, has poor detection performance, and does not propose solutions for evidence that cannot be automatically annotated for clinicians without bioinformatics background. Summary of the Invention
[0005] The present invention provides a clinical annotation method and device for the pathogenicity of copy number variation, which can solve the defects that the existing CNV intelligent interpretation software is complex to operate, has poor detection performance, and does not propose solutions for evidence that cannot be automatically annotated, and can ensure the convenient operation of the CNV intelligent interpretation software and improve the accuracy of the CNV intelligent interpretation software.
[0006] In a first aspect, the present invention provides a clinical annotation method for the pathogenicity of copy number variations, including: obtaining the location information and the type of variation of the copy number variation; the location information includes the chromosome where the copy number variation is located, the starting position of the copy number variation, and the ending position of the copy number variation; the type of variation includes copy number deletion and copy number duplication; according to the location information and the type of variation, determining first information of the copy number variation, the first information including whether it contains protein-coding genes, the number of protein-coding genes contained, the overlapping information between the copy number variation and common population variations, the overlapping information between the copy number variation and historically common pathogenic copy number variations, the overlapping information between the copy number variation and historically common benign copy number variations, and the overlapping information between the copy number variation and haploinsufficient genes or regions, or the overlapping information between the copy number variation and triplo-sensitive genes or regions; based on the copy number variation interpretation guidelines, comparing the first information of the copy number variation with a preset database according to the type of variation, scoring the first information, and obtaining a first evaluation result; the preset database is established based on the ensembl, ClinGen, dbVar, DECIPHER, DGV, ISCA, and OMIM databases; clinically significantly grading the pathogenicity of the copy number variation according to the first evaluation result to obtain the result of the clinical significance grading, and the result of the clinical significance grading includes benign, likely benign, of uncertain clinical significance, likely pathogenic, and pathogenic.
[0007] According to the clinical annotation method for the pathogenicity of copy number variations provided by the present invention, the comparing the first information of the copy number variation with a preset database according to the type of variation and scoring the first information includes: when the type of variation is copy number deletion, scoring whether the copy number deletion region contains protein-coding genes, the number of protein-coding genes contained, the overlapping information between the copy number variation and haploinsufficient genes or regions, the overlapping information between the copy number variation and common population variations, the overlapping information between the copy number variation and historically common pathogenic copy number variations, and the overlapping information between the copy number variation and historically common benign copy number variations.
[0008] According to the clinical annotation method for the pathogenicity of copy number variations provided by the present invention, comparing the first information of the copy number variation with a preset database according to the variation type and scoring the first information includes: when the variation type is copy number duplication, scoring whether the copy number duplication region contains protein-coding genes, the number of the contained protein-coding genes, the overlapping information between the copy number variation and the genes or regions sensitive to triple dosage, the overlapping information between the copy number variation and the common population variations, the overlapping information between the copy number variation and the historically common pathogenic copy number variations, and the overlapping information between the copy number variation and the historically common benign copy number variations.
[0009] The clinical annotation method for the pathogenicity of copy number variations provided by the present invention further includes: receiving the second information of the copy number variation; the second information includes the characteristics of the reported phenotypes of the genes or regions where the copy number variation occurs, the co-segregation characteristics of the pedigree, the statistical difference situation of case-control studies, the disease inheritance pattern, and the parental origin; receiving the score given by the user based on experience for the second information to obtain a second scoring result; and clinically significantly grading the pathogenicity of the copy number variation according to the first scoring result and the second scoring result to obtain the result of the clinical significance grading.
[0010] The clinical annotation method for the pathogenicity of copy number variations provided by the present invention further includes: displaying the first scoring result, the second scoring result, and the result of the clinical significance grading.
[0011] The clinical annotation method for the pathogenicity of copy number variations provided by the present invention further includes: receiving the viewing permissions set by the user for the first information, the second information, the first scoring result, the second scoring result, and the result of the clinical significance grading of the copy number variation; and receiving the instruction to print or download the first information, the second information, the first scoring result, the second scoring result, and the result of the clinical significance grading with the viewing permissions.
[0012] Second aspect, the present invention also provides a clinical annotation device for the pathogenicity of copy number variations, comprising: an acquisition module, configured to acquire the location information and the variation type of the copy number variation; the location information includes the chromosome where the copy number variation is located, the start position of the copy number variation, and the end position of the copy number variation; the variation type includes copy number deletion and copy number duplication; a first determination module, configured to determine first information of the copy number variation according to the location information and the variation type, the first information including whether it contains protein-coding genes, the number of protein-coding genes contained, the overlap information between the copy number variation and common population variations, the overlap information between the copy number variation and historically common pathogenic copy number variations, the overlap information between the copy number variation and historically common benign copy number variations, and the overlap information between the copy number variation and haploinsufficient genes or regions, or the overlap information between the copy number variation and triplo-sensitive genes or regions; a first scoring module, configured to compare the first information of the copy number variation with a preset database based on a copy number variation interpretation guide according to the variation type, score the first information, and obtain a first evaluation result; the preset database is established based on the ensembl, ClinGen, dbVar, DECIPHER, DGV, ISCA, and OMIM databases; a first classification module, configured to clinically classify the pathogenicity of the copy number variation according to the first evaluation result to obtain the result of the clinical significance classification, and the result of the clinical significance classification includes benign, likely benign, uncertain significance, likely pathogenic, and pathogenic.
[0013] Third aspect, the present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the steps of the clinical annotation method for the pathogenicity of copy number variations as described in the first aspect are implemented.
[0014] Fourth aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the clinical annotation method for the pathogenicity of copy number variations as described in the first aspect are implemented.
[0015] Fifth aspect, the present invention also provides a computer program product, on which executable instructions are stored, and when the instructions are executed by a processor, the processor is caused to implement the steps of the clinical annotation of the pathogenicity of copy number variations as described in the first aspect.
[0016] A clinical annotation method and device for the pathogenicity of copy number variations provided by the present invention obtain the position information and variation type of the copy number variation; the position information includes the chromosome where the copy number variation is located, the starting position of the copy number variation, and the ending position of the copy number variation; the variation type includes copy number deletion and copy number duplication; according to the position information and the variation type, the first information of the copy number variation is determined, and the first information includes whether it contains protein-coding genes, the number of protein-coding genes contained, the overlapping information between the copy number variation and common population variations, the overlapping information between the copy number variation and historically common pathogenic copy number variations, the overlapping information between the copy number variation and historically common benign copy number variations, and the overlapping information between the copy number variation and haploinsufficient genes or regions, or the overlapping information between the copy number variation and triplo-sensitive genes or regions; based on the copy number variation interpretation guidelines, the first information of the copy number variation is compared with a preset database according to the variation type, and the first information is scored to obtain a first evaluation result; the preset database is established based on the ensembl, ClinGen, dbVar, DECIPHER, DGV, ISCA, and OMIM databases; according to the first evaluation result, the pathogenicity of the copy number variation is clinically significantly graded to obtain the result of clinical significance grading, and the result of clinical significance grading includes benign, likely benign, of uncertain clinical significance, likely pathogenic, and pathogenic. A scoring system is established according to the copy number variation interpretation guidelines, and each piece of evidence of the copy number variation is scored item by item. Finally, the pathogenicity of the copy number variation is annotated according to the total score, realizing one-key automated annotation analysis of copy number variations, greatly reducing the workload of manual interpretation, and greatly improving the efficiency of genetic variation analysis and clinical interpretation of copy number variations; in the scoring process, databases such as ensembl, ClinGen, dbVar, DECIPHER, DGV, ISCA, and OMIM are comprehensively considered, making the annotation results more reliable. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the implementation examples or the prior art descriptions. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 It is a schematic flowchart of an embodiment of a clinical annotation method for the pathogenicity of copy number variations provided by the present invention;
[0019] Figure 2 It is a schematic diagram of an annotation result provided by the present invention;
[0020] Figure 3 Schematic diagram of the detection results of UAnno1.0 provided by the present invention;
[0021] Figure 4 Schematic diagram of the detection results of ClassifyCNV provided by the present invention;
[0022] Figure 5 Another schematic diagram of the detection results of UAnno1.0 provided by the present invention;
[0023] Figure 6 Another schematic diagram of the detection results of ClassifyCNV provided by the present invention;
[0024] Figure 7 Schematic diagram of the UAnno1.0 interface provided by the present invention;
[0025] Figure 8 Schematic diagram of the display page of the historical database provided by the present invention;
[0026] Figure 9 Schematic diagram of the interpretation page provided by the present invention;
[0027] Figure 10 Schematic diagram of the manual evidence adjustment page provided by the present invention;
[0028] Figure 11 Schematic diagram of the interpretation result page provided by the present invention;
[0029] Figure 12 Schematic diagram of the gene annotation page provided by the present invention;
[0030] Figure 13 Schematic diagram of the dosage sensitivity annotation page provided by the present invention;
[0031] Figure 14 Schematic diagram of the gene link page provided by the present invention;
[0032] Figure 15 Schematic diagram of the common population CNV annotation page provided by the present invention;
[0033] Figure 16 Schematic diagram of the historical CNV annotation page provided by the present invention;
[0034] Figure 17 Schematic diagram of the genome browser page provided by the present invention;
[0035] Figure 18 Another schematic diagram of the UAnno1.0 interface provided by the present invention;
[0036] Figure 19Schematic diagram of the CNV summary result page provided by the present invention;
[0037] Figure 20 Schematic diagram of the composition structure of an embodiment of a clinical annotation device for the pathogenicity of copy number variations provided by the present invention;
[0038] Figure 21 Schematic diagram of the physical structure of an electronic device provided by the present invention. Detailed implementation manners
[0039] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0040] In the present invention, the clinical annotation method for the pathogenicity of copy number variations is applied to the clinical annotation tool for the pathogenicity of copy number variations - UAnno1.0. This software is divided into two parts: the front end and the back end. The back end uses the R language to write the rules for interpreting evidence, and the front end uses the shiny interface for input and display, so that the interpreters can use this software more conveniently.
[0041] Figure 1 Schematic diagram of the process of an embodiment of a clinical annotation method for the pathogenicity of copy number variations provided by the present invention. As Figure 1 shown, the clinical annotation method for the pathogenicity of copy number variations includes the following steps:
[0042] S101, obtaining the position information and the variation type of the copy number variation; the position information includes the chromosome where the copy number variation is located, the starting position of the copy number variation, and the ending position of the copy number variation; the variation types include copy number deletion and copy number duplication.
[0043] In step S101, the user inputs the position information and the variation type of each copy number variation in the front-end interface of UAnno1.0, and the software can obtain the position information and the variation type of the copy number variation.
[0044] S102. Determine the first information of the copy number variation according to the location information and the type of variation. The first information includes whether it contains protein-coding genes, the number of protein-coding genes it contains, the overlapping information between the copy number variation and the variations in the general population, the overlapping information between the copy number variation and the historically common pathogenic copy number variations, the overlapping information between the copy number variation and the historically common benign copy number variations, the overlapping information between the copy number variation and the genes or regions with haploinsufficiency, or the overlapping information between the copy number variation and the genes or regions sensitive to triploid dosage.
[0045] In step S102, according to the location information of the copy number variation, that is, the chromosome of the copy number variation, the starting position of the copy number variation, and the ending position of the copy number variation, the gene fragment where the copy number variation occurs can be accurately located, as well as whether the fragment contains protein-coding genes and the number of protein-coding genes it contains. According to the type of variation, the overlapping information between the copy number variation on this fragment and the variations in the general population, the overlapping information between the copy number variation and the historically common pathogenic copy number variations, the overlapping information between the copy number variation and the historically common benign copy number variations, the overlapping information between the copy number variation and the genes or regions with haploinsufficiency, or the overlapping information between the copy number variation and the genes or regions sensitive to triploid dosage can be obtained.
[0046] S103. Based on the copy number variation interpretation guidelines, compare the first information of the copy number variation with a preset database according to the type of variation, score the first information, and obtain the first evaluation result. The preset database is established based on the ensembl, ClinGen, dbVar, DECIPHER, DGV, ISCA, and OMIM databases.
[0047] In step S103, the interpretation guidelines were promulgated by the American College of Medical Genetics and Genomics (ACMG) in 2019. The interpretation guidelines established a scoring system to facilitate users to evaluate the pathogenicity evidence in a more quantitative method and also provided convenient evidence for the automation of CNV interpretation.
[0048] Due to the differences in the types of variations, the corresponding first information of the copy number variations also varies. Compare the first information of the copy number variation with the preset database, and score the first information according to the interpretation guidelines. Sum up the scores of each piece of information in the first information to obtain the first evaluation result.
[0049] S104. Clinically significant grading of the pathogenicity of the copy number variation is performed according to the first evaluation result to obtain the result of clinically significant grading. The result of clinically significant grading includes benign, likely benign, uncertain significance, likely pathogenic, and pathogenic.
[0050] In step S104, the basis for clinically grading the pathogenicity of copy number variations according to the first evaluation result can be as follows: when the first evaluation result is less than or equal to -0.99 points, it is determined that the copy number variation is a benign variation; when the first evaluation result is between -0.9 and -0.98, it is determined that the copy number variation is a likely benign variation; when the first evaluation result is between -0.9 and 0.9, it is determined that the copy number variation is a variant of uncertain clinical significance; when the first evaluation result is between 0.9 and 0.98, it is determined that the copy number variation is a likely pathogenic variation; when the first evaluation result is greater than or equal to 0.99, it is determined that the copy number variation is a pathogenic variation.
[0051] A method for clinically annotating the pathogenicity of copy number variations provided by the present invention includes obtaining the position information and variation type of the copy number variation; the position information includes the chromosome where the copy number variation is located, the starting position of the copy number variation, and the ending position of the copy number variation; the variation type includes copy number deletion and copy number duplication; according to the position information and variation type, the first information of the copy number variation is determined, and the first information includes whether it contains protein-coding genes, the number of protein-coding genes, the overlapping information between the copy number variation and common population variations, the overlapping information between the copy number variation and historically common pathogenic copy number variations, the overlapping information between the copy number variation and historically common benign copy number variations, and the overlapping information between the copy number variation and haploinsufficient genes or regions, or the overlapping information between the copy number variation and triplo-sensitive genes or regions; based on the copy number variation interpretation guidelines, the first information of the copy number variation is compared with a preset database according to the variation type, and the first information is scored to obtain the first evaluation result; the preset database is established based on the ensembl, ClinGen, dbVar, DECIPHER, DGV, ISCA, and OMIM databases; the pathogenicity of the copy number variation is clinically graded according to the first evaluation result to obtain the result of clinical significance grading, and the result of clinical significance grading includes benign, likely benign, uncertain clinical significance, likely pathogenic, and pathogenic. A scoring system is established according to the copy number variation interpretation guidelines, and each item of evidence for the copy number variation is scored item by item. Finally, the pathogenicity of the copy number variation is annotated based on the total score, realizing one-key automated annotation analysis of copy number variations, greatly reducing the workload of manual interpretation, and significantly improving the efficiency of genetic variation analysis and clinical interpretation of copy number variations; during the scoring process, databases such as ensembl, ClinGen, dbVar, DECIPHER, DGV, ISCA, and OMIM are comprehensively considered, making the annotation results more reliable.
[0052] In some alternative embodiments, comparing the first information of the copy number variation with a preset database according to the type of variation and scoring the first information may include: when the type of variation is copy number deletion, scoring whether the copy number deletion region contains protein-coding genes, the number of protein-coding genes contained, the overlapping information between the copy number variation and the genes or regions with haploinsufficiency, the overlapping information between the copy number variation and the variations in the general population, the overlapping information between the copy number variation and the historically common pathogenic copy number variations, and the overlapping information between the copy number variation and the historically common benign copy number variations. The scoring criteria are shown in Table 1, where the HI gene is a gene with haploinsufficiency.
[0053] Table 1 Scoring Criteria for Copy Number Deletion with Different Types
[0054]
[0055]
[0056] In some alternative embodiments, comparing the first information of the copy number variation with a preset database according to the type of variation and scoring the first information includes: when the type of variation is copy number duplication, scoring whether the copy number duplication region contains protein-coding genes, the number of protein-coding genes contained, the overlapping information between the copy number variation and the genes or regions sensitive to triplo dosage, the overlapping information between the copy number variation and the variations in the general population, the overlapping information between the copy number variation and the historically common pathogenic copy number variations, and the overlapping information between the copy number variation and the historically common benign copy number variations. The scoring criteria are shown in Table 2, where the TS gene is a gene with haploinsufficiency.
[0057] Table 1 Scoring Criteria for Copy Number Deletion with Different Types
[0058]
[0059]
[0060] In some alternative embodiments, the clinical annotation method for the pathogenicity of copy number variations further includes: receiving the second information of the copy number variation; the second information includes the characteristics of the reported phenotypes of the genes or regions where the copy number variation occurs, the co-segregation characteristics of the pedigree, the statistical differences in case-control studies, the mode of inheritance of the disease, and the parental origin; receiving the score of the second information based on the user's experience to obtain the second scoring result; and grading the clinical significance of the pathogenicity of the copy number variation according to the first scoring result and the second scoring result to obtain the result of the clinical significance grading.
[0061] Among them, the disease inheritance patterns include dominant inheritance and recessive inheritance, and the parental origin includes de novo and inherited from parents. The second information is the evidence in the interpretation guide. However, since there is no automated interpretation evidence database, the interpreter needs to add it based on their own experience and score the added second information to obtain the second scoring result. The first scoring result and the second scoring result are summed up, and the pathogenicity level of the copy number variation is judged according to the sum of the two.
[0062] In some alternative embodiments, the method for clinically annotating the pathogenicity of copy number variations further includes: presenting the first scoring result, the second scoring result, and the result of clinical significance classification. Among them, the above results can be presented through the front-end shiny interface.
[0063] In some alternative embodiments, the first information, the second information, the first scoring result, the second scoring result, and the result of clinical significance classification of the copy number variation queried by the user are saved to the background database according to the user's logged-in account.
[0064] In some alternative embodiments, the method for clinically annotating the pathogenicity of copy number variations further includes: receiving the viewing permissions set by the user for the first information, the second information, the first scoring result, the second scoring result, and the result of clinical significance classification of the copy number variation; receiving an instruction to print or download the first information, the second information, the first scoring result, the second scoring result, and the result of clinical significance classification with viewing permissions.
[0065] For the protection of user privacy, the user can set viewing permissions for the first information, the second information, the first scoring result, the second scoring result, and the result of clinical significance classification of the copy number variation on the UAnno1.0 software. Since each input copy number variation position information and variation type generates corresponding first information, second information, first scoring result, second scoring result, and result of clinical significance classification, the permissions for the first information, second information, first scoring result, second scoring result, and result of clinical significance classification corresponding to the same input copy number variation position information and variation type are the same.
[0066] After receiving the permissions set by the user, the UAnno1.0 software can present the above information with viewing permissions and receive an instruction to print or download the above information.
[0067] In some alternative embodiments, by comparing the UAnno1.0 software with a commercially available CNV automated interpretation software ClassifyCNV from two aspects of detection consistency and detection efficiency, the main contents are as follows:
[0068] First, 1,670 CNVs with interpretation results were re-annotated using UAnno1.0 and ClassifyCNV respectively, and the annotation results are as follows Figure 2 shown. Through Figure 2 it can be obtained that: in terms of pathogenicity and likely pathogenicity, the consistency between the annotation results of UAnno1.0 and the known results is 92.42%, and the consistency between the annotation results of ClassifyCNV and the known results is 86.36%; in terms of benignity and likely benignity, the consistency between the annotation results of UAnno1.0 and the known results is 72.34%, and the consistency between the annotation results of ClassifyCNV and the known results is 40.96%. There are more results of unknown clinical significance in the annotation results of ClassifyCNV.
[0069] Secondly, the CNVs detected by NIPTPlus of 2 - 10M were used to compare the detection performance of the two, and the detection results are respectively as Figure 3 and Figure 4 shown. From Figure 3 and Figure 4 it can be seen that the proportion of pathogenic and likely pathogenic CNVs detected by the UAnno1.0 software is 25.6%, and the proportion of pathogenic and likely pathogenic CNVs detected by ClassifyCNV is 15.54%; the CNVs detected by NIPTPlus above 10M were used to compare the detection performance of the two, and the detection results are respectively as Figure 5 and Figure 6 shown. The proportion of pathogenic and likely pathogenic CNVs detected by UAnno1.0 is 98.94%, and the proportion of pathogenic and likely pathogenic CNVs detected by ClassifyCNV is 92.55%.
[0070] From the above data, it can be obtained that: UAnno1.0 has a higher accuracy rate for CNV annotation.
[0071] In some alternative embodiments, the UAnno1.0 software includes two modules: variant interpretation and user data management, as Figure 7 shown. After entering the interpretation system, click on the module pointed by the arrow to enter the interpretation system interface for CNV interpretation. The specific interpretation is as follows:
[0072] Step 1: Input the chromosome of the CNV, such as chr22, the start coordinate of the CNV is 18912231, and the end coordinate is 21465672. Select the mutation type as CNV deletion, and click "Mutation Interpretation" for interpretation. If there is an interpretation of this CNV in the database, the historical interpretation of this CNV will be displayed on the "Historical Interpretation Database Page" at this time. What is shown in the module is the discrimination of CNVs in the database that have been judged by predecessors and have an Overlap and Ratio both of 50% or more compared with the CNV interpreted by the user. The user can use this as a reference. The display page of the historical database is as Figure 8 shown.
[0073] Step 2: As Figure 9 shown, the interpretation page shows the evidence relied on for the interpretation, the scoring range of each piece of evidence, and the basis for scoring the evidence. Among them, the basis for scoring can be automatically filled according to the interpretation or manually edited and adjusted. According to the basis for scoring, the evidence can be automatically scored or manually adjusted. In the initial state, the left side is in the open state, and the score sliders on the right side are all defaulted to 0 points. After clicking the "Perform Interpretation" button, the automatic scoring situation of each piece of evidence and the basis for scoring will be shown on the interpretation page. For example, in this case, both Evidence 2A and 2B are scored 1 point, Evidence 3 scores 0.45 points, and Evidence XA scores 0.9 points. As Figure 10 shown, click "Manual Evidence Adjustment" to expand the evidence that has not been automated. These evidences can be supplemented and added relying on the user's own experience. The green buttons on the left side of this part are all in the closed state in the initial state, and the score sliders on the right side are all defaulted to 0 points. When using a certain piece of evidence, first open the green switch on the left side of the evidence, and then slide the score slider on the right side to supplement the scoring of an evidence. If the switch is not opened and only the score slider on the right side is slid, the use of this piece of evidence is invalid. If the evidence is not adjusted, the "Generate Interpretation Result" can be directly clicked. If the evidence needs to be adjusted, after all the evidence is adjusted and confirmed to be correct, click the "Generate Interpretation Result" button to obtain the discrimination of the final result of the CNV. When the sum of the scores of all evidences is greater than or equal to 1 point, the final score given here is 1 point. When the sum of the scores of all evidences is less than or equal to -1 point, the final score given here is -1 point. The schematic diagram of the interpretation result page is as Figure 11 shown.
[0074] Step 3: The database details page is a detailed display of the case evidence database used for the interpretation. Figure 7 The icons under "Link" in
[0075] are some databases relied on by the interpretation system. Clicking on the icon can enter the database. Figure 12As shown, the first part of gene annotation is used to determine whether there are protein-coding genes based on the database. 1 is a text description of the judgment on whether this CNV contains protein-coding genes. 2 can adjust the number of CNVs displayed on the page. 3 can perform ascending or descending sorting. 4 can perform precise searches. 5 can perform paging. In this example, the length of this CNV segment is 2.6 Mb, located in the region of 22q11.21 on chromosome 22, and contains protein-coding genes.
[0076] As Figure 13 shown, the second part of dosage sensitivity annotation is used to determine the coverage of known haploinsufficient (triploinsensitive) genes (regions) and common high-prevalence syndromes based on the database. Among them, genes with a score of 3 for haploinsufficient (triploinsensitive) genes are dosage-sensitive genes, and regions with a score of 3 for haploinsufficient (triploinsensitive) regions are dosage-sensitive regions. In this example, this CNV segment contains 0 triploinsensitive genes and 2 triploinsensitive regions (coverage above 80%). After querying the Decipher database for common high-prevalence syndromes, this CNV segment covers 100% of the pathogenic region of the known syndrome 22q11 duplication syndrome.
[0077] As Figure 14 shown, the third part of gene links is used to determine the number of protein-coding genes in the CNV segment. In this example, the CNV segment contains a total of 48 protein-coding genes, and 45 genes come from a single Gene Family. By querying the OMIM database, this CNV segment contains 39 OMIM genes, of which 13 are OMIM Morbid genes.
[0078] As Figure 15 shown, the fourth part of common population CNV annotation is used to determine the overlap of the CNV segment with common population variations. In this example, the results of population CNV analysis found 1 population CNV that completely covers this CNV segment, from DGV, with the highest population frequency of 0.07%.
[0079] As Figure 16 shown, the fifth part of historical CNV annotation is used to determine the situation of this CNV segment covering historical pathogenic and benign CNVs. In this example, this CNV segment covers 47 known CNVs in all pathogenic regions, including 8 from the Decipher database and 39 from the ISCA database. The results of benign CNV analysis show that no benign CNV that completely covers this CNV segment was found.
[0080] Step 4. The genome browser page is a visual display of CNV in the evidence library, as Figure 17As shown, clicking on "Tracks" in the upper left corner allows you to select the database to be displayed in the browser. Clicking the button on the right can also be used to make adjustments such as moving left and right, zooming in and out, etc.
[0081] Step 5. Click Figure 18 on the module pointed by the arrow in to enter the user's personal data management interface. This interface can be used to query and download user data. After entering this interface, you can see as Figure 19 shown. First, you can count the total number of CNVs queried by the user in this system, the number and proportion of deletions, duplications, as well as the number and proportion of (possibly) benign, of uncertain significance, (possibly) pathogenic CNVs. Second, you can also view the detailed information of each queried CNV, and at the same time, you can download or print it.
[0082] Figure 20 is a schematic diagram of the composition structure of an embodiment of a clinical annotation device for the pathogenicity of copy number variations provided by the present invention. As Figure 20 shown, the device includes:
[0083] An acquisition module 2001, configured to acquire the location information and variation type of the copy number variation; the location information includes the chromosome where the copy number variation is located, the start position of the copy number variation, and the end position of the copy number variation; the variation type includes copy number deletion and copy number duplication;
[0084] A first determination module 2002, configured to determine the first information of the copy number variation according to the location information and the variation type, where the first information includes whether it contains protein-coding genes, the number of protein-coding genes contained, the overlap information between the copy number variation and common population variations, the overlap information between the copy number variation and historical common pathogenic copy number variations, the overlap information between the copy number variation and historical common benign copy number variations, and the overlap information between the copy number variation and haploinsufficient genes or regions, or the overlap information between the copy number variation and triplo-sensitive genes or regions;
[0085] A first scoring module 2003, configured to compare the first information of the copy number variation with a preset database based on the copy number variation interpretation guidelines according to the variation type, score the first information, and obtain a first evaluation result; the preset database is established based on the ensembl, ClinGen, dbVar, DECIPHER, DGV, ISCA, OMIM databases;
[0086] A first classification module 2004, configured to clinically classify the pathogenicity of the copy number variation according to the first evaluation result to obtain the result of the clinical significance classification, where the result of the clinical significance classification includes benign, likely benign, of uncertain clinical significance, likely pathogenic, and pathogenic.
[0087] Optionally, the first scoring module 2003 includes:
[0088] A first scoring unit for scoring, when the mutation type is copy number deletion, whether the copy number deletion region contains protein-coding genes, the number of protein-coding genes contained, the overlapping information between the copy number variation and genes or regions sensitive to haploinsufficiency, the overlapping information between the copy number variation and common population variations, the overlapping information between the copy number variation and historical common pathogenic copy number variations, and the overlapping information between the copy number variation and historical common benign copy number variations.
[0089] Optionally, the first scoring module 2003 further includes:
[0090] A second scoring unit for scoring, when the mutation type is copy number duplication, whether the copy number duplication region contains protein-coding genes, the number of protein-coding genes contained, the overlapping information between the copy number variation and genes or regions sensitive to triploinsufficiency, the overlapping information between the copy number variation and common population variations, the overlapping information between the copy number variation and historical common pathogenic copy number variations, and the overlapping information between the copy number variation and historical common benign copy number variations.
[0091] Optionally, the device further includes:
[0092] A first receiving module for receiving the second information of the copy number variation; the second information includes the characteristics of the reported phenotypes of the genes or regions with copy number variation, the co-segregation characteristics of the pedigree, the statistical difference situation of the case-control study, the mode of inheritance of the disease, and the parental origin;
[0093] A second receiving module for receiving the score of the second information by the user based on experience to obtain a second scoring result;
[0094] A second grading module for clinically grading the pathogenicity of the copy number variation according to the first scoring result and the second scoring result to obtain the result of the clinical significance grading.
[0095] Optionally, the device further includes:
[0096] A display module for displaying the first scoring result, the second scoring result, and the result of the clinical significance grading.
[0097] Optionally, the device further includes:
[0098] A third receiving module for receiving the viewing permissions set by the user for the first information, the second information, the first scoring result, the second scoring result, and the result of the clinical significance grading of the copy number variation;
[0099] A fourth receiving module, configured to receive an instruction to print or download the first information, the second information, the first scoring result, the second scoring result, and the result of the clinical significance grading that have viewing permissions.
[0100] Figure 21 A schematic physical structure diagram of an electronic device provided by the present invention, as Figure 21 shown, the electronic device may include: a processor 2101, a communication interface 2102, a memory 2103, and a communication bus 2104. Among them, the processor 2101, the communication interface 2102, and the memory 2103 communicate with each other through the communication bus 2104. The processor 2101 can call the logical instructions in the memory 2103 to execute a method for clinically annotating the pathogenicity of copy number variations, and the method includes:
[0101] Obtaining the location information and the type of copy number variation; the location information includes the chromosome where the copy number variation is located, the start position of the copy number variation, and the end position of the copy number variation; the type of copy number variation includes copy number deletion and copy number duplication; according to the location information and the type of copy number variation, determining the first information of the copy number variation, the first information including whether it contains protein-coding genes, the number of protein-coding genes contained, the overlapping information between the copy number variation and the common population variation, the overlapping information between the copy number variation and the historically common pathogenic copy number variation, the overlapping information between the copy number variation and the historically common benign copy number variation, and the overlapping information between the copy number variation and the genes or regions with haploinsufficiency, or the overlapping information between the copy number variation and the genes or regions sensitive to triploid dosage; based on the copy number variation interpretation guidelines, comparing the first information of the copy number variation with a preset database according to the type of copy number variation, scoring the first information, and obtaining a first evaluation result; the preset database is established based on the ensembl, ClinGen, dbVar, DECIPHER, DGV, ISCA, and OMIM databases; clinically grading the pathogenicity of the copy number variation according to the first evaluation result to obtain the result of the clinical significance grading, and the result of the clinical significance grading includes benign, likely benign, uncertain clinical significance, likely pathogenic, and pathogenic.
[0102] In addition, when the logical instructions in the above-mentioned memory 2103 are implemented in the form of software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0103] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the method for clinically annotating the pathogenicity of copy number variations provided by the above-mentioned various methods. The method includes:
[0104] Obtain the location information and variation type of the copy number variation; the location information includes the chromosome where the copy number variation is located, the start position of the copy number variation, and the end position of the copy number variation; the variation type includes copy number deletion and copy number duplication; according to the location information and variation type, determine the first information of the copy number variation. The first information includes the number of encoded gene proteins, the overlapping information between the copy number variation and the variations in the general population, the overlapping information between the copy number variation and the historically common pathogenic copy number variations, and the overlapping information between the copy number variation and the genes or regions with haploinsufficiency, or the overlapping information between the copy number variation and the genes or regions sensitive to triploid dosage; based on the copy number variation interpretation guidelines, compare the first information of the copy number variation with a preset database according to the variation type, score the first information, and obtain a first evaluation result; the preset database is established based on the ensembl, ClinGen, dbVar, DECIPHER, DGV, ISCA, and OMIM databases; according to the first evaluation result, clinically significantly grade the pathogenicity of the copy number variation to obtain the result of the clinical significance grading. The result of the clinical significance grading includes benign, likely benign, uncertain clinical significance, likely pathogenic, and pathogenic.
[0105] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the method for clinically annotating the pathogenicity of copy number variations provided by the above-mentioned various methods. The method includes:
[0106] Obtain the location information and mutation type of the copy number variation; the location information includes the chromosome where the copy number variation is located, the start position of the copy number variation, and the end position of the copy number variation; the mutation type includes copy number deletion and copy number duplication; according to the location information and mutation type, determine the first information of the copy number variation, and the first information includes whether it contains protein-coding genes, the number of protein-coding genes contained, the overlapping information of the copy number variation with common population variations, the overlapping information of the copy number variation with historically common pathogenic copy number variations, the overlapping information of the copy number variation with historically common benign copy number variations, and the overlapping information of the copy number variation with haploinsufficient genes or regions, or the overlapping information of the copy number variation with triplo-sensitive genes or regions; based on the copy number variation interpretation guidelines, compare the first information of the copy number variation with a preset database according to the mutation type, score the first information, and obtain the first evaluation result; the preset database is established based on the ensembl, ClinGen, dbVar, DECIPHER, DGV, ISCA, and OMIM databases; according to the first evaluation result, clinically significant grading of the pathogenicity of the copy number variation is performed to obtain the result of clinically significant grading, and the result of clinically significant grading includes benign, likely benign, of uncertain clinical significance, likely pathogenic, and pathogenic.
[0107] The system embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative effort.
[0108] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product, and this computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or some parts of the embodiments.
[0109] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A clinical annotation method for the pathogenicity of copy number variations, characterized in that, Including: Obtaining the location information and mutation type of the copy number variation; the location information includes the chromosome where the copy number variation is located, the starting position of the copy number variation, and the ending position of the copy number variation; the mutation type includes copy number deletion and copy number duplication; According to the location information and the mutation type, determining the first information of the copy number variation, the first information including whether a protein-coding gene is included or the number of protein-coding genes included determined by the location information, and the overlapping information of the copy number variation with common population variations, the overlapping information of the copy number variation with historically common pathogenic copy number variations, the overlapping information of the copy number variation with historically common benign copy number variations, and the overlapping information of the copy number variation with haploinsufficient genes or regions, or the overlapping information of the copy number variation with triplo-sensitive genes or regions; Based on the copy number variation interpretation guidelines, comparing the first information of the copy number variation with a preset database according to the mutation type, and scoring the first information to obtain a first scoring result; the preset database is established based on the ensembl, ClinGen, dbVar, DECIPHER, DGV, ISCA, and OMIM databases; Clinically significantly grading the pathogenicity of the copy number variation according to the first scoring result to obtain the result of the clinical significance grading, the result of the clinical significance grading including benign, likely benign, of uncertain clinical significance, likely pathogenic, and pathogenic; The comparing the first information of the copy number variation with a preset database according to the mutation type and scoring the first information includes: When the mutation type is the copy number deletion, scoring whether the region of the copy number deletion contains a protein-coding gene, the number of protein-coding genes included, the overlapping information of the copy number variation with haploinsufficient genes or regions, the overlapping information of the copy number variation with common population variations, and the overlapping information of the copy number variation with historically common pathogenic copy number variations and the overlapping information of the copy number variation with historically common benign copy number variations; When the mutation type is the copy number duplication, scoring whether the region of the copy number duplication contains a protein-coding gene, the number of protein-coding genes included, the overlapping information of the copy number variation with triplo-sensitive genes or regions, the overlapping information of the copy number variation with common population variations, and the overlapping information of the copy number variation with historically common pathogenic copy number variations and the overlapping information of the copy number variation with historically common benign copy number variations.
2. The clinical annotation method for the pathogenicity of copy number variations according to claim 1, characterized in that, Also including: Receiving the second information of the copy number variation; the second information includes the characteristics of the reported phenotypes of the genes or regions where the copy number variation occurs, the co-segregation characteristics of the pedigree, the statistical difference situation of the case-control study, the mode of inheritance of the disease, and the parental origin; Receiving the score of the second information by the user based on experience to obtain a second scoring result; Based on the first scoring result and the second scoring result, the clinical significance of the copy number variation pathogenicity is classified to obtain the result of the clinical significance classification.
3. The method for clinical annotation of the pathogenicity of copy number variation according to claim 2, wherein Further included are: Displaying the first scoring result, the second scoring result, and the result of the clinical significance classification.
4. The method for clinically annotating the pathogenicity of copy number variations according to claim 3, wherein Further included are: Receiving the viewing permissions set by the user for the first information, second information, first scoring result, second scoring result, and the result of the clinical significance classification of the copy number variation; Receiving instructions to print or download the first information, second information, first scoring result, second scoring result, and the result of the clinical significance classification with the viewing permissions.
5. A clinical annotation device for the pathogenicity of copy number variations, characterized in that, Included are: An acquisition module for acquiring the location information and the variation type of the copy number variation; the location information includes the chromosome where the copy number variation is located, the starting position of the copy number variation, and the ending position of the copy number variation; the variation type includes copy number deletion and copy number duplication; A first determination module for determining the first information of the copy number variation according to the location information and the variation type, the first information including whether it contains protein-coding genes or the number of protein-coding genes determined by the location information, and the overlapping information of the copy number variation with common population variations, the overlapping information of the copy number variation with historically common pathogenic copy number variations, the overlapping information of the copy number variation with historically common benign copy number variations, and the overlapping information of the copy number variation with haploinsufficient genes or regions, or the overlapping information of the copy number variation with triplo-sensitive genes or regions; A first scoring module for comparing the first information of the copy number variation with a preset database based on the copy number variation interpretation guidelines according to the variation type, and scoring the first information to obtain a first scoring result; the preset database is established based on the ensembl, ClinGen, dbVar, DECIPHER, DGV, ISCA, OMIM databases; A first classification module for classifying the clinical significance of the copy number variation pathogenicity according to the first scoring result to obtain the result of the clinical significance classification, the result of the clinical significance classification including benign, likely benign, of uncertain clinical significance, likely pathogenic, and pathogenic; The first scoring module includes: a first scoring unit and a second scoring unit; The first scoring unit is used to score whether the region of copy number deletion contains protein-coding genes, the number of protein-coding genes, the overlapping information of the copy number variation with haploinsufficient genes or regions, the overlapping information of the copy number variation with common population variations, the overlapping information of the copy number variation with historically common pathogenic copy number variations, and the overlapping information of the copy number variation with historically common benign copy number variations when the variation type is copy number deletion; The second scoring unit is used to score whether the region of copy number duplication contains protein-coding genes, the number of protein-coding genes contained, the overlapping information between the copy number variation and genes or regions sensitive to triple dosage, the overlapping information between the copy number variation and common population variations, the overlapping information between the copy number variation and historical common pathogenic copy number variations, and the overlapping information between the copy number variation and historical common benign copy number variations when the variation type is copy number duplication.
6. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the clinical annotation method for the pathogenicity of copy number variations according to any one of claims 1 to 4.
7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the clinical annotation method for the pathogenicity of copy number variations according to any one of claims 1 to 4.
8. A computer program product having executable instructions stored thereon, characterized in that, When the instruction is executed by the processor, the processor is caused to implement the steps of the clinical annotation method for the pathogenicity of copy number variations according to any one of claims 1 to 4.
Citation Information
Patent Citations
Method for detecting copy number variation by means of single sample based on second-generation sequencing technology, and computer system
CN110246543A
Copy number variant analysis method and system and computer readable storage medium
CN110570902A