Methods, systems, devices, and computer-readable storage media for phenotypic analysis-based gene mutation

CN119229965BActive Publication Date: 2026-09-15AEGICARE (SHENZHEN) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411419811.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-11
Publication Date
2026-09-15
Estimated Expiration
2044-10-11

AI Technical Summary

Technical Problem

[0003]对受检者进行基因检测,通常会检测到较多基因突变,目前虽然有蛋白质功能预测软件(如SIFT、PolyPhen2、MutationTaster、DANN、CADD、primateAI等软件)、人群频率数据库(如1000Genomes Project、Exome Variant Server、The Exome AggregationConsortium等数据库)等功能预测工具对基因突变的致病性进行预测,但这些方法主要是从基因突变的功能性进行考虑,并未考虑患者表型与突变的关联,经功能预测后可能仍然有大量意义不明的突变,如何从这些大量的基因突变中查找可能与受检者所患疾病相关的突变,医生或医学解读人员需要花费大量时间去查阅文献和数据库,这对于解读人员的经验要求较高,且耗时较长

Benefits of technology

[0019] This application analyzes the phenotypic correlation between a subject's phenotype and the phenotypic genes containing their mutations. Gene mutations are scored based on this correlation, and then ranked according to the scores. This allows medical interpreters to prioritize gene mutations that may be related to the subject's phenotype, and these mutations are highly likely to be pathogenic mutations causing the phenotype, thus improving the efficiency of medical interpreters. This application considers the association between patient phenotypes and the genes containing their mutations, saving manpower and time costs, reducing the influence of subjective human interpretation, and making the results more standardized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119229965B_ABST
    Figure CN119229965B_ABST
Patent Text Reader

Abstract

The application discloses a method, system, device and computer readable storage medium for analyzing gene mutation based on phenotype, and the method comprises the following steps: data acquisition, acquiring data of phenotype and gene mutation of a subject; phenotype description, including describing the phenotype of the subject as a standard entry; gene scoring, including scoring each gene of the gene mutation of the subject according to genetic heterogeneity, genetic mode and phenotype comparison, to obtain a score of each gene; and sorting, sorting the gene mutation of the subject based on the score. According to the application, by analyzing the correlation between the phenotype of the subject and the phenotype of the gene where the mutation is located, scoring the gene mutation according to the correlation, and then sorting the mutation according to the score, medical interpreters can pay more attention to pathogenic gene mutations that may be related to the phenotype of the subject, and the working efficiency of the medical interpreters can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of genetic analysis technology, and in particular to a method, system, apparatus and computer-readable storage medium for phenotypic analysis of gene mutations. Background Technology

[0002] Genetic diseases are illnesses caused by genetic factors, typically involving mutations in one or more genes. Genetic testing is a technique that analyzes specific genes or regions in an individual's genome to determine genetic information, gene variations, and their potential impact. Genetic testing has a wide range of applications in the diagnosis of genetic diseases; it can help diagnose patients who are already showing symptoms, detect carriers, and assess the risk to family members.

[0003] Genetic testing of individuals typically detects numerous gene mutations. While functional prediction tools exist, such as protein function prediction software (e.g., SIFT, PolyPhen2, MutationTaster, DANN, CADD, primateAI) and population frequency databases (e.g., 1000 Genomes Project, Exome Variant Server, The Exome Aggregation Consortium), to predict the pathogenicity of gene mutations, these methods primarily consider the functionality of gene mutations and do not take into account the association between patient phenotype and mutations. Even after functional prediction, a large number of mutations of undetermined significance may still remain. Identifying mutations that may be related to the patient's disease from among these numerous gene mutations requires doctors or medical interpreters to spend a significant amount of time reviewing literature and databases. This demands a high level of experience from the interpreters and is time-consuming. Summary of the Invention

[0004] To address the aforementioned issues, this application provides a method, system, apparatus, and computer-readable storage medium for phenotypic analysis of gene mutations.

[0005] The following technical solution is adopted in this application:

[0006] The first aspect of this application discloses a method for gene mutation analysis based on phenotype, comprising: data acquisition, acquiring data on the phenotype and gene mutations of a subject; phenotypic description, including standardizing the phenotype description of the subject into terms; gene scoring, including assigning a score to each gene of the subject's gene mutation based on genetic heterogeneity, inheritance pattern, and phenotypic comparison, obtaining a score for each gene; and ranking, ranking the subject's gene mutations based on the scores, with genes having higher scores being more likely to be disease-causing genes of the subject; wherein, the gene scoring includes the following cases: if the disease corresponding to the gene mutation has genetic heterogeneity, then the gene is assigned a score of 0; if the disease corresponding to the gene does not have genetic heterogeneity, and the inheritance pattern of the disease corresponding to the gene is dominant: if all the standardized terms of the subject's phenotype are included in the standardized terms of the gene's corresponding phenotype, then the gene is assigned a score of 2; if the standardized terms of the subject's phenotype are partially included in the gene... If the standardized phenotype terminology for the subject's phenotype does not overlap with the standardized phenotype terminology for the corresponding gene, then the similarity between the subject's phenotype and the standardized phenotype terminology for the gene containing the subject's gene mutation is compared. If the similarity is greater than a predetermined value, then the gene is assigned 1 point. If the disease corresponding to the gene does not exhibit genetic heterogeneity, and the disease's inheritance pattern is recessive: if the standardized phenotype terminology for the subject's phenotype is entirely contained within the standardized phenotype terminology for the corresponding gene, then the gene is assigned 1 point; if the standardized phenotype terminology for the subject's phenotype is partially contained within the standardized phenotype terminology for the corresponding gene, then the gene is assigned 0.5 points. If the standardized phenotype terminology for the subject's phenotype does not overlap with the standardized phenotype terminology for the corresponding gene, then the similarity between the subject's phenotype and the standardized phenotype terminology for the gene containing the subject's gene mutation is compared. If the similarity is greater than a first predetermined value, then the gene is assigned 0.5 points.

[0007] It should be noted that in this application, by analyzing the phenotypic correlation between the subject's phenotype and the phenotypic genes containing the mutations, assigning scores to gene mutations based on the correlation, and then ranking the mutations according to the scores, medical interpreters can prioritize gene mutations that may be related to the subject's phenotype. These mutations are highly likely to be pathogenic mutations causing the subject's phenotype, thus improving the efficiency of medical interpreters. This application considers the association between the patient's phenotype and the gene containing the mutation, saving manpower and time costs, reducing the influence of subjective human interpretation, and making the results more standardized.

[0008] In one implementation of this application, the standardized terms are HPO terms. It should be noted that HPO (Human Phenotype Ontology) is a standardized vocabulary used to describe human phenotypic characteristics. HPO terms refer to specific phenotypic characteristics or symptoms defined in HPO. Each term has a unique identifier (ID), usually beginning with "HP:" followed by a series of numbers. For example, "HP:0000010" refers to the term "Recurrenturinary tract infections". It should also be noted that HPO terms can be Chinese terms downloaded from CHPO (Chinese Human Phenotype Ontology), a terminology set similar to HPO used to translate and optimize existing HPO terms from abroad. CHPO aims to provide a standardized terminology system for the diagnosis of genetic diseases, gene function research, and clinical data analysis in the Chinese population.

[0009] One implementation of this application further includes: filtering the gene mutations before data acquisition, or filtering the gene mutations after sorting, wherein the filtering includes excluding benign mutations and suspected benign mutations among the gene mutations. It should be noted that gene mutations can be filtered using databases such as gnomAD, clinvar, dbSNP, ExAC, and the 1000GenomesProject to remove benign and suspected benign mutations, thereby reducing the number of gene mutations that interpreters need to focus on and improving interpretation efficiency.

[0010] In one implementation of this application, the gene mutation is a de novo mutation, meaning a gene mutation that is appearing for the first time in the subject and is not inherited from the parents. It should be noted that when a child has a genetic disease while the parents are healthy, a de novo mutation is often the cause of the child's illness. Accurately assessing the pathogenicity of a de novo mutation is crucial for the diagnosis and treatment of genetic diseases. Therefore, analyzing the de novo mutation in the subject is beneficial for identifying the pathogenic gene in the subject.

[0011] In one implementation of this application, the similarity is calculated using the Python open-source package pyhpo, and the first predetermined value is 0.1. It should be noted that pyhpo is a Python library specifically designed for processing Human Phenotypic Title Sets (HPOs), available at: https: / / pyhpo.readthedocs.io / en / latest / tutorial / basics.html and GitHub: https: / / github.com / anergictcell / pyhpo. It supports calculating phenotypic similarity. In bioinformatics, calculating phenotypic similarity is crucial for genetic disease diagnosis, gene function research, and clinical data analysis. The phenotypic similarity between two cases can be calculated using the `calculate_similarity` function provided by pyhpo: `similarity_score = calculate_similarity(phenotypes_case_1, phenotypes_case_2, ontology) print(f"Similarity score between Case 1 and Case 2:{similarity_score}")`.

[0012] One implementation of this application further includes: storing data from multiple previously analyzed subjects in a local database; after gene scoring and before sorting, calculating the similarity between the normalized terms of the current subject's phenotype and the normalized terms of the phenotypes of previous subjects in the local database; adding the scores of genes from previous subjects with similarity greater than a second predetermined value to the scores of the same genes in the current subject to obtain the score for each gene in the current subject. It should be noted that by first scoring mutations in a large number of subjects and storing the information in a database, when scoring subsequent patients (the current subject), the database is first searched for cases with similar diseases and mutations or new mutations in the same genes. The more such cases there are, the more likely the mutation or new mutation is to be related to the disease. Adding the gene scores of similar cases to the current mutation to obtain the final gene score can further enhance the association between the subject's phenotype and the gene containing the new mutation.

[0013] In one implementation of this application, the similarity is 2 * the number of identical terms in the normalized terms of the current subject's phenotype and the normalized terms of the past subject's phenotype / (the number of normalized terms of the current subject's phenotype + the number of normalized terms of the past subject's phenotype), and the second predetermined value is 0.5.

[0014] One implementation of this application further includes: converting the score of each gene into a gene pathogenicity level; if the gene score is less than 1, the gene pathogenicity level is extremely low; if the gene score is greater than or equal to 1 and less than 2, the gene pathogenicity level is low; if the gene score is greater than or equal to 2 and less than 4, the gene pathogenicity level is relatively high; and if the gene score is greater than or equal to 4, the gene pathogenicity level is extremely high. It should be noted that this allows for the classification of gene mutations in the tested individual, enabling interpreters to prioritize mutations with higher pathogenicity.

[0015] The second aspect of this application discloses a system for gene mutation analysis based on phenotypic characteristics, comprising: a data acquisition module for acquiring data on the phenotype and gene mutations of a subject; a phenotypic description module for describing the subject's phenotype as standardized terms; a gene scoring module for assigning a score to each gene of the subject's gene mutations based on genetic heterogeneity, inheritance pattern, and phenotypic comparison, thereby obtaining a score for each gene; and a ranking module for ranking the subject's gene mutations based on the scores, wherein genes with higher scores are more likely to be the subject's disease-causing genes; wherein the gene scoring module includes: if the disease corresponding to the gene mutation has genetic heterogeneity, then assigning 0 points to that gene; if the disease corresponding to the gene does not have genetic heterogeneity, and the inheritance pattern of the disease corresponding to the gene is dominant; if all standardized terms of the subject's phenotype are included in the standardized terms of the gene's corresponding phenotype, then assigning 2 points to that gene; if the standardized terms of the subject's phenotype are partially included in the standardized terms of the gene's corresponding phenotype. If the standardized phenotype terminology for the gene is included in the gene's corresponding phenotype standard entry, then the gene is assigned 1 point. If there is no overlap between the standardized phenotype terminology for the subject's phenotype and the standardized phenotype terminology for the gene corresponding to the subject's gene mutation, then the similarity between the subject's phenotype and the standardized phenotype terminology for the gene mutation is compared. If the similarity is greater than a predetermined value, then the gene is assigned 1 point. If the disease corresponding to the gene does not exhibit genetic heterogeneity and the inheritance pattern of the disease is recessive: if the standardized phenotype terminology for the subject is entirely included in the standardized phenotype terminology for the gene, then the gene is assigned 1 point; if the standardized phenotype terminology for the subject is partially included in the standardized phenotype terminology for the gene, then the gene is assigned 0.5 points. If there is no overlap between the standardized phenotype terminology for the subject's phenotype and the standardized phenotype terminology for the gene, then the similarity between the subject's phenotype and the standardized phenotype terminology for the gene mutation is compared. If the similarity is greater than a first predetermined value, then the gene is assigned 0.5 points.

[0016] A third aspect of this application discloses an apparatus for gene mutation analysis based on phenotypic analysis, comprising a memory and a processor, wherein the memory is used to store a program, and the processor executes the program stored in the memory to implement the gene mutation analysis method based on phenotypic analysis as described in the first aspect of this application.

[0017] The fourth aspect of this application discloses a computer-readable storage medium storing a program that can be executed by a processor to implement the method for phenotypic analysis of gene mutations as described in the first aspect of this application.

[0018] The beneficial effects of this application are as follows:

[0019] This application analyzes the phenotypic correlation between a subject's phenotype and the phenotypic genes containing their mutations. Gene mutations are scored based on this correlation, and then ranked according to the scores. This allows medical interpreters to prioritize gene mutations that may be related to the subject's phenotype, and these mutations are highly likely to be pathogenic mutations causing the phenotype, thus improving the efficiency of medical interpreters. This application considers the association between patient phenotypes and the genes containing their mutations, saving manpower and time costs, reducing the influence of subjective human interpretation, and making the results more standardized. Attached Figure Description

[0020] Figure 1 This is a flowchart of the method for gene mutation based on phenotypic analysis involved in this application.

[0021] Figure 2 This is a schematic diagram illustrating the principle of phenotypic similarity threshold calculation involved in this application.

[0022] Figure 3 This is the Sanger verification result of the DYNC1H1 gene mutation in the subject sample in the embodiments of this application.

[0023] Figure 4 This is the Sanger verification result of the DYNC1H1 gene mutation in the mother sample of the subject in this application embodiment.

[0024] Figure 5 This is the Sanger verification result of the DYNC1H1 gene mutation in the father's sample of the subject in this application embodiment.

[0025] Figure 6 This is the Sanger verification result of the BICD2 gene mutation in the subject sample in the embodiments of this application.

[0026] Figure 7 This is the Sanger verification result of the BICD2 gene mutation in the mother sample of the subject in this application embodiment.

[0027] Figure 8 This is the Sanger verification result of the BICD2 gene mutation in the father's sample of the subject in this application embodiment. Detailed Implementation

[0028] The present invention will now be described in further detail with reference to specific embodiments and accompanying drawings. In the following embodiments, many details are described to facilitate a better understanding of the present application. However, those skilled in the art will readily recognize that some features may be omitted in different situations, or may be replaced by other materials or methods. In some cases, certain operations related to the present application are not shown or described in the specification. This is to avoid obscuring the core parts of the present application with excessive description. For those skilled in the art, detailed description of these related operations is not necessary; the relevant operations can be fully understood based on the description in the specification and general technical knowledge in the art.

[0029] Furthermore, the features, operations, or characteristics described in the specification can be combined in any suitable manner to form various embodiments. At the same time, the steps or actions in the method description can be rearranged or adjusted in a manner obvious to those skilled in the art. Therefore, the various orders in the specification and drawings are only for the clear description of a particular embodiment and do not imply a necessary order, unless otherwise stated that a particular order must be followed.

[0030] Genetic testing of individuals typically detects numerous gene mutations. While functional prediction tools such as protein function prediction software and population frequency databases exist to predict the pathogenicity of gene mutations, these methods primarily consider the functionality of the mutations without taking into account the association between the patient's phenotype and the mutation. Even after functional prediction, a large number of mutations of unknown significance may still remain. Identifying mutations that may be related to the patient's disease from among these numerous gene mutations requires doctors or medical interpreters to spend a significant amount of time reviewing literature and databases. This demands a high level of experience from the interpreters and is time-consuming.

[0031] De novo mutations are mutations not inherited from parents. In cases where a child has a genetic disease while the parents are healthy, de novo mutations are often the cause of the child's illness. Accurately assessing the pathogenicity of de novo mutations is crucial for the diagnosis and treatment of genetic diseases. However, de novo mutations are usually numerous, requiring doctors or medical interpreters to spend considerable time searching literature and databases to identify pathogenic mutations. This method is time-consuming and generally requires experienced personnel. Currently, there are no computer programs specifically designed to sort de novo mutations. Most common methods rely on protein function prediction software and population frequency data. These methods have limitations and do not consider the association between patient phenotype and de novo mutations.

[0032] This application first extracts standardized phenotypic information from patients' medical records or disease descriptions. It analyzes the association between this phenotypic information and the gene containing the novel mutation, and scores the novel mutation by combining it with similar cases from a database. This database case analysis enhances the association between the examinee's phenotype and their gene mutation (finding cases in the database with similar diseases and novel mutations in the same gene; the more such cases, the more likely the novel mutation is related to the disease). These scores are accumulated to obtain a final score for each novel mutation. All novel mutations are ranked according to the final score, prioritizing those most relevant to the patient's disease. This reduces the time doctors and medical interpreters spend searching for pathogenic mutations and is more reasonable and reliable than existing computer methods. Focusing on the top-ranked novel mutations and conducting further validation (such as Sanger validation) reveals novel mutations directly related to the disease.

[0033] This application provides a method, system, apparatus, and computer-readable storage medium for phenotypic analysis of gene mutations.

[0034] Figure 1 This is a flowchart of the method for gene mutation based on phenotypic analysis involved in this application.

[0035] like Figure 1 As shown, in a specific embodiment, the method for analyzing gene mutations based on phenotypes may include: data acquisition, acquiring data on the phenotype and gene mutations of the subject; phenotypic description, describing the subject's phenotype as standardized terms; gene scoring, assigning a score to each gene mutation of the subject based on genetic heterogeneity, inheritance pattern, and phenotypic comparison; scoring based on a local database, searching the local database for cases with similar phenotypes to the subject and mutations in the same genes, adding the gene scores of the cases to the subject's gene scores to obtain the final score for each gene of the subject; and sorting, sorting the subject's gene mutations based on the scores.

[0036] The present application will be further described in detail below through specific embodiments. The following embodiments are only for further illustration of the present application and should not be construed as limiting the present application.

[0037] It should be noted that this embodiment primarily analyzes newly emerging mutations, but it can also be applied to other mutation geometries, such as for scoring and ranking known mutation sets associated with a disease. This embodiment does not filter the mutation list; however, filtering can be performed according to the actual research objectives.

[0038] Example 1:

[0039] (1) Download the HPO Chinese terminology from the CHPO official website (https: / / www.chinahpo.net / ). The file contains 16,493 lines and three columns. The first column is the HPO English terminology, the second column is the HPO number, and the third column is the HPO Chinese terminology, as shown in Table 1 below (due to copyright restrictions, only 10 terms are shown in this example):

[0040] Table 1

[0041] Recurrenturinarytractinfections HP:0000010 Recurrent urinary tract infection Neurogenic bladder HP:0000011 Neurogenic bladder dysfunction Urinaryurgency HP:0000012 Urgency Hypoplasia of theuterus HP:0000013 Uterine hypoplasia Abnormality of the bladder HP:0000014 Bladder abnormalities Bladderdiverticulum HP:0000015 bladder diverticulum Urinary retention HP:0000016 Urinary retention Nocturia HP:0000017 Nocturia Urinaryhesitancy HP:0000019 Hesitation in urination Urinaryincontinence HP:0000020 Urinary incontinence Megacystis HP:0000021 Megabladder

[0042] Download the gene and phenotype corresponding files from the OMIM website (https: / / www.omim.org / ), and associate the gene with the HPO Chinese terminology to generate a new file. This file contains 8962 records. The first column is the OMIM ID of the phenotype, the second column is the English name of the disease, the third column is the mode of disease inheritance, the fourth column is the gene name, the fifth column is the OMIM ID of the gene, and the sixth column is the HPO ID and Chinese terminology, as shown in Table 2 below (due to copyright restrictions, only 5 records are shown in this example):

[0043] Table 2

[0044]

[0045]

[0046]

[0047] (2) The subject was a fetus with abnormalities detected by fetal ultrasound. Amniotic fluid samples were taken from the fetus, and blood samples were drawn from the parents for family whole-exome sequencing. The sequencing data were compared to obtain the VCF files of the three family members. After merging the family VCFs, a list of newly discovered mutations in the subject was obtained. The first column was formatted as chromosome number of the mutation / hg19 coordinate / reference base / mutated base, and the second column was the gene name of the mutation. The list of newly discovered mutations is shown in Table 3 below:

[0048] Table 3

[0049]

[0050]

[0051]

[0052]

[0053]

[0054]

[0055] (3) Extract standardized HPO terms from the examinee's medical records or disease descriptions:

[0056] 1) Matching is performed in the original text using a self-built thesaurus. Each synonym corresponds to an HPO entry; that is, when any synonym appears in the original text, its corresponding HPO entry will be extracted. For example, 'NT thickening' corresponds to...

[0057] If 'NT thickening' appears in a patient's medical record, then 'HP:0010880 increased nuchal translucency thickness' will be included as one of the HPO entries for that patient.

[0058] 2) First, segment the patient's medical record into different sentences according to Chinese punctuation marks (such as periods and commas). Then, use open-source software to segment the Chinese HPO terms. If each segmented word can be completely matched in any sentence of the original text, then that HPO term will be one of the patient's HPO terms. For example, 'HP:0007340 lower limb muscle weakness' will be segmented into 'lower limb', 'muscle', and 'weakness'. If the patient's medical record contains 'the patient reported slight weakness in the lower limb muscles', then 'HP:0007340 lower limb muscle weakness' will be one of the patient's HPO terms.

[0059] 3) In some cases, neither of the above two methods can match HPO terms. In this case, fuzzy matching is used to compare the word similarity of each sentence with the Chinese HPO terms, and the one with the highest similarity is selected. For example, if the patient's medical record only has a simple description of 'oligospermia', then all Chinese HPO terms are taken, and each Chinese term is compared with 'oligospermia'. Finally, 'HP:0000798 Oligospermia' has the highest similarity to the original text, so 'HP:0000798 Oligospermia' is selected as one of the patient's HPO terms.

[0060] The phenotypic and normalized descriptions of the subjects in this embodiment are shown in Table 4 below:

[0061] Table 4

[0062]

[0063]

[0064] (4) Gene scoring, local database comparison, and sorting of the newly discovered mutations of the subjects:

[0065] 1) Record relevant information in the OMIM file (i.e., the file shown in Table 2): the mode of inheritance of each gene, i.e., dominant or recessive inheritance; analyze the genetic heterogeneity of the diseases corresponding to the genes: record all diseases corresponding to each gene, and record all genes corresponding to each disease. When a disease corresponds to multiple genes, the disease has genetic heterogeneity. When a disease corresponds to only one gene and is unrelated to other genes, there is no genetic heterogeneity.

[0066] 2) Phenotypic comparisons were performed between the HPO phenotype combination corresponding to each newly discovered mutation gene and the patient's HPO phenotype combination, and genetic heterogeneity was scored according to the mode of inheritance:

[0067] i. The disease corresponding to this gene is dominantly inherited: a) If all of the patient's phenotypic HPO terms are included in the phenotypic HPO terms corresponding to this gene, or if all of the phenotypic HPO terms corresponding to this gene are included in the patient's phenotypic HPO terms, then the phenotypic similarity is high, and 2 points are assigned; b) If more than one (not all) of the patient's phenotypic HPO terms appear in the phenotypic HPO terms corresponding to this gene, then the phenotypic similarity is moderate, and 1 point is assigned; c) If the patient's phenotypic HPO terms have no overlap with the phenotypic HPO terms corresponding to this gene, then the similarity between the two HPO groups is calculated using the Python open-source package pyhpo. If the similarity score is greater than 0.1 (the threshold is described later), then 1 point is assigned;

[0068] ii. The disease corresponding to this gene is recessive: a) If all of the patient's phenotypic HPO terms are included in the phenotypic HPO terms corresponding to this gene, or if all of the phenotypic HPO terms corresponding to this gene are included in the patient's phenotypic HPO terms, then the phenotypic similarity is high, and 1 point is assigned; b) If more than one (not all) of the patient's phenotypic HPO terms appear in the phenotypic HPO terms corresponding to this gene, then the phenotypic similarity is moderate, and 0.5 points are assigned; c) If the patient's phenotypic HPO terms have no overlap with the phenotypic HPO terms corresponding to this gene, then the similarity between the two HPO groups is calculated using the Python open-source package pyhpo. If the similarity score is greater than 0.1, 0.5 points are assigned.

[0069] iii. If the disease corresponding to this gene exhibits genetic heterogeneity, then all mutations on this gene are assigned a score of 0.

[0070] The similarity threshold is obtained using the following method: Figure 2 This is a schematic diagram illustrating the principle of phenotypic similarity threshold calculation involved in this application, combined with... Figure 2The study selected a number of patient samples who had undergone family-based three-person whole-exome sequencing. The phenotypes of these patients had been standardized into HPO terminology combinations by experienced medical personnel, denoted as patient_HPO. For each patient, at least one pathogenic gene was identified based on their phenotype and the pathogenic effect of the de novo mutation. Each pathogenic gene could be found in the OMIM database as a corresponding disease syndrome. These disease syndromes corresponded to several HPO terms, with each disease's corresponding HPO terminology combination denoted as gene_HPO1, gene_HPO2, ..., gene_HPOn. The similarity score S between each gene_HPO and patient_HPO was calculated using pypho (the phenotype similarity between two cases was calculated using the `calculate_similarity` function provided by pypho).

[0071] = `similarity_score = calculate_similarity(phenotypes_case_1, phenotypes_case_2, ontology)` print(f"Similarity score between Case 1 and Case 2: {similarity_score}"))`, finally taking the maximum similarity score `Smax` for each patient. Collect the `Smax` of all patients. Assuming these scores follow a normal distribution, the expected value plus or minus three standard deviations contains the scores of 99.7% of the patients. Therefore, the expected value minus three standard deviations is taken as the score threshold `Scutoff`. In this embodiment, by calculating the distribution of scores from 100 samples, the score threshold `Scutoff` is set to 0.1.

[0072] 3) Search the local database for HPO term combinations that contain any HPO combination matching the current patient's HPO combination, narrowing the search scope. Then, calculate the similarity between the current patient's HPO combination and these HPO combinations by iterating through them. The similarity calculation formula is: 2 * (number of patient HPO combinations matching the HPO combination of a current database case) / (number of patient HPO combinations + number of current database case HPO combinations). For example, if the current patient's HPO combinations are HP:0001155: hand abnormality, HP:0001838: rocking foot, HP:0002533: posture abnormality, HP:0006101: finger syndactyly, and there is a case in the database with the same HPO combination: HP:0002533: posture abnormality, HP:0006101: finger syndactyly, the similarity is 2 * 2 / (4 + 2) = 0.6. Select records with a phenotypic similarity greater than 0.5. If no record has a similarity greater than 0.5, do not proceed to the next step. Next, the record with the highest similarity is selected. If multiple records have the same similarity, the one with the highest mutation score is selected, and its mutation score is added to the current patient's mutation score to obtain the final score. After scoring is completed, the current patient's sample number, the mutation information (including chromosome number, genomic coordinates (hg19 version), reference bases, mutated bases, the gene in which they are located, the patient's HPO entry and Chinese text, and the pathogenicity score of the new mutation are stored in the database.

[0073] (5) Output the final scores and intermediate results as shown in Table 5 below. This table has been sorted in descending order according to the gene scores:

[0074] Table 5

[0075]

[0076]

[0077]

[0078]

[0079]

[0080]

[0081]

[0082]

[0083]

[0084]

[0085]

[0086] In Table 5, the first column contains mutation information (chromosome number of the mutation / hg19 coordinates / reference base / mutated base); the second column contains the name of the mutated gene; the third column contains the gene's inheritance pattern, AD for dominant inheritance and AR for recessive inheritance; the fourth column indicates whether the disease corresponding to the gene is genetically heterogeneous, with 1 for yes and 0 for no; columns five through seven are related to gene scoring, with column five showing the number of HPOs in the patient's phenotype that are identical to the HPOs corresponding to the gene, and column six showing the similarity between the patient's HPO term combinations calculated by the pyhpo package and the HPO combinations corresponding to the gene; column seven shows the score given to the current mutation based on phenotypic comparison, gene inheritance pattern, and genetic heterogeneity; columns eight through nine are related to local database searches, with column eight showing the medical record number and mutation information of similar cases in the database, and column nine showing the score of the gene in that case; and column ten shows the final score of the mutation, with the list of newly discovered mutations sorted in descending order according to the final score.

[0087] The following example uses the scoring process for de novo mutations in the DYNC1H1 gene to illustrate the analysis of the table above:

[0088] 1) The patient’s HPO combination is HP:0001155: hand abnormality, HP:0001838: rocking foot, HP:0002533: postural abnormality, HP:0006101: finger syndactyly. There are three diseases related to DYNC1H1 in the OMIM database, with OMIM IDs #614228, #614563 and #158600 respectively.

[0089] 2) The number of HPO entries corresponding to these three OMIM IDs that are completely consistent with the patient's HPO entries is 0, so the value in the fifth column is 0.

[0090] 3) Among them, the HPO combination corresponding to #614563 is "HP:0002365 brainstem dysplasia, HP:0001249 intellectual disability, HP:0001302 megalencephalopathy, HP:0001250 epileptic seizures, HP:0000252 microcephaly, HP:0003477 axonal peripheral neuropathy, HP:0002079 corpus callosum agenesis, HP:0007359 focal epileptic seizures, HP:0200055 microcephaly, HP:0001288 gait instability, HP:0000494 hyposlant palpebral fissure, HP:0000006 autosomal dominant inheritance, The combination "HP:0002510 Spastic quadriplegia, HP:0001357 Plagiocerebellar deformity, HP:0001760 Foot abnormality, HP:0001321 Cerebellar hypoplasia, HP:0011220 Forehead protrusion, HP:0001252 Hypotonia, HP:0001265 Weakened tendon reflexes" has a similarity score of 0.26 with the patient's HPO combination "HP:0001155: Hand abnormality, HP:0001838: Rocking chair foot, HP:0002533: Postural abnormality, HP:0006101: Syndactyly" calculated using pyhpo. The similarity scores of the other two OMIM IDs are 0.11 and 0.004, respectively. Therefore, the maximum value of 0.26 is taken as the value of the sixth column.

[0091] 4) Diseases related to the DYNC1H1 gene do not exhibit genetic heterogeneity. The DYNC1H1 gene is inherited in an autosomal dominant manner. Furthermore, the similarity score between the HPO combination of the DYNC1H1 gene-related disease and the patient's own HPO combination is 0.26, which is greater than 0.1. Therefore, according to the scoring rules, the DYNC1H114 / 102476743 / C / T mutation is assigned 1 point, which is the value in the seventh column.

[0092] 5) A similar disease was found in the local database, and it was a newly discovered mutation case that also occurred on the DYNC1H1 gene: AS23184784:chr14:102467292. The score of the DYNC1H1 gene in this case is 4 (the value in the ninth column). This score is added to the current case, and the final score of 14 / 102476743 / C / T is 5, which is the value in the tenth column.

[0093] For the mutations listed first (DYNC1H1) and second (BICD2) in the table above, samples from all three family members have undergone Sanger validation, confirming that these two mutations are indeed de novo mutations. The Sanger validation results are as follows: Figures 3 to 8 As shown, Figure 3The Sanger verification results of the DYNC1H1 gene mutation in the subject sample in this application embodiment show that the subject carries the DYNC1H1 gene mutation: 14 / 102476743 / C / T; Figure 4 The Sanger verification results of the DYNC1H1 gene mutation in the mother sample of the subject in this application embodiment show that there is no mutation at position 14 / 102476743 in the mother sample of the subject; Figure 5 The Sanger verification results of the DYNC1H1 gene mutation in the father's sample of the subject in this application embodiment show that there is no mutation at position 14 / 102476743 in the father's sample. Figure 6 The Sanger verification results of the BICD2 gene mutation in the subject sample in this application embodiment show that the subject carries the BICD2 gene mutation: 9 / 95484981 / T / G; Figure 7 The Sanger verification results of the BICD2 gene mutation in the mother's sample of the subject in this application embodiment show that there is no mutation at position 9 / 95484981 in the mother's sample. Figure 8 The Sanger verification results of the BICD2 gene mutation in the father's sample of the subject in this application embodiment show that there is no mutation at position 9 / 95484981 in the father's sample.

[0094] Mutations in the DYNC1H1 gene are typically associated with complex cortical dysplasia with other brain malformations type 13, axonal peroneal muscular atrophy type 20, and spinal muscular atrophy type 1 (lower limb affected). Mutations in the BICD2 gene are typically associated with autosomal dominant lower limb affected spinal muscular atrophy type 2B and autosomal dominant lower limb affected spinal muscular atrophy type 2A. Experienced medical interpreters have determined that these two mutations are likely pathogenic and closely related to the fetal phenotype. This demonstrates that this application effectively prioritizes newly discovered mutations related to the patient's phenotype, ensuring their rapid identification by physicians and medical interpreters.

[0095] The above description, in conjunction with specific embodiments, provides a further detailed explanation of this application and should not be construed as limiting the specific implementation of this application to these descriptions. Those skilled in the art to which this application pertains can make several simple deductions or substitutions without departing from the concept of this application.

Claims

1. A method for analyzing a genetic mutation based on a phenotype, characterized by, include: Data acquisition involves obtaining data on the phenotype and gene mutations of the subjects. Phenotypic description, including a standardized terminology describing the phenotype of the subject; Gene scoring includes assigning a score to each gene with a gene mutation in the subject based on genetic heterogeneity, inheritance pattern, and phenotypic comparison, thereby obtaining a score for each gene; The gene mutations of the examinee are sorted according to the scores, and the higher the score, the higher the probability that the gene is the disease-causing gene of the examinee. The gene scoring includes the following cases: If the disease corresponding to the mutated gene exhibits genetic heterogeneity, then the gene is assigned a score of 0. If the disease corresponding to the gene does not exhibit genetic heterogeneity, and the inheritance pattern of the disease corresponding to the gene is dominant: if all the standardized terms of the subject's phenotype are included in the standardized terms of the phenotype corresponding to the gene, then the gene is assigned 2 points; if some of the standardized terms of the subject's phenotype are included in the standardized terms of the phenotype corresponding to the gene, then the gene is assigned 1 point; if there is no overlap between the standardized terms of the subject's phenotype and the standardized terms of the phenotype corresponding to the gene, then the similarity between the subject's phenotype and the standardized terms of the phenotype corresponding to the gene where the subject's gene mutation is located is compared, and if the similarity is greater than a predetermined value, then the gene is assigned 1 point. If the disease corresponding to the gene does not exhibit genetic heterogeneity, and the inheritance pattern of the disease corresponding to the gene is recessive: if all the standardized terms of the subject's phenotype are included in the standardized terms of the phenotype corresponding to the gene, then the gene is assigned a score of 1; if some of the standardized terms of the subject's phenotype are included in the standardized terms of the phenotype corresponding to the gene, then the gene is assigned a score of 0.5; if there is no overlap between the standardized terms of the subject's phenotype and the standardized terms of the phenotype corresponding to the gene, then the similarity between the subject's phenotype and the standardized terms of the phenotype corresponding to the gene where the subject's gene mutation is located is compared. If the similarity is greater than a first predetermined value, then the gene is assigned a score of 0.

5.

2. The method of claim 1, wherein, The standardized term is the HPO term.

3. The method of claim 1, wherein, Also includes: The gene mutations are filtered before the data is acquired, or after sorting, and the filtering includes excluding benign mutations and suspected benign mutations among the gene mutations.

4. The method of claim 1, wherein, The gene mutation is a de novo mutation, which means that the gene mutation is appearing for the first time in the subject and is not inherited from the parents.

5. The method according to claim 1, characterized in that, The similarity is calculated using the Python open-source package pyhpo, and the first predetermined value is 0.

1.

6. The method according to any one of claims 1 to 5, characterized in that, Also includes: Data from multiple previously analyzed subjects are stored in a local database. After gene scoring and before sorting, the similarity between the normalized terms of the current subject's phenotype and the normalized terms of the phenotypes of previous subjects is calculated by traversing the local database. The scores of genes from previous subjects with similarity greater than a second predetermined value are added to the scores of the same genes in the current subject to obtain the score of each gene in the current subject.

7. The method according to claim 6, characterized in that, The similarity is calculated as 2 * the number of terms that are the same in the normalized terms of the current subject's phenotype and the normalized terms of the past subject's phenotype / (the number of normalized terms of the current subject's phenotype + the number of normalized terms of the past subject's phenotype), and the second predetermined value is 0.

5.

8. A system for analyzing gene mutations based on phenotypic characteristics, characterized in that, include: The data acquisition module is used to acquire data on the phenotype and gene mutations of the subjects. The phenotypic description module is used to describe the phenotypic description of the subject as a standardized term; The gene scoring module assigns a score to each gene with a gene mutation in the subject based on genetic heterogeneity, inheritance pattern, and phenotypic comparison, thus obtaining a score for each gene. The sorting module sorts the gene mutations of the examinee based on the score, and the higher the score, the higher the probability that the gene is the disease-causing gene of the examinee. The gene scoring module includes: If the disease corresponding to the mutated gene exhibits genetic heterogeneity, then the gene is assigned a score of 0. If the disease corresponding to the gene does not exhibit genetic heterogeneity, and the inheritance pattern of the disease corresponding to the gene is dominant: if all the standardized terms of the subject's phenotype are included in the standardized terms of the phenotype corresponding to the gene, then the gene is assigned 2 points; if some of the standardized terms of the subject's phenotype are included in the standardized terms of the phenotype corresponding to the gene, then the gene is assigned 1 point; if there is no overlap between the standardized terms of the subject's phenotype and the standardized terms of the phenotype corresponding to the gene, then the similarity between the subject's phenotype and the standardized terms of the phenotype corresponding to the gene where the subject's gene mutation is located is compared, and if the similarity is greater than a predetermined value, then the gene is assigned 1 point. If the disease corresponding to the gene does not exhibit genetic heterogeneity, and the inheritance pattern of the disease corresponding to the gene is recessive: if all the standardized terms of the subject's phenotype are included in the standardized terms of the phenotype corresponding to the gene, then the gene is assigned a score of 1; if some of the standardized terms of the subject's phenotype are included in the standardized terms of the phenotype corresponding to the gene, then the gene is assigned a score of 0.5; if there is no overlap between the standardized terms of the subject's phenotype and the standardized terms of the phenotype corresponding to the gene, then the similarity between the subject's phenotype and the standardized terms of the phenotype corresponding to the gene where the subject's gene mutation is located is compared. If the similarity is greater than a first predetermined value, then the gene is assigned a score of 0.

5.

9. A device for analyzing gene mutations based on phenotypic characteristics, characterized in that, The method includes a memory and a processor, the memory being used to store a program, and the processor executing the program stored in the memory to implement the method for gene mutation analysis based on phenotypic analysis as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The storage medium stores a program that can be executed by a processor to implement the method for gene mutation based on phenotypic analysis as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Sorting method and device for genetic mutations

    CN108710781A

  • Gene variation and phenotype information association analysis method and system

    CN116612813A