Scoring and sorting method and system for pathogenic mutation of single-gene genetic disease

By extracting standardized phenotypic terms using a hybrid strategy and deep learning model, and combining gene variant annotation and semantic similarity calculation, efficient and automated sorting of pathogenic mutations in single-gene hereditary diseases is achieved. This solves the problems of low automation and low recall in existing technologies and improves diagnostic accuracy.

CN120977385AActive Publication Date: 2025-11-18HANGZHOU BOSHENG BIOTECHNOLOGY CO LTD +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511491841.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-20
Publication Date
2025-11-18
Estimated Expiration
2045-10-20

AI Technical Summary

Technical Problem

Current technologies have low levels of automation in the diagnosis of single-gene genetic diseases, lack the ability to analyze patients' unstructured clinical phenotypic information, and the scoring system fails to fully integrate the correlation between phenotype and gene, resulting in a low recall rate for variant screening.

Method used

A hybrid strategy combining dictionary methods and deep learning models is adopted to extract standardized phenotypic terms. Through the OWL similarity semantics algorithm and gene variation annotation, a supervised learning model is used to perform multi-dimensional pathogenicity comprehensive scoring, and a scoring and ranking system for pathogenic mutations of single-gene hereditary diseases is constructed.

Benefits of technology

It improves the recall rate and automation of detecting pathogenic mutations in genetic diseases, and enhances the accuracy and efficiency of diagnosing single-gene genetic diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120977385A_ABST
    Figure CN120977385A_ABST
Patent Text Reader

Abstract

The invention discloses a scoring and sorting method and system for pathogenic mutation of a single-gene genetic disease, and relates to the technical field of biomedical treatment, the method comprises the following steps: based on pre-acquired phenotypic data and literature abstracts, performing standardized phenotypic term extraction on pre-acquired symptom description by using a mixed strategy, performing phenotype and gene association degree scoring on the standardized phenotype terms and pre-acquired gene mutation data through a similarity comparison method; performing gene variation annotation on a pre-acquired gene variation database, and performing gene variation scoring on the pre-acquired gene variation data according to a gene variation annotation result; and performing multi-dimensional pathogenicity comprehensive scoring evaluation by using a supervised learning model and a weighted summation method to obtain a pathogenic mutation sorting scoring result. By recommending the most probable pathogenic mutation, the grading and sorting of the pathogenic mutation of the genetic disease have the advantages of high detection recall rate and high automation degree.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of biomedicine, in particular to a scoring and ranking method and system for pathogenic mutations of monogenic genetic diseases. BACKGROUND

[0002] Mendelian monogenic genetic diseases are caused by mutations in a single gene, have a clear pathogenic mechanism, and often involve rare or specific variations in individual patients. It is estimated that there are more than 8000 known monogenic genetic diseases, affecting about 1 / 20 of the global population. With the rapid development of second-generation high-throughput sequencing (NGS) technology, especially the widespread application of whole exome sequencing (WES) in clinical practice, molecular diagnosis of monogenic genetic diseases has become more efficient and feasible. Through WES technology, all the variation information in the exonic region of the patient can be obtained at one time, providing an important basis for the identification of pathogenic mutations.

[0003] In practical applications, WES can generate tens of thousands to hundreds of thousands of genetic variation sites, of which only a few are truly pathogenic mutations (Pathogenic), and the rest are mostly benign (Benign) or variants of uncertain significance (VUS). Therefore, how to quickly and accurately screen out the most likely pathogenic candidate mutations from the vast amount of variation data is one of the core challenges in the current molecular diagnosis of genetic diseases. The current mainstream analysis process usually relies on multiple databases (such as ClinVar, HGMD, COSMIC) and annotation tools (such as ANNOVAR, VEP, Transvar) to annotate the variations, but these methods often lack effective integration of patient phenotype information, resulting in low recall rate and easy omission of phenotype-related but not yet fully reported pathogenic mutations.

[0004] On the other hand, the clinical symptoms of patients are often recorded in a free text manner, which is not uniform in description and lacks standardization, and is difficult to be directly used for automatic analysis of computer programs. At present, the Human Phenotype Ontology (HPO) and the China Human Phenotype Ontology (CHPO) are gradually popularized as standardized terminology systems in the world, so as to standardize and structure the phenotype description, and better compare with the known gene-phenotype association database. Some data analysis institutions have introduced HPO into their systems, but it is difficult to link with the clinic due to lack of Chinese version. Non-clinical professionals are prone to bias in the process of converting HPO by secondary analysis of natural language phenotype information provided by doctors. In addition, it is still a challenge to automatically extract accurate HPO terms from unstructured text in practical application, and most processes still rely on manual annotation and review, which is time-consuming and prone to errors.

[0005] At present, some tools such as Phenolyzer, Phenomizer and Exomiser can screen candidate pathogenic genes on the basis of standardized phenotypes. These tools generally judge the matching degree of patient phenotypes and known diseases based on the semantic similarity calculation between HPO terms, and realize the similarity evaluation between HPOs through MICA (Maximum Information Common Ancestor) algorithm. However, the accuracy of such methods is limited by the hierarchical organization of HPO terms in anatomical structure, and cannot fully reflect the relevance of phenotypes in disease mechanism. For example, two phenotypes that are structurally similar (such as tricuspid valve anomaly HP:0001702 and tricuspid valve prolapse HP:0001704) may have a high semantic similarity, but their corresponding pathogenic genes are completely different, which may lead to an increase in false positives and false negatives.

[0006] In summary, the existing genetic disease diagnosis systems based on whole exome sequencing generally have the following problems: (1) The variant screening process relies on manual interpretation, and the degree of automation is low; (2) There is a lack of automatic analysis capability for patient unstructured clinical phenotype information; (3) The scoring system fails to comprehensively integrate the association of phenotype and gene, harmfulness of variation, frequency of occurrence, literature support and other multi-dimensional evidence. Therefore, there is an urgent need for a high-automation scoring and sorting method and system that can integrate and analyze clinical phenotypes and NGS variation data, which can combine phenotype matching, variation annotation, database information and other factors to automatically identify the most likely pathogenic mutations, so as to improve the diagnosis efficiency and accuracy of single gene genetic diseases.

[0007] In view of the problems in the related art, no effective solution has been proposed so far. SUMMARY

[0008] In view of the problems in the related art, the present application provides a scoring and ranking method and system for pathogenic mutations of monogenic genetic diseases to overcome the above technical problems existing in the prior art.

[0009] To this end, the present application adopts the following specific technical solutions: According to an aspect of the present application, a scoring and ranking method for pathogenic mutations of monogenic genetic diseases is provided, which comprises the following steps: S1, based on the pre-acquired phenotype data and literature abstracts, using a hybrid strategy to extract standardized phenotype terms from the pre-acquired symptom descriptions, and using a similarity comparison method to score the association degree between the standardized phenotype terms and the pre-acquired gene mutation data; S2, annotating the pre-acquired gene variation database, and scoring the pre-acquired gene mutation data based on the gene variation annotation results; S3, based on the scoring results of the association degree between the phenotype and the gene and the gene variation scoring results, using a supervised learning model and a weighted summation method to perform multi-dimensional pathogenicity comprehensive scoring and evaluation on the pre-acquired gene mutation data, and obtaining the pathogenic mutation ranking and scoring results.

[0010] Further, based on the pre-acquired phenotype data and literature abstracts, using a hybrid strategy to extract standardized phenotype terms from the pre-acquired symptom descriptions, and using a similarity comparison method to score the association degree between the standardized phenotype terms and the pre-acquired gene mutation data includes: S11, based on the pre-acquired phenotype data and literature abstracts, using a dictionary method and a deep learning model to extract phenotype terms from the pre-acquired symptom descriptions, obtaining dictionary matching results and phenotype term recognition results; S12, using a weighted summation method to integrate the dictionary matching results and the phenotype term recognition results to obtain standardized phenotype terms; S13, using an OWL similarity semantic algorithm to compare the semantic similarity between the standardized phenotype terms and the pre-acquired gene mutation data, and scoring the association degree between the phenotype and the gene based on the similarity comparison results, to obtain the association degree scoring results between the phenotype and the gene.

[0011] Further, based on the pre-acquired phenotype data and literature abstracts, using a dictionary method and a deep learning model to extract phenotype terms from the pre-acquired symptom descriptions, obtaining dictionary matching results and phenotype term recognition results includes: S111, based on the pre-acquired phenotype data, using a dictionary method to match phenotype terms from the clinical symptom descriptions, to obtain dictionary matching results; S112, phenotype term matching of the pre-acquired literature abstracts and the pre-acquired phenotype data is performed on positive and negative sample divisions, and a supervised training set is constructed based on the positive and negative sample division results; S113, the supervised training set is used to train a deep learning model, and a phenotype named entity recognition is performed on the clinical symptom description based on the trained deep learning model and a preset probability threshold, to obtain a phenotype recognition result.

[0012] Further, the standardized phenotype terms and the pre-acquired gene mutation data are compared in semantic similarity by an OWL similarity semantic algorithm, and a phenotype-gene association degree scoring is performed based on the similarity comparison result, to obtain a phenotype-gene association degree scoring result, including: S131, the standardized phenotype terms and the pre-acquired gene mutation data are compared in semantic similarity by an OWL similarity semantic algorithm, and a phenotype-gene association degree scoring is performed based on the similarity comparison result, to obtain a phenotype-gene association degree scoring result, including: S132, the appearance frequency of each phenotype and gene is counted according to the phenotype-gene association set, and the most specific common ancestor is found based on the appearance frequency of each phenotype and gene. S133, the Resnik similarity is used to combine the most specific common ancestor for semantic similarity comparison, and a phenotype-gene association degree scoring is performed based on the similarity comparison result, to obtain a phenotype-gene association degree scoring result.

[0013] Further, the pre-acquired gene variation database is annotated for gene variation, and the pre-acquired gene mutation data is scored for gene variation according to the gene variation annotation result, including: S21, the pre-acquired gene variation database is annotated for gene variation, to obtain a gene variation annotation result. The pre-acquired gene variation database includes: a variation harmfulness database, a population frequency database, a human gene mutation database, a clinical variation database, and a variation classification system database. The gene variation annotation result includes: a variation harmfulness annotation result, a gene frequency annotation result, and a variation classification annotation result. S22, the pre-acquired gene mutation data is scored for gene variation according to the gene variation annotation result by using a weighted summation method and a preset scoring rule, to obtain a gene variation scoring result.

[0014] Further, the gene variation scoring result includes: a variation harmfulness scoring result, a variation occurrence rate scoring result, a variation report scoring result, and a variation pathogenicity scoring result.

[0015] Furthermore, based on the phenotypic-gene association score and the gene variation score, a supervised learning model and a weighted summation method are used to perform a multi-dimensional pathogenicity comprehensive scoring evaluation on the pre-acquired gene mutation data, resulting in a pathogenic mutation ranking score, including: S31. Based on the scoring results of the degree of association between phenotype and gene and the scoring results of gene variation, the weights of the comprehensive scoring assessment of multi-dimensional pathogenicity are set. S32. Utilize the weight training mechanism of ranking learning combined with pre-acquired gene mutation data to optimize the set weights and obtain the weight optimization results. S33. Based on the weight optimization results, the pre-acquired gene mutation data is evaluated using a weighted summation method to achieve a multi-dimensional comprehensive pathogenicity score, and the pathogenic mutation ranking score is obtained.

[0016] Furthermore, the weights are optimized using a ranking learning-based weight training mechanism combined with pre-acquired gene mutation data. The optimized weights include: S321. Training sample pairs are constructed using pre-acquired gene mutation data to obtain several positive and negative sample pairs with pathogenic sites. S322. Construct a pairwise ranking loss function using positive and negative sample pairs with pathogenic sites, and use the pairwise ranking loss function as the training target of the weight training mechanism for ranking learning. S323. Based on the pairwise sorting loss function, the optimizer is used to optimize and iteratively update the set weights to obtain the weight optimization results.

[0017] Furthermore, the expression for the pairwise sorting loss function is as follows: ; In the formula, This represents the pairwise sorting loss function; Indicates the first i Pathogenic variants in each sample; It is the first i The overall score of pathogenic variants in each sample; It is the first i Benign variant scores paired with pathogenic variants in each sample; Indicates a benign variant paired with a pathogenic variant; β The hyperparameters that indicate the steepness of the loss surface; N This represents the total number of positive and negative sample pairs.

[0018] According to another aspect of the present invention, a scoring and ranking system for pathogenic mutations in single-gene hereditary diseases is provided. The scoring and ranking system for pathogenic mutations in single-gene hereditary diseases includes: a phenotype-gene association scoring module, a gene variation scoring module, and a pathogenicity comprehensive assessment module. The phenotype and gene association scoring module is used for scoring the association degree between the standardized phenotype terms and the pre-obtained gene mutation data by using a similarity comparison method based on the pre-obtained phenotype data and literature abstracts, and extracting the standardized phenotype terms from the pre-obtained symptom description by using a hybrid strategy. The gene variation scoring module is used for annotating the pre-obtained gene variation database, and scoring the pre-obtained gene mutation data according to the gene variation annotation result. The pathogenicity comprehensive evaluation module is used for comprehensively scoring and evaluating the pathogenicity of the pre-obtained gene mutation data by using a supervised learning model and a weighted summation method according to the scoring result of the association degree between the phenotype and the gene and the scoring result of the gene variation, to obtain a scoring result of pathogenic mutation ranking.

[0019] The present application has the following advantages: The present application extracts standardized phenotype terms (HPO, human phenotype ontology) from pre-obtained symptom descriptions, combines multi-dimensional annotation and scoring of gene variations, and recommends the most likely pathogenic mutations from massive genetic disease detection variation data (i.e., pre-obtained gene mutation data), so that the scoring and ranking of pathogenic mutations of genetic diseases have the advantages of high detection recall rate and high automation degree, and solve the problems of single consideration factor, low automation degree and poor recall rate in current clinical methods. BRIEF DESCRIPTION OF DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0021] Figure 1 is a flowchart of a single gene genetic disease pathogenic mutation scoring and ranking method according to an embodiment of the present application; Figure 2 is a principle block diagram of a single gene genetic disease pathogenic mutation scoring and ranking system according to an embodiment of the present application; Figure 3 is a deep learning training process flowchart of the association degree between the phenotype and the gene in a single gene genetic disease pathogenic mutation scoring and ranking method according to an embodiment of the present application; Figure 4 is a schematic diagram of identifying standardized HPO phenotype terms from unstructured clinical symptom descriptions in a single gene genetic disease pathogenic mutation scoring and ranking method according to an embodiment of the present application; Figure 5This is a schematic diagram of the scoring principle in a scoring and ranking method for pathogenic mutations in a single-gene hereditary disease according to an embodiment of the present invention.

[0022] In the picture: 1. Phenotypic and gene association scoring module; 2. Gene variation scoring module; 3. Pathogenicity comprehensive assessment module. Detailed Implementation

[0023] To further illustrate the various embodiments, the present invention provides accompanying drawings, which are part of the disclosure of the present invention. These drawings are mainly used to illustrate the embodiments and can be used in conjunction with the relevant descriptions in the specification to explain the operating principles of the embodiments. With reference to these drawings, those skilled in the art should be able to understand other possible implementation methods and the advantages of the present invention.

[0024] According to embodiments of the present invention, a scoring and ranking method and system for pathogenic mutations in single-gene hereditary diseases are provided.

[0025] The present invention will now be further described in conjunction with the accompanying drawings and specific embodiments, such as... Figure 1 As shown, according to an embodiment of the present invention, a scoring and ranking method for pathogenic mutations in single-gene hereditary diseases is provided, the method comprising the following steps: S1. Based on pre-acquired phenotypic data and literature abstracts, a hybrid strategy is used to extract standardized phenotypic terms from the pre-acquired symptom descriptions, and a similarity comparison method is used to score the degree of phenotypic-gene association between the standardized phenotypic terms and the pre-acquired gene mutation data.

[0026] Specifically, based on pre-acquired phenotypic data and literature abstracts, a hybrid strategy is used to extract standardized phenotypic terms from the pre-acquired symptom descriptions. Then, a similarity comparison method is used to score the degree of phenotypic-gene association between the standardized phenotypic terms and the pre-acquired gene mutation data, including: S11. Based on the pre-acquired phenotypic data and literature abstracts, phenotypic terms are extracted from the pre-acquired symptom descriptions using dictionary methods and deep learning models to obtain dictionary matching results and phenotypic term recognition results.

[0027] Specifically, based on pre-acquired phenotypic data and literature abstracts, phenotypic terms are extracted from pre-acquired symptom descriptions using dictionary methods and deep learning models. The dictionary matching results and phenotypic term recognition results include: S111. Based on the pre-acquired phenotypic data, the dictionary method is used to perform phenotypic term matching on the description of clinical symptoms to obtain dictionary matching results. S112. Perform positive and negative sample division by matching the pre-acquired literature abstracts with the pre-acquired phenotypic data using phenotypic terms, and construct a supervised training set based on the positive and negative sample division results. S113, train the deep learning model using the supervised training set, and perform phenotype named entity recognition on the clinical symptom description based on the trained deep learning model and a preset probability threshold to obtain a phenotype recognition result.

[0028] S12, integrate the phenotype terms by using a weighted summation method on the dictionary matching result and the phenotype term recognition result to obtain standardized phenotype terms.

[0029] S13, perform semantic similarity comparison on the standardized phenotype terms and the pre-acquired gene mutation data by using an OWL similarity semantic algorithm, and score the association degree between the phenotype and the gene based on the similarity comparison result to obtain a scoring result of the association degree between the phenotype and the gene.

[0030] Specifically, the semantic similarity comparison on the standardized phenotype terms and the pre-acquired gene mutation data is performed by using an OWL similarity semantic algorithm, and the scoring of the association degree between the phenotype and the gene is performed based on the similarity comparison result to obtain a scoring result of the association degree between the phenotype and the gene, and the scoring result of the association degree between the phenotype and the gene includes: S131, construct a semantic association network of the standardized phenotype terms and the pre-acquired gene mutation data by using an OWL similarity semantic algorithm to obtain a phenotype and gene association set; S132, according to the phenotype and gene association set, statistics the occurrence frequency of each phenotype and gene, and find the most specific common ancestor based on the occurrence frequency of each phenotype and gene; S133, perform semantic similarity comparison by using a Resnik similarity combined with the most specific common ancestor, and score the association degree between the phenotype and the gene based on the similarity comparison result to obtain a scoring result of the association degree between the phenotype and the gene.

[0031] Specifically, in actual application, the whole exome sequencing (WES) experimental procedure and library on-machine sequencing include the following steps: (a) sample collection and DNA extraction, specifically, collecting peripheral blood (EDTA tube), oral swab, amniotic fluid or tissue sample, using Eppendorf centrifuge (5810R and 5427R, Germany); centrifuging at 4°C for 10 min at 1600g, taking the supernatant, centrifuging at 16000g for 10 min, and then taking the supernatant, i.e. blood plasma. gDNA extraction, specifically, using TIANamp Blood DNA Kit (TIANGEN) or QIAamp DNA Blood Mini Kit (QIAGEN) for conventional extraction. (b) DNA quality and concentration detection. Concentration detection, specifically, using Qubit 3.0 fluorometer (Thermo, USA). Fragment size detection, specifically, using Nanodrop or Agilent TapeStation / Fragment Analyzer to detect integrity. (c) WES library construction, specifically, library construction kit, i.e. Twist Bioscience Human Core Exome Kit. DNA starting amount, specifically, gDNA starting amount ≥ 100-200 ng. End repair & A tail addition (End Repair / A-Tailing); ligation of adapters (Illumina double-end adapters); library purification using AMPure XP beads (Beckman Coulter); exon capture, specifically, hybridization capture after mixing the library with probes; target region enrichment after hybridization; PCR amplification, specifically, amplifying the library to obtain sufficient product. Library quality control, specifically, KAPA Library Quant Kit (Roche) for qPCR quantification; library fragment size, specifically, Fragment Analyzer (Agilent) for detection. (d) Sequencing platform, specifically, Illumina NovaSeq 6000.

[0032] Specifically, in practical applications, the basic analysis of whole exome sequencing (WES) data includes the following steps: (a) the sequencer data is converted from BCL files to FASTQ files using bcl2fastq software; (b) low-quality reads, reads containing more than 5% N bases, and reads shorter than 50 bp are removed from the sequencing data using cutadapt software; (c) the above reads are aligned to the hg19 reference genome using bwa software to remove redundant PCR sequences; and (d) DNA fragments with low alignment quality, those that are not aligned, and those whose paired-end reads are not perfectly paired are removed using samtools. (e) Sort the filtered DNA fragments according to their alignment positions; (f) Use the Recalibration algorithm from the GATK software package to correct for instrument biases and errors specific to the Illumina sequencer; (g) Realign INDEL regions: For hotspot regions with frequent insertion and deletion mutations, use the IndelRealigner algorithm from the GATK software package to construct debrujin maps, perform local assembly, and reduce false positive mutations; (c) Use the GATK HaplotypeCaller algorithm to detect SNV and INDEL variants and generate VCF files.

[0033] Specifically, to address the issues that most current HPO-based phenotypic recognition methods rely on: dictionary-based methods—which have high precision but low recall; or supervised machine learning models—which require a large amount of manually labeled data and are difficult to scale, this invention utilizes a hybrid approach, combining dictionary methods with unsupervised / weakly supervised deep learning models to construct an automated HPO concept recognition system that can achieve efficient phenotypic extraction without manually labeled training sets.

[0034] Specifically, first, all HPO terms and their synonyms and definitions are extracted; ambiguous abbreviations (such as ASD, which may stand for Atrial Septal Defect or Autism Spectrum Disorder) are filtered out. Second, a remote supervised training set is constructed by matching PubMed abstracts (i.e., document abstracts), totaling 27 million, with HPO terms to automatically generate positive samples; unmatched n-grams are randomly sampled as negative samples; BioBERT is used for multi-class classification, with n-grams as input and HPO IDs as output. A Trie tree structure is used to accelerate the search process. Third, the BioBERT model is trained using a pre-trained BioBERT model (trained on PubMed); each candidate phrase is classified; a probability threshold is set to filter the final recognition results. Fourth, the dictionary and deep learning results are integrated and merged; a POS label filter is introduced to remove meaningless phrases; overlapping concepts are handled to improve recognition coverage.

[0035] The scoring of phenotypic-gene association combines dictionary methods with unsupervised / weakly supervised deep learning methods for efficient phenotypic extraction. It also assesses the degree of matching between patient phenotypes and known diseases based on semantic similarity calculations between HPO terms, and uses algorithms such as MICA (Maximum Information Common Ancestor) to evaluate the similarity between HPOs. The dictionary-based method is calculated as follows: ; In the formula, h i This refers to a standardized phenotypic term in HPO; Match ( h i , T () indicates that the term was successfully matched precisely or fuzzily in the text; T This represents the input text. The BioBERT-based deep learning method can be quantified as follows: ; In the formula, w 1... w n It is a sequence of words in the text; Given a sequence of context words, a certain HPO concept h j The probability is calculated, and then BioBERT is used to classify each candidate n-gram, outputting the probability score for each HPO concept. Finally, the combination method is as follows: for any two candidate HPO concepts... h a and h bIf they overlap in the text (e.g., share vocabulary), the following rule applies: Specifically, if... h a and h b No overlap → Keep both; if h a and h b Shared ID → Retain the highest scorer; if h a and h b Same start and end positions but different IDs → retain the highest scorer; if h a and h b Different start and end positions with different IDs → both are retained. The final comprehensive score can be quantified as follows: ; In the formula, S dict ( h j () is the matching score of the dictionary method for the HPO concept. h j (0 or 1); S BioBERT ( h j ) is the softmax probability value output by the BioBERT model; α ∈[0,1] is a hyperparameter that balances the weights of the two models.

[0036] See the detailed HPO Normalized Named Entity (NER) extraction flowchart. Figure 3 . Figure 3 The HPO dictionary was constructed by extracting term names and synonyms from the HPO dictionary. The term names and synonyms were also constructed by combining the patient clinical description dictionary. The BERT deep learning model was trained using PubMed literature and phenotype-NER was performed using the patient clinical description to obtain the deep learning results. The HPO phenotypes were merged based on the dictionary results and the deep learning results to obtain the standardized HPO.

[0037] A schematic diagram illustrating the extraction of HPO from a real-world case using a doctor's original, unstructured clinical description for diagnosis is shown below. Figure 4 . Figure 4The clinical description is bilateral hereditary retinal degeneration. Vision loss appeared in both eyes 8 years prior without any obvious cause, with loss of foveal reflex and patchy atrophy; extensive peripheral retinal pigmentation was also observed. The chief complaint was decreased white blood cell and platelet counts for 7-8 years. This includes both leukopenia and thrombocytopenia. The decrease in white blood cells and platelets is the primary concern.

[0038] The extracted standardized phenotypic terms are: retinal degeneration, HP:0000546; vision loss, HP:0000572; extensive peripheral retinal pigmentation, HP:0001106; leukopenia, HP:0001882; thrombocytopenia, HP:0001873; thrombocytopenia, HP:0001873.

[0039] Finally, by combining the patient's standardized HPO ID and VCF mutation file, a weighted score was calculated to determine the association between any gene mutation and the patient's clinical phenotype.

[0040] The semantic similarity between the HPO term list and the known HPO terms for each disease or gene in the database was compared using the OWLsim semantic algorithm, which can be quantified as follows: ; In the formula, H patient This represents the set of HPOs for a given patient, and the set of HPOs associated with each gene-related disease (via OMIM / Orphanet). H gene Using Resnik similarity + MICA (Maximum Information Content Ancestor) as the core, for any two HPO terms h i ∈ H patient and h j ∈ H gene Their similarity is the amount of information about their most specific common ancestor, quantified as follows: ; In the formula, P ( h ) is an HPO term h The frequency of occurrence in the entire disease database. The final calculated phenotype score ranges from 0 to 1, with higher scores indicating a closer match between the patient's phenotype and the disease associated with that gene.

[0041] S2. Annotate the gene variation in the pre-acquired gene variation database, and score the gene variation in the pre-acquired gene mutation data based on the gene variation annotation results.

[0042] Specifically, the process involves annotating pre-acquired gene variation databases and scoring pre-acquired gene mutation data based on the annotation results, including: S21. Perform gene variation annotation on the pre-acquired gene variation database to obtain gene variation annotation results; The pre-acquired gene variant databases include: variant harmfulness database, population frequency database, human gene mutation database, clinical variant database, and variant grading system database; Gene variation annotation results include: variation harmfulness annotation results, gene frequency annotation results, and variation classification annotation results. S22. Based on the gene variation annotation results, use the weighted summation method and preset scoring rules to score the gene variation in the pre-acquired gene variation data and obtain the gene variation scoring results.

[0043] Specifically, the gene mutation scoring results include: mutation harmfulness scoring results, mutation incidence scoring results, mutation reporting scoring results, and mutation pathogenicity scoring results.

[0044] Specifically, in the field of NGS (Next Generation Sequencing), determining whether a mutation is deleterious or pathogenic is a crucial step in variant annotation. Defect-based harmfulness prediction software primarily assesses the potential impact of mutations on protein function through sequence conservation, structural effects, evolutionary information, or machine learning methods. The following are commonly used harmfulness prediction tools and databases widely used in research and clinical practice: SIFT (Sorting Intolerant From Tolerant) determines whether a mutation affects protein function based on sequence homology and amino acid conservation; PolyPhen-2 (Polymorphism Phenotyping v2) combines sequence, structure, and protein functional domain information to predict the impact of amino acid substitutions on protein function; MutationTaster integrates multi-dimensional information such as sequence, conservation, protein function, and splice site alterations for classification; and CADD (Combined Annotation Dependent Depletion) uses machine learning models to comprehensively assess the harmfulness of mutations based on dozens of annotation features (conservation, epigenetic modifications, regulatory information, etc.). REVEL (Rare ExomeVariant Ensemble Learner) integrates results from multiple existing prediction tools (including SIFT, PolyPhen-2, MutationAssessor, etc.) and trains them using an ensemble learning method to derive the final prediction. Alphamissense (developed by DeepMind) predicts the impact of all possible missense variants on protein function based on AlphaFold protein structure and a deep learning model. MetaSVM / MetaLR is based on machine learning (SVM or logistic regression) and integrates the outputs of tools such as SIFT, PolyPhen, LRT, and MutationTaster. SpliceAI is based on a deep neural network and learns from upstream and downstream 10kbp sequences to predict whether variants affect normal splicing (gain / loss). MaxEntScan is based on a maximum entropy model and calculates the probability of base combinations around standard donor / recipient splicing sites (+ / - 3~6bp). dbscSNV (Database of SplicingConsensus SNVs) provides two machine learning models (AdaBoost, RandomForest) to score each splicing-related SNV.

[0045] This invention employs an ensemble-based pathogenicity prediction system, constructing an automated weighted scoring mechanism. By combining the results of multiple existing variant pathogenicity prediction tools, it comprehensively scores each candidate mutation, thereby improving the accuracy of identifying truly pathogenic mutations. This ensemble-based prediction system can serve as an auxiliary decision-making tool in clinical diagnosis, genetic disease screening, and cancer mutation analysis. Specifically, for the pathogenicity prediction of individual variants, this invention constructs a quantitative scoring system based on multiple mainstream variant functional annotation tools. Specifically, it incorporates the outputs of multiple independent variant prediction algorithms or databases, discretizing the results into "harmful" and "harmless" categories. Each score is weighted and summed according to a preset threshold, ultimately generating a comprehensive pathogenicity score for the variant.

[0046] Specifically, variant harmfulness scoring is an ensemble-based pathogenicity prediction method. It constructs an automated weighted scoring system that combines the results of multiple existing variant harmfulness prediction tools to comprehensively score each candidate mutation, thereby improving the accuracy of identifying truly pathogenic mutations. For example, the scoring rules for each tool are as follows: SIFT scores 1 point for a "Damaging" prediction and 0 points for a "Tolerated" prediction; PolyPhen-2 scores 1 point for a "Probably Damaging" prediction and 0 points for a "Possibly Damaging" or "Benign" prediction; REVEL scores 1 point for a score less than 0.5 and 0 points for a score greater than or equal to 0.5; MutationTaster scores 1 point for "Disease causing" and 0 points for "Polymorphism"; CADD (phred-scaled) scores 1 point for "Disease causing" and 0 points for "Polymorphism"; and CADD (phred-scaled) scores 1 point for "Disease causing" and 0 points for "Polymorphism". For the score, 1 point is awarded if it is less than 15, and 0 points are awarded if it is greater than or equal to 15; for AlphaMissense, 1 point is awarded if the score is greater than or equal to 0.8, and 0 points are awarded if it is less than 0.8; for M-CAP, 1 point is awarded if the score is less than 0.025, and 0 points are awarded if it is greater than or equal to 0.025; for MetaSVM and MetaLR, 1 point is awarded if the score is less than 0.5, and 0 points are awarded if it is greater than or equal to 0.5; for SpliceAI, 1 point is awarded if the splicing site prediction score is less than 0.2, and 0 points are awarded otherwise; for MaxEntScan, 1 point is awarded if the model predicts that the variant has a significant impact on the splicing signal, and 0 points are awarded if there is no significant impact; for dbscSNV, 1 point is awarded if the AdaBoost or RandomForest score is greater than or equal to 0.6, and 0 points are awarded otherwise. The results of each of the above tools are converted into a binary score under a unified standard and then summed as a basic indicator for predicting the pathogenicity of variants, which can be further used to rank candidate mutation sites or for joint analysis with phenotypic association scores. This method maintains its biological explanatory power while exhibiting strong scalability and compatibility. The calculation method for the harmfulness score of the variant is as follows: ; In the formula, k i Assign weights to each tool (e.g., SIFT: 1.0, REVEL: 1.2, AlphaMissense: 1.5). f i ( xThe output is a binary number (0 or 1); the final score can be a floating-point number. Specifically, Total_Score ≥ 2.5: highly suspicious pathogenic mutation, Total_Score 1.5~2.5: moderate risk mutation, Total_Score < 1.5: possibly harmless mutation. Therefore, it can be used to quantify the harmfulness of all detected mutations for subsequent screening.

[0047] In NGS analysis, assessing the population frequency of gene variants is one of the important bases for identifying potential pathogenic variants. The variant frequency analysis method used in this system is based on multiple authoritative public databases. These databases integrate large-scale genomic data from different ethnic groups around the world and are widely used in genetic disease research, cancer mutation analysis, and clinical molecular diagnostics. The main population frequency databases and their functions used in this invention include gnomAD (Genome Aggregation Database), maintained by the Broad Institute, which integrates more than 150,000 whole-exome and whole-genome sequencing data, covering multiple ethnic groups worldwide. The 1000 Genomes Project, as one of the earliest large-scale population genome databases, contains WGS data from approximately 2,500 individuals from 26 populations worldwide. ExAC (Exome Aggregation Consortium), the predecessor of gnomAD, mainly focuses on the statistical analysis of variant frequencies in exon regions. It also includes an internally constructed normal human genome frequency (MAF) database, specifically including MAF information from 3,000 healthy subjects. The mutation incidence rate score is a harmlessness scoring function based on population frequency, used for automated mutation screening. Lower frequency indicates a higher likelihood of pathogenicity and a higher score, while higher frequency suggests a higher likelihood of benign polymorphism and a lower score. Specifically, this invention defines a "Population Frequency Pathogenicity Score" (PFPS), calculated as follows: ; In the formula, f : This represents the maximum population frequency of this variant in databases such as gnomAD and 1000G (e.g., GMAF: Global Minor Allele Frequency). f max The threshold is set at 0.001 (0.1%). A frequency higher than this is considered benign, while a frequency lower than this is considered pathogenic. ε :for, ε =1 e -6 The scoring mechanism of this invention uses a very small constant value to avoid division-by-zero errors. It effectively identifies low-frequency or rare variants, thereby improving the recall and accuracy of pathogenic mutations.

[0048] For variant case reporting scoring (i.e., variant reporting scoring), this invention further incorporates internationally recognized variant pathogenicity databases, including the Human Gene Mutation Database (HGMD), ClinVar, and the Leiden Open Variation Database (LOVD), to assist in determining the pathogenicity probability of candidate variants. These databases provide different levels of pathogenicity classification labels, such as "DM," "FP," and "R" in HGMD, "Pathogenic" and "Benign" in ClinVar, and "DM" and "DP" in LOVD. This invention employs a weighted scoring mechanism to quantify and score the variant pathogenicity labels provided by each database, and calculates the overall pathogenicity score of the variant through a linear combination method. Specifically, in the HGMD database, if a variant is marked as "DM" or "FP", it is considered a clearly pathogenic variant and receives 1 point; if it is "DM" or "R", it receives 0.5 points and 0.2 points respectively. In the ClinVar database, if a variant is marked as "Pathogenic" or "Likely Pathogenic", it receives 1 point and 0.8 points respectively; if it is "Uncertain Significance", it receives 0.2 points; otherwise, no points are awarded. In the LOVD database, if a variant is marked as "DM", it receives 1 point; if it is "DP" or "DFP", it receives 0.1 points; otherwise, no points are awarded. Then, the scores from each database are weighted and summed according to preset weights (HGMD: 0.4, ClinVar: 0.4, LOVD: 0.2) to obtain the Multi-Database Pathogenicity Score (MDPS). Optionally, the calculation method is as follows: ; In the formula, S HGMD , S ClinVar , S LOVD These represent the pathogenicity scores of the variant in these three databases, respectively. a 1, a 2, a 3. The weights set for each database can be adjusted based on its authority and data integrity (optionally, for...). a 1 = 0.4 a 2 = 0.4, a (3=0.2). This scoring mechanism can effectively improve the accuracy and recall of pathogenic mutation identification, and is especially suitable for clinical screening of single-gene hereditary diseases and cancer mutations.

[0049] For pathogenicity scoring, this invention further introduces the internationally recognized classification standard for gene variant pathogenicity—the American College of Medical Genetics and Genomics (ACMG) variant grading system—to automatically score and rank the pathogenicity of candidate variants. The ACMG standard provides multi-dimensional pathogenicity evidence items, including pathogenicity support items (such as PM1, PM2, PP1, PP3) and benign support items (such as BA1, BS1, BP4, BP7), and assigns different weights according to the strength of the evidence. Ultimately, ACMG generates a pathogenicity grading result, including: Pathogenic (P)—Definitely pathogenic, indicating that the variant has sufficient evidence to support a clear causal relationship with a certain genetic disease; Likely Pathogenic (LP)—Highly likely to be pathogenic, indicating that the variant has a strong probability of being pathogenic, but has not yet reached the standard of "definitely pathogenic"; Uncertain Significance (VUS)—Uncertain significance, indicating that currently available data are insufficient to determine whether the variant is pathogenic or benign. Possible reasons include insufficient experimental data, uncertain population frequency information, missing pedigree information, or inconsistent results from multiple prediction algorithms. Likely Benign (LB) – This category indicates that the variant is likely not pathogenic, but not entirely ruling out pathogenicity. It is typically based on benign-related ACMG entries (such as BS1, BP1–BP7) or high-frequency observations from large-scale population databases (such as gnomAD, ExAC). While not absolutely benign, the probability of pathogenicity is very low. Benign (B) – This category indicates that the variant has been proven by substantial evidence not to cause disease. It typically occurs at extremely high population frequencies (e.g., >5%), is widespread in healthy individuals, and has been classified as benign polymorphism by multiple studies or expert groups. This invention proposes an automated pathogenicity scoring method based on the ACMG variant pathogenicity grading criteria. By assigning different numerical weights to the ACMG classification labels (such as Pathogenic, Likely Pathogenic, Uncertain Significance, Likely Benign, and Benign) corresponding to candidate variants, a comprehensive pathogenicity score function with strong interpretability and high standardization is constructed. Specifically, if the variant... LA variant classified as pathogenic is assigned the highest pathogenicity score (e.g., 3 points); a variant classified as Likely Pathogenic is assigned the second highest score (e.g., 2 points); a variant classified as Uncertain Significance (VUS) is assigned a neutral score (e.g., 1 point); and a variant classified as Likely Benign or Benign is assigned a lower score or zero score (e.g., 0.5 points and 0 points) based on the strength of its benign evidence, respectively. Through this weighting mechanism, this invention can rapidly and accurately rank the pathogenicity of variants detected by large-scale NGS without human intervention. Combined with annotation information from other dimensions (such as population frequency, splicing effects, and structural prediction), it achieves multidimensional assessment and prioritization of variant pathogenicity. This method significantly improves the automation level and diagnostic efficiency in clinical genetic testing, and is particularly suitable for scenarios involving single-gene diseases, tumor mutation screening, and personalized medicine guidance. Specifically, this invention defines a mapping function. f ( L Each ACMG tag is converted into a real score, representing the pathogenicity probability of that variant. The pathogenicity score is calculated as follows: ; In the formula, P: Pathogenic, LP: Likely Pathogenic, VUS: Variant of UncertainSignificance, LB: Likely Benign, B: Benign.

[0050] S3. Based on the phenotypic and gene association scores and gene variation scores, a multi-dimensional pathogenicity comprehensive scoring evaluation is performed on the pre-acquired gene mutation data using a supervised learning model and a weighted summation method to obtain the pathogenic mutation ranking scores.

[0051] Specifically, based on the phenotypic-gene association score and the gene variation score, a supervised learning model and a weighted summation method are used to perform a multi-dimensional pathogenicity comprehensive scoring evaluation on the pre-acquired gene mutation data, resulting in a pathogenic mutation ranking score, including: S31. Based on the scoring results of the degree of association between phenotype and gene and the scoring results of gene variation, the weights of the comprehensive scoring assessment of multi-dimensional pathogenicity are set. S32. Utilize the weight training mechanism of ranking learning combined with pre-acquired gene mutation data to optimize the set weights and obtain the weight optimization results.

[0052] Specifically, the weights are optimized using a ranking learning-based weight training mechanism combined with pre-acquired gene mutation data. The optimized weights include: S321. Training sample pairs are constructed using pre-acquired gene mutation data to obtain several positive and negative sample pairs with pathogenic sites. S322. Construct a pairwise ranking loss function using positive and negative sample pairs with pathogenic sites, and use the pairwise ranking loss function as the training target of the weight training mechanism for ranking learning. S323. Based on the pairwise sorting loss function, the optimizer is used to optimize and iteratively update the set weights to obtain the weight optimization results.

[0053] Specifically, the expression for the pairwise sorting loss function is: ; In the formula, This represents the pairwise sorting loss function; Indicates the first i Pathogenic variants in each sample; It is the first i The overall score of pathogenic variants in each sample; It is the first i Benign variant scores paired with pathogenic variants in each sample; Indicates a benign variant paired with a pathogenic variant; β The hyperparameters that indicate the steepness of the loss surface; N This represents the total number of positive and negative sample pairs.

[0054] S33. Based on the weight optimization results, the pre-acquired gene mutation data is evaluated using a weighted summation method to achieve a multi-dimensional comprehensive pathogenicity score, and the pathogenic mutation ranking score is obtained.

[0055] Specifically, the pathogenic mutation ranking score output by the weighted score of multi-dimensional pathogenicity assessment (i.e., multi-dimensional pathogenicity comprehensive score assessment) is used to determine the phenotypic association score (i.e., the score of the degree of association between the phenotype and the gene). S phenotype ), the score for assessing the impact of variant function (i.e., the score for the harmfulness of the variant). S functional ), the variation population frequency analysis score (i.e., the obtained variation incidence score) S population ), multi-database pathogenicity score (i.e., the reported variant score) S database ) and ACMG evidence item mapping scores (i.e., the obtained pathogenicity score of the variant) S acmgPhenotypic association scoring involves extracting standardized Human Phenotypic Ontology (HPO) terms from patients' free text descriptions using hybrid natural language processing technologies such as BERT, and calculating semantic matching scores between these terms and known genetic disease phenotypes to quantify the clinical relevance between variants and phenotypes. Variant functional impact assessment involves comprehensively judging the functional impact of candidate variants based on multiple existing predictions (such as SIFT, PolyPhen-2, REVEL, AlphaMissense, etc.), using ensemble learning to weighted sum the outputs of each algorithm to form a unified functional harmfulness score. Variant population frequency analysis relies on large-scale population genome databases (such as Gnom). Based on variant frequency information from AD, 1000 Genomes, and ExAC, a logarithmic transformation-based harmfulness scoring function was designed, resulting in higher scores for low-frequency or unreported variants and lower scores for high-frequency benign variants. A multi-database pathogenicity scoring system was implemented, integrating annotation information on whether a variant is included as pathogenic from authoritative databases such as HGMD, ClinVar, and LOVD, assigning different weights based on their inclusion level and performing a linear combination. ACMG evidence item mapping involved converting pathogenicity evidence items (such as PM, PP, BS, and BP) as defined by the Association for Medical Genetics and Genomics (ACMG) into quantitative scores and introducing a weighting mechanism to reflect the strength of evidence for different items. Combining multiple independent scores, covering dimensions such as clinical phenotype matching, variant functional impact, population frequency distribution, database inclusion, and ACMG pathogenicity evidence, a highly interpretable and automated variant pathogenicity identification system was constructed. The outputs of the above five systems were fused through a weighted linear combination to obtain a unified comprehensive variant pathogenicity score. S pathogenicity : ; Comprehensive pathogenicity scores can effectively improve the accuracy of identifying mutations associated with single-gene genetic diseases, and are especially suitable for clinical genetic diagnosis, personalized medicine guidance, and research-level variant screening tasks.

[0056] The weights in the multi-dimensional pathogenicity assessment weighted score ranking of pathogenic mutations are output. w 1, w 2,…, w 5. The settings can be customized according to the specific application scenario, such as using manual experience-based assignment or optimized training through supervised learning models. Specifically, a weight training mechanism based on Learning to Rank (LTR) is provided to optimize the various weights in the comprehensive pathogenicity scoring function (such as phenotypic association score, variant functional impact score, variant population frequency score, multi-database pathogenicity score, ACMG evidence item mapping score, etc., i.e., each weight). w 1,w 2,…, w 5) The aim is to use supervised learning to ensure that known pathogenic variants are ranked first or as high as possible in the candidate variant list. Specifically, the training process includes the following steps: 1) Constructing training sample pairs: for each patient sample, all candidate variants are extracted, and pathogenic variants are paired with other benign variants to form multiple positive and negative sample pairs; 2) Defining a scoring function: [The text abruptly ends here, so the translation stops as well.] V The overall score is: ; 3) Define the ranking loss function, that is, use pairwise ranking loss as the training objective. Specifically, for each sample, there is one pathogenic site and multiple benign / non-pathogenic variants. For each pair (pathogenic site...) V p , benign site V q We hope that: S ( V p > S ( V q Therefore, a loss function can be constructed, the formula of which is: ; In the formula, It is the first i Pathogenic variants in each sample; It is the first i The overall score of pathogenic variants in each sample; It is the first i Benign variant scores paired with pathogenic variants in each sample; It is a benign variant that is paired with it; β It is a hyperparameter that controls the steepness of the loss surface (specifically, it is 1 in practical applications). N This represents the total number of positive and negative sample pairs. This loss function encourages pathogenic variants to score higher than benign variants. If the score of a pathogenic variant is greater than that of a benign variant, the loss approaches 0; otherwise, the loss increases. 4) Weight optimization and iterative update: using the Adam optimizer, the weights of each module are gradually adjusted by minimizing the above ranking loss function, ultimately giving pathogenic variants a higher ranking priority in the overall score. This training mechanism does not require manually labeled training data; it only relies on existing pathogenic variants and benign variants in their context to complete model training, exhibiting high automation and scalability. After loss function optimization, the optimal weight vector is obtained. w =[0.30,0.25,0.15,0.15,0.15], which makes the pathogenic variant recall rate achieve the expected effect.

[0057] Example 1: Specifically, in this invention, test set 1 is used, which consists of 133 cases of hypercholesterolemia with known pathogenic sites and their ordering, 56 cases of tuberous sclerosis with known pathogenic sites and their ordering, and 99 cases of phenylketonuria with known pathogenic sites and their ordering, as shown in Table 1.

[0058] Table 1. Basic information of clinical samples:

[0059] The recall rate of pathogenic variants in test set 1 achieved the expected results, as shown in Table 2.

[0060] Table 2 compares the recall rates of the default weights with those optimized by this invention:

[0061] Wherein, Recall@5 is the proportion of pathogenic variants ranked in the top 5 in each sample; Recall@10 is the proportion ranked in the top 10; Recall@20 is the proportion ranked in the top 20; the 95% CI (confidence interval) is estimated using the Wilson Score method, reflecting statistical uncertainty; the optimized version of this invention represents the optimal weight combination obtained through ranking learning method (i.e., the Pairwise Ranking Loss method mentioned above) (specifically, here...). w =[0.30,0.25,0.15,0.15,0.15]). Experimental results show that using the optimized weight configuration ( w When the inequality is [0.30, 0.25, 0.15, 0.15, 0.15], the system achieves the best performance across all three metrics.

[0062] For example, in phenylketonuria, Recall@20 reached 97.6%, significantly higher than other settings.

[0063] The above results indicate that by setting weight parameters appropriately or through automated training, the accuracy and stability of pathogenicity identification of variants can be effectively improved, especially in large-scale NGS data analysis and clinical genetic diagnosis scenarios.

[0064] See the overall flowchart of the module of this invention. Figure 5 . Figure 5The system extracts standardized phenotypic terms from clinical symptom descriptions, then performs a scoring module based on phenotype and genes; it performs harmfulness annotation using a variant harmfulness database, then performs a scoring module based on variant harmfulness; it performs minor allele frequency annotation using a healthy population frequency database, then performs a scoring module based on variant incidence; it performs minor allele frequency annotation using a diseased population frequency database, then performs a scoring module based on variant case reports; it performs ACMG grading annotation using an ACMG database, then performs a scoring module based on variant pathogenicity; and finally, it outputs a pathogenic mutation ranking score based on a multi-dimensional pathogenicity assessment weighted score from the above modules.

[0065] Example 2: The clinical sample information used in the examples is shown in Table 3.

[0066] Table 3 Basic information of clinical samples:

[0067] Test set 2 was used to compare the recall of the present invention with that of currently used clinical ranking and scoring software such as exomiser and phenolyzer (Recall@5 (95% CI) represents the recall of pathogenic sites within the top 5 and its 95% confidence interval).

[0068] Recall@10 (95% CI) represents the recall rate and its 95% confidence interval for pathogenic sites within the top 10, and Recall@20 (95% CI) represents the recall rate and its 95% confidence interval for pathogenic sites within the top 20. The comparison results are shown in Table 4.

[0069] Table 4 Comparison Results:

[0070] As shown in Table 4, the overall recall rate of this invention is significantly better than that of currently used clinical ranking and scoring software such as Exomiser and Phenoyzer.

[0071] The comprehensive pathogenicity scoring system can effectively improve the accuracy of identifying mutations related to single-gene genetic diseases, and is especially suitable for clinical genetic diagnosis, personalized medicine guidance and research-level variant screening tasks.

[0072] like Figure 2 As shown, according to another embodiment of the present invention, a scoring and ranking system for pathogenic mutations of single-gene hereditary diseases is provided. The scoring and ranking system for pathogenic mutations of single-gene hereditary diseases includes: a phenotype-gene association scoring module 1, a gene variation scoring module 2, and a pathogenicity comprehensive assessment module 3. Phenotypic and gene association scoring module 1 is used to extract standardized phenotypic terms from pre-acquired phenotypic data and literature abstracts using a hybrid strategy, and to score the degree of phenotypic and gene association between the standardized phenotypic terms and pre-acquired gene mutation data using a similarity comparison method. Gene variation scoring module 2 is used to annotate the gene variation in the pre-acquired gene variation database and score the gene variation in the pre-acquired gene mutation data based on the gene variation annotation results. The pathogenicity comprehensive assessment module 3 is used to perform multi-dimensional pathogenicity comprehensive scoring assessment on pre-acquired gene mutation data based on the scoring results of phenotypic and gene association degree and gene variation, and to obtain pathogenic mutation ranking scoring results.

[0073] In summary, this invention utilizes the automatic identification of combinations of the most likely pathogenic mutations in patients in the preparation of products for the detection or auxiliary detection of Mendelian monogenic genetic diseases. The combinations include any one or more of the following six categories, specifically any three, four, five, or six of the six categories: (1) scoring based on the degree of association between phenotype and gene; (2) scoring based on the harmfulness of the variant; (3) scoring based on the incidence of the variant; (4) scoring based on the reported cases of the variant; (5) scoring based on the pathogenicity of the variant; and (6) output of a ranking score of pathogenic mutations based on a weighted score of multidimensional pathogenicity assessment.

[0074] In summary, by utilizing the above-mentioned technical solution of this invention, this invention extracts standardized phenotypic terms (HPO, Human Phenotypic Ontology) from pre-acquired symptom descriptions, combines multi-dimensional annotation and scoring of gene variations, and recommends the most likely pathogenic mutations from massive genetic disease detection variation data (i.e., pre-acquired gene mutation data). This makes the scoring and ranking of pathogenic mutations in genetic diseases have the advantages of high detection recall and high degree of automation, solving the problems of current clinical methods that consider only one factor, have low degree of automation, and poor recall.

[0075] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A scoring and ranking method for pathogenic mutations in single-gene hereditary diseases, characterized in that, The method includes: S1. Based on pre-acquired phenotypic data and literature abstracts, a hybrid strategy is used to extract standardized phenotypic terms from the pre-acquired symptom descriptions, and a similarity comparison method is used to score the degree of phenotypic-gene association between the standardized phenotypic terms and the pre-acquired gene mutation data. S2. Annotate the gene variation in the pre-acquired gene variation database, and score the gene variation in the pre-acquired gene mutation data based on the gene variation annotation results. S3. Based on the phenotypic and gene association scores and gene variation scores, a multi-dimensional pathogenicity comprehensive scoring evaluation is performed on the pre-acquired gene mutation data using a supervised learning model and a weighted summation method to obtain the pathogenic mutation ranking scores.

2. The scoring and ranking method for pathogenic mutations in single-gene hereditary diseases according to claim 1, characterized in that, The process involves extracting standardized phenotypic terms from pre-acquired phenotypic data and literature abstracts using a hybrid strategy, and then scoring the degree of phenotypic-gene association between the standardized phenotypic terms and pre-acquired gene mutation data using a similarity comparison method. S11. Based on the pre-acquired phenotypic data and literature abstracts, phenotypic terms are extracted from the pre-acquired symptom descriptions using dictionary methods and deep learning models to obtain dictionary matching results and phenotypic term recognition results. S12. Use the weighted summation method to integrate the dictionary matching results and phenotypic terminology recognition results to obtain standardized phenotypic terms. S13. The semantic similarity of standardized phenotypic terms and pre-acquired gene mutation data is compared using the OWL similarity semantics algorithm. Based on the similarity comparison results, the degree of association between phenotype and gene is scored to obtain the score of the degree of association between phenotype and gene.

3. The scoring and ranking method for pathogenic mutations in single-gene hereditary diseases according to claim 2, characterized in that, The process involves extracting phenotypic terms from pre-acquired phenotypic data and literature summaries using dictionary methods and deep learning models. The resulting dictionary matching and phenotypic term recognition results include: S111. Based on the pre-acquired phenotypic data, the dictionary method is used to perform phenotypic term matching on the description of clinical symptoms to obtain dictionary matching results. S112. Perform positive and negative sample division by matching the pre-acquired literature abstracts with the pre-acquired phenotypic data using phenotypic terms, and construct a supervised training set based on the positive and negative sample division results. S113. The deep learning model is trained using a supervised training set, and the trained deep learning model is combined with a preset probability threshold to perform phenotypic named entity recognition on clinical symptom descriptions to obtain phenotypic recognition results.

4. The scoring and ranking method for pathogenic mutations in single-gene hereditary diseases according to claim 2, characterized in that, The OWL similarity semantics algorithm is used to compare the semantic similarity between standardized phenotypic terms and pre-acquired gene mutation data, and the degree of association between phenotype and gene is scored based on the similarity comparison results. The phenotypic-gene association score results include: S131. Using the OWL similarity semantics algorithm, construct a semantic association network between standardized phenotypic terms and pre-acquired gene mutation data to obtain a set of phenotypic and gene associations. S132. Calculate the frequency of occurrence of each phenotype and gene based on the phenotype-gene association set, and find the most specific common ancestor based on the frequency of occurrence of each phenotype and gene. S133. Using Resnick similarity combined with the most specific common ancestor, semantic similarity is compared, and the degree of association between phenotype and gene is scored based on the similarity comparison results to obtain the score of the degree of association between phenotype and gene.

5. The scoring and ranking method for pathogenic mutations in single-gene hereditary diseases according to claim 1, characterized in that, The step of annotating the pre-acquired gene variation database and scoring the pre-acquired gene mutation data based on the gene variation annotation results includes: S21. Perform gene variation annotation on the pre-acquired gene variation database to obtain gene variation annotation results; The pre-acquired gene variation database includes: a variation harmfulness database, a population frequency database, a human gene mutation database, a clinical variation database, and a variation grading system database; The gene variation annotation results include: variation harmfulness annotation results, gene frequency annotation results, and variation classification annotation results. S22. Based on the gene variation annotation results, use the weighted summation method and preset scoring rules to score the gene variation in the pre-acquired gene variation data and obtain the gene variation scoring results.

6. The scoring and ranking method for pathogenic mutations in single-gene hereditary diseases according to claim 5, characterized in that, The gene mutation scoring results include: mutation harmfulness scoring results, mutation incidence scoring results, mutation reporting scoring results, and mutation pathogenicity scoring results.

7. The scoring and ranking method for pathogenic mutations in single-gene hereditary diseases according to claim 1, characterized in that, The process involves using a supervised learning model and a weighted summation method to perform a multi-dimensional comprehensive pathogenicity assessment on the pre-acquired gene mutation data based on the scoring results of phenotypic and gene association and gene variation, resulting in a pathogenic mutation ranking score. S31. Based on the scoring results of the degree of association between phenotype and gene and the scoring results of gene variation, the weights of the comprehensive scoring assessment of multi-dimensional pathogenicity are set. S32. Utilize the weight training mechanism of ranking learning combined with pre-acquired gene mutation data to optimize the set weights and obtain the weight optimization results. S33. Based on the weight optimization results, the pre-acquired gene mutation data is evaluated using a weighted summation method to achieve a multi-dimensional comprehensive pathogenicity score, and the pathogenic mutation ranking score is obtained.

8. The scoring and ranking method for pathogenic mutations in single-gene hereditary diseases according to claim 7, characterized in that, The weight training mechanism using ranking learning, combined with pre-acquired gene mutation data, optimizes the set weights to obtain the following weight optimization results: S321. Training sample pairs are constructed using pre-acquired gene mutation data to obtain several positive and negative sample pairs with pathogenic sites. S322. Construct a pairwise ranking loss function using positive and negative sample pairs with pathogenic sites, and use the pairwise ranking loss function as the training target of the weight training mechanism for ranking learning. S323. Based on the pairwise sorting loss function, the optimizer is used to optimize and iteratively update the set weights to obtain the weight optimization results.

9. The scoring and ranking method for pathogenic mutations in single-gene hereditary diseases according to claim 8, characterized in that, The expression for the pairwise sorting loss function is: ; In the formula, This represents the pairwise sorting loss function; Indicates the first i Pathogenic variants in each sample; It is the first i The overall score of pathogenic variants in each sample; It is the first i Benign variant scores paired with pathogenic variants in each sample; Indicates a benign variant paired with a pathogenic variant; β The hyperparameters that indicate the steepness of the loss surface; N This represents the total number of positive and negative sample pairs.

10. A scoring and ranking system for pathogenic mutations in single-gene hereditary diseases, used to implement the scoring and ranking method for pathogenic mutations in single-gene hereditary diseases as described in any one of claims 1-9, characterized in that, The scoring and ranking system for pathogenic mutations in this single-gene hereditary disease includes: a phenotype-gene association scoring module, a gene variation scoring module, and a comprehensive pathogenicity assessment module; The phenotype-gene association scoring module is used to extract standardized phenotypic terms from pre-acquired phenotypic data and literature abstracts using a hybrid strategy, and to score the degree of phenotypic-gene association between the standardized phenotypic terms and pre-acquired gene mutation data using a similarity comparison method. The gene variation scoring module is used to annotate the gene variation in the pre-acquired gene variation database and score the gene variation in the pre-acquired gene mutation data based on the gene variation annotation results. The pathogenicity comprehensive assessment module is used to perform multi-dimensional pathogenicity comprehensive scoring assessment on pre-acquired gene mutation data based on the phenotypic and gene association degree scoring results and gene variation scoring results, using a supervised learning model and weighted summation method, to obtain pathogenic mutation ranking scoring results.

Citation Information

Patent Citations

  • Gene variation scoring and sequencing method for genetic variation analysis

    CN117877578A

  • Covariate correction of time data from phenotypic measurements of different drug usage patterns

    CN118451511A

  • Method, system and device for analyzing gene mutation based on phenotype and computer readable storage medium

    CN119229965A

  • System and method for identifying mutation and phenotype association

    CN119452417A

  • Notification service server capable of providing access notification service to harmful sites and operating method thereof

    KR102421572B1