Methods, devices, equipment, and media for predicting pathogenic genes based on phenotypic fingerprinting
By constructing gene-specific phenotypic fingerprints and multidimensional clinical phenotypic scores, the problem of low efficiency in identifying pathogenic genes for complex diseases in existing technologies has been solved, achieving efficient and objective prediction of pathogenic genes and reducing reliance on expert experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- XIANGYA HOSPITAL CENT SOUTH UNIV
- Filing Date
- 2026-03-16
- Publication Date
- 2026-05-26
AI Technical Summary
In complex disease scenarios, existing technologies struggle to quickly and accurately identify the most likely pathogenic genes from a large number of candidate genes, especially when patients present with a combination of multi-system and multi-dimensional phenotypes. The sheer number of candidate genes, unclear priorities, low interpretation efficiency, and reliance on expert experience for consistent results make it difficult to achieve the desired outcome.
The pathogenic gene prediction method based on phenotypic fingerprints constructs gene-specific phenotypic fingerprints and transforms multidimensional clinical phenotypic information into gene-centered quantitative scores at the object level. This enables the assessment of the risk of carrying pathogenic variants and the prioritization of candidate pathogenic genes, reducing reliance on expert experience.
It significantly improves the efficiency and objectivity of pathogenic gene identification for complex diseases, provides a systematic, quantitative and efficient pathogenic gene prediction scheme, reduces reliance on expert experience, and improves the accuracy and reliability of pathogenic gene identification.
Smart Images

Figure CN121838892B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of biomedical informatics, and in particular to a method, apparatus, device, and medium for predicting pathogenic genes based on phenotypic fingerprinting. Background Technology
[0002] In recent years, the rapid development of high-throughput sequencing technology has led to an exponential increase and accumulation of a large amount of human genetic variation data. For complex diseases (such as neurodevelopmental disorders), the phenotypic profiles corresponding to individual pathogenic genes often exhibit incomplete penetrance, variable expression levels, and significant individual differences. At the same time, different genes have different patterns of influence on different phenotypic dimensions, making it difficult to directly establish an explanatory chain from variation to phenotype.
[0003] Current clinical workflows typically rely on the following methods to screen and rank candidate genes: 1. Initial screening of genes based on existing gene-disease knowledge bases or guidelines; 2. Grading of candidate variants based on evidence of pathogenicity; 3. Determining potential pathogenic genes and forming a diagnostic conclusion by combining patient phenotype with previous reports.
[0004] However, in complex disease scenarios, the above process often faces challenges such as an excessive number of candidate genes, unclear priorities, low interpretation efficiency, and reliance on expert experience for consistent results, due to limitations such as incomplete phenotypic data collection, inconsistent phenotypic descriptions, and insufficient knowledge base coverage. In particular, when a patient presents with a combination of multi-system and multi-dimensional phenotypic traits, identifying the most likely pathogenic gene from a large pool of candidate genes within a short period has become a pressing issue. Summary of the Invention
[0005] This application aims to at least address the technical problems existing in the prior art. To this end, this application proposes a method, apparatus, device, and medium for predicting pathogenic genes based on phenotypic fingerprints. This method constructs gene-specific phenotypic fingerprints at the population level and transforms multidimensional clinical phenotypic information into gene-centered quantitative scores at the object level. By achieving two key outputs—risk assessment of pathogenic variants and priority ranking of candidate pathogenic genes—it can significantly improve the efficiency and objectivity of identifying pathogenic genes in complex diseases, reduce reliance on expert experience, and has significant clinical application value.
[0006] To achieve the above objectives, in a first aspect, this application provides a method for predicting pathogenic genes based on phenotypic fingerprints, the method comprising:
[0007] Based on at least one subject's multiple phenotypes and multiple genes associated with the target disease, multiple gene-phenotype association pairs are constructed, and the association information of each gene across all phenotypic dimensions is determined based on the multiple gene-phenotype association pairs to construct a phenotypic fingerprint for each gene according to the association information; wherein any gene-phenotype association pair consists of any phenotype and any pathogenic gene, and any two gene-phenotype association pairs are different; each gene includes at least one pathogenic variant or potentially pathogenic variant;
[0008] In response to the pathogenic gene prediction instruction of the target object, determine the observed phenotype of the target object associated with the target disease;
[0009] Based on the phenotypic fingerprint and the observed phenotype, a gene-centered multiphenotypic score is calculated for each candidate gene of the target object, in order to predict the risk value of the target object carrying the corresponding pathogenic variant and determine the priority of each candidate gene based on the multiphenotypic score of each candidate gene.
[0010] The pathogenic gene prediction method based on phenotypic fingerprinting provided in the first aspect of this application has at least the following beneficial effects:
[0011] This method constructs a phenotypic fingerprint for each gene, systematically integrating and quantifying the association information between genes and multiple phenotypic dimensions. It enables the construction of gene-specific phenotypic fingerprints at the population level. Unlike existing technologies that simply associate genes with diseases one-to-one, this method more comprehensively captures the complex effects of genes across different phenotypic dimensions, providing more refined and richer background information for subsequent individualized predictions. Furthermore, this method calculates gene-centric multiphenotypic scores through gene phenotypic fingerprints, transforming the multidimensional clinical phenotypic information of the target subject into a gene-centric quantitative score at the object level. It utilizes the observed phenotypes of the target subject to achieve two key outputs: risk assessment of pathogenic variants and priority ranking of candidate pathogenic genes. This data-driven scoring mechanism can objectively and efficiently evaluate the fit between each candidate gene and the actual phenotype of the target subject. By introducing the concept of phenotypic fingerprints and combining them with a multiphenotypic scoring mechanism, this method provides a systematic, quantitative, and efficient pathogenic gene prediction scheme, significantly improving the efficiency and objectivity of pathogenic gene identification for complex diseases, reducing reliance on expert experience, and possessing significant clinical application value.
[0012] To achieve the above objectives, a second aspect of this application provides a pathogenic gene prediction device based on phenotypic fingerprinting, the device comprising:
[0013] The phenotypic fingerprint construction module is used to construct multiple gene-phenotype association pairs based on multiple phenotypes and multiple genes associated with at least one subject and the target disease, and to determine the association information of each gene across all phenotypic dimensions based on the multiple gene-phenotype association pairs, so as to construct a phenotypic fingerprint for each gene according to the association information; wherein any gene-phenotype association pair consists of any phenotype and any pathogenic gene, and any two gene-phenotype association pairs are different; each gene includes at least one pathogenic variant or a potentially pathogenic variant;
[0014] The prediction instruction response module is used to respond to the pathogenic gene prediction instruction of the target object and determine the observed phenotype of the target object related to the target disease.
[0015] The risk prediction and ranking module is used to calculate a gene-centered multiphenotypic score for each candidate gene of the target object based on the phenotypic fingerprint and the observed phenotype, so as to predict the risk value of the target object carrying the corresponding pathogenic variant and determine the priority of each candidate gene based on the multiphenotypic score of each candidate gene.
[0016] To achieve the above objectives, a third aspect of this application provides an electronic device, including at least one control processor and a memory for communicatively connecting to the at least one control processor; the memory stores instructions executable by the at least one control processor, which, when executed by the at least one control processor, enables the at least one control processor to perform the pathogenic gene prediction method based on phenotypic fingerprints described in the first aspect.
[0017] To achieve the above objectives, in a fourth aspect of this application, a computer-readable storage medium is provided, the computer-readable storage medium storing computer-executable instructions, the computer-executable instructions being used to cause a computer to execute the pathogenic gene prediction method based on phenotypic fingerprints described in the first aspect. Attached Figure Description
[0018] Figure 1 This is a flowchart illustrating the pathogenic gene prediction method based on phenotypic fingerprinting provided in this application embodiment;
[0019] Figure 2 This is a schematic diagram of the sample screening process provided in the embodiments of this application;
[0020] Figure 3 This is a schematic diagram of a single high-confidence gene-level assessment provided in an embodiment of this application;
[0021] Figure 4 This is a schematic diagram of a single extended gene-level assessment provided in an embodiment of this application;
[0022] Figure 5 This is a schematic diagram of the multi-gene level assessment provided in the embodiments of this application;
[0023] Figure 6 This is a schematic diagram of the independent external verification queue provided in an embodiment of this application;
[0024] Figure 7 This is a schematic diagram of the structure of the pathogenic gene prediction device based on phenotypic fingerprinting provided in the embodiments of this application;
[0025] Figure 8 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0026] The embodiments of this application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application.
[0027] In the description of this application, the use of terms such as "first," "second," etc., is for the purpose of distinguishing technical features only and should not be construed as indicating or implying relative importance or implicitly indicating the number of technical features indicated or the order of the technical features indicated.
[0028] like Figure 1 As shown in one embodiment of this application, a method for predicting pathogenic genes based on phenotypic fingerprints is provided. This method includes the following steps:
[0029] Step S110: Based on at least one subject's multiple phenotypes and multiple genes associated with the target disease, construct multiple gene-phenotype association pairs, and determine the association information of each gene across all phenotype dimensions based on the multiple gene-phenotype association pairs, so as to construct the phenotypic fingerprint of each gene according to the association information.
[0030] Step S120: In response to the pathogenic gene prediction instruction of the target object, determine the observed phenotype of the target object related to the target disease.
[0031] Step S130: Based on the phenotypic fingerprint and observed phenotype, calculate the gene-centered multiphenotypic score for each candidate gene of the target object, so as to predict the risk value of the target object carrying the corresponding pathogenic variant and determine the priority of each candidate gene based on the multiphenotypic score of each candidate gene.
[0032] To facilitate understanding, some key terms in this embodiment will be explained below:
[0033] The subjects are the individuals from whom the data is used to construct the phenotypic fingerprint. They are typically patients diagnosed with a specific disease or exhibiting a particular disease phenotype, whose genomic and phenotypic data have been collected and analyzed. It is important to note that this is not limited to a specific target disease; it can include neurodevelopmental disorders, mental illnesses, and other multifactorial diseases.
[0034] Phenotype refers to the observable traits of an organism, including morphological, physiological, biochemical, and behavioral characteristics. In this embodiment, phenotype can be manifested as disease symptoms, clinical indicators, imaging features, etc.
[0035] A gene is a segment of DNA that carries genetic information and is the basic unit of heredity in an organism. Genes influence an organism's traits by encoding proteins or other functional molecules. Pathogenic or probable pathogenic variants refer to DNA sequence alterations in the genome that may lead to disease or increase the risk of disease. Pathogenic variants are those that have been definitively proven to be associated with disease, while probable pathogenic variants are those that are highly suspected of being associated with disease based on existing evidence. A detailed classification of pathogenic and probable pathogenic variants in genes will be provided in subsequent examples and will not be elaborated upon here.
[0036] The target population refers to individuals for whom disease-causing gene prediction is required. Their observed phenotypes will be used to assess their risk of carrying disease-causing variants. The observed phenotypes refer to phenotypic data related to the target disease that are actually measured or observed in the target population.
[0037] The multiphenotypic score is a quantitative score calculated by comprehensively considering the association information of a candidate gene across multiple phenotypic dimensions and the observed phenotypes of the target subject. Priority refers to the ranking of multiple candidate genes, reflecting the likelihood of each candidate gene being a pathogenic gene, which helps clinicians quickly focus on the most relevant genes.
[0038] This embodiment provides a method for predicting pathogenic genes based on phenotypic fingerprints, the method including the following main features:
[0039] The core of this application lies in constructing gene-specific phenotypic fingerprints at the population level and transforming multidimensional clinical phenotypic information into gene-centered quantitative scores at the object level, thereby achieving two key outputs: risk assessment of pathogenic variants and priority ranking of candidate pathogenic genes.
[0040] In step S110 of this embodiment, data is first collected from the subjects. The data mainly includes genetic data and phenotypic data. Taking the target disease as autism spectrum disorder (ASD) as an example, phenotypes (such as self-feeding with a spoon, Tourette syndrome, epilepsy, etc.) are collected from multiple ASD patients. It should be noted that phenotypes can be classified into binary phenotypes (such as Tourette syndrome, epilepsy) and continuous phenotypes (such as self-feeding with a spoon). At the same time, corresponding genetic data are collected. It should be noted that the genes here include at least one pathogenic variant or potentially pathogenic variant. The classification of gene variants will be described in detail in subsequent embodiments and will not be elaborated here.
[0041] When constructing multiple gene-phenotype association pairs, one approach is to manually review literature and expert experience to organize known gene-phenotype relationships into gene-phenotype association pairs. For example, for a hereditary disease, gene mutation information and various clinical symptoms exhibited by patients can be manually collected, pairing each gene with each symptom. Another approach is to extract gene and phenotype data from large biomedical databases, perform preliminary screening, and pair each gene with each phenotype to form an initial set of gene-phenotype association pairs.
[0042] When determining the association information of each gene across all phenotypic dimensions and constructing a phenotypic fingerprint, one approach is to use statistical methods, such as calculating the correlation coefficient or performing a chi-square test, for each gene-phenotype association pair to assess the strength and direction of the association between the gene and the phenotype. These statistical results are then directly used as association information and combined to form the gene's phenotypic fingerprint. Another approach is to describe the relationship between each gene and all relevant phenotypes through manual analysis or based on predefined rules, such as "strong association," "weak association," or "no association." These descriptions are then encoded as numerical values or symbols as association information, and used to construct the gene's phenotypic fingerprint.
[0043] In step S120 of this embodiment, when responding to a pathogenic gene prediction instruction for a target object, one possible approach is that the system receives a manually input instruction from a user or clinician, which explicitly specifies the target object for which pathogenic gene prediction needs to be performed. For example, a doctor selects a patient in the diagnostic system and clicks "Start Prediction." Another possible approach is that the system automatically scans the queue of target objects to be processed through a scheduled task or batch processing mechanism. Once a new prediction request is detected, the prediction process is automatically triggered. For example, when new patient genome and phenotype data are uploaded to the database, the system automatically identifies and initiates prediction. When determining the observational phenotypes of a target subject related to a target disease, one possible approach is manual input, where medical staff or data administrators manually enter phenotypic information from paper or electronic documents such as clinical diagnostic reports and physical examination results into the system. For example, doctors input data such as the patient's height, weight, blood pressure, and specific symptoms item by item based on the patient's medical records. Another possible approach is for the system to directly obtain the target subject's structured phenotypic data from an electronic medical record system or laboratory information management system through a preset data interface. For example, the system can automatically extract recorded disease symptoms, biochemical indicators, and other information from the patient's electronic health record.
[0044] In one possible implementation, any gene in step S110 can be used as a candidate gene. Candidate genes can also be genes with multiple undetermined variants detected by sequencing. When calculating the gene-centric multiphenotypic score for each candidate gene of the target object in step S130, one possible approach is to calculate a score for each candidate gene by analyzing the portion of its phenotypic fingerprint corresponding to the observed phenotype of the target object. Another possible approach is to logically compare the association information in the phenotypic fingerprint with the observed phenotype of the target object. For example, if a gene is associated with a certain phenotype and the target object observes the phenotype, a fixed score is assigned; otherwise, no score is added. Then, the scores of all phenotypes are summed to form the multiphenotypic score.
[0045] When predicting the risk of a target subject carrying a corresponding pathogenic variant and determining the priority of each candidate gene based on its multiphenotypic score, one approach is to directly use the calculated multiphenotypic score (0-1) as the risk value and sort the candidate genes according to the score from high to low to determine their priority. For example, the higher the score, the greater the risk and the higher the priority. Another approach is to set a fixed threshold. If the multiphenotypic score of a candidate gene exceeds the threshold, the target subject is considered to have a high risk of carrying the corresponding pathogenic variant and is marked as high priority; otherwise, it is marked as low priority.
[0046] This method constructs a phenotypic fingerprint for each gene, systematically integrating and quantifying the association information between genes and multiple phenotypic dimensions. It enables the construction of gene-specific phenotypic fingerprints at the population level. Unlike existing technologies that simply associate genes with diseases one-to-one, this method more comprehensively captures the complex effects of genes across different phenotypic dimensions, providing more refined and richer background information for subsequent individualized predictions. Furthermore, by calculating gene-centered multiphenotypic scores, this method quantitatively matches the observed phenotype of the target subject with the pre-constructed gene phenotypic fingerprint. In other words, it transforms the multidimensional clinical phenotypic information of the target subject into a gene-centered quantitative score at the subject level, achieving two key outputs: risk assessment of pathogenic variants and priority ranking of candidate pathogenic genes. This data-driven scoring mechanism avoids over-reliance on the experience of a single expert and can objectively and efficiently assess the fit between each candidate gene and the actual phenotype of the target subject. This method introduces the concept of phenotypic fingerprinting and combines it with a multiphenotypic scoring mechanism to provide a systematic, quantitative, and efficient pathogenic gene prediction scheme. It significantly improves the efficiency and objectivity of pathogenic gene identification for complex diseases, reduces reliance on expert experience, and has important clinical application value.
[0047] In some embodiments of this application, step S110, which determines the association information of each gene across all phenotypic dimensions based on multiple gene-phenotype association pairs, to construct a phenotypic fingerprint for each gene based on the association information, includes:
[0048] Step S210: Perform independent association tests on multiple gene-phenotype association pairs to obtain the association test results.
[0049] Step S220: Based on the association test results, draw an association map of multiple gene phenotype association pairs.
[0050] Step S230: Based on the association map, determine the association direction, effect size, and significance information of each gene across all phenotypic dimensions.
[0051] Step S240: Construct the phenotypic fingerprint of each gene based on the association direction, effect size, and significance information.
[0052] In this embodiment, independent association tests are performed on multiple gene-phenotype association pairs to assess whether a statistically significant association exists between each gene and a specific phenotype. By performing separate statistical analyses on each gene-phenotype association pair, the strength and reliability of their correlations can be quantified. For example, a chi-square test can be used to assess the association between categorical variables (such as genes and phenotypes), or logistic regression or linear regression models can be used to analyze the association between gene variation and phenotype based on phenotype type (binary or continuous), while controlling for potential confounding factors.
[0053] Based on the association test results, an association map of multiple gene-phenotype association pairs is created. An association map is a visualization tool used to visually display the association patterns between genes and various phenotypes. It helps researchers quickly identify which genes are significantly associated with which phenotypes, as well as the strength and direction of these associations. For example, association test results (such as p-values, effect sizes, etc.) can be presented as a heatmap, where rows represent genes, columns represent phenotypes, and color intensity indicates association strength; or a network graph can be constructed, where nodes represent genes and phenotypes, edges represent the associations between them, and the thickness or color of the edges reflects the association strength.
[0054] Based on the association map, the association direction, effect size, and significance information of each gene across all phenotypic dimensions are determined. This step extracts the key information needed to construct the phenotypic fingerprint from the association map. The association direction indicates the trend of the gene's influence on the phenotype (e.g., increasing or decreasing a certain phenotypic trait), the effect size quantifies the magnitude of this influence, and the significance information indicates the statistical reliability of this influence. For example, for heatmaps, significant associations can be identified using color gradients and preset thresholds, and the corresponding effect sizes (such as odds ratios and regression coefficients) and p-values can be extracted from the original test results as significance information. For network graphs, the connection strength and direction between nodes can be analyzed, and the effect size and significance p-value of each association can be directly obtained from the statistical report of the association test.
[0055] A phenotypic fingerprint is constructed for each gene based on association direction, effect size, and significance information. A phenotypic fingerprint is a unique, multi-dimensional vector of phenotypic association features for each gene. It integrates the association direction, effect size, and significance between the gene and all related phenotypes, forming a unique signature for subsequent pathogenic gene prediction. For example, the association direction, effect size, and significance information (such as p-values or corrected p-values) of each gene across all phenotypic dimensions can be combined into a vector or matrix as the phenotypic fingerprint of that gene; alternatively, the effect size and significance information can be weighted or transformed, for example, by multiplying the effect size by the inverse or logarithm of the significance level to highlight stronger and more reliable associations, forming a more discriminative phenotypic fingerprint.
[0056] This embodiment can accurately extract biologically significant association information from complex gene and phenotype data. Through independent association tests, this embodiment ensures the statistical reliability of the association between each gene and phenotype, avoiding the interference of false positive associations on subsequent predictions. The creation of association maps provides researchers with an intuitive and comprehensive perspective to quickly understand the complex relationship between genes and multiple phenotypes. Finally, the phenotypic fingerprint constructed by integrating association direction, effect size, and significance information can more accurately capture the unique biological function of each gene and its role in disease occurrence and development, thereby significantly improving the accuracy and reliability of subsequent disease-causing gene prediction.
[0057] In some embodiments of this application, the association test results include the test P-value corresponding to each gene phenotype association pair;
[0058] Step S220 involves drawing an association map of multiple gene-phenotype association pairs based on the association test results, including:
[0059] Step S2210: Based on the test P value corresponding to each gene-phenotype association pair, select multiple first gene-phenotype association pairs from multiple gene-phenotype association pairs whose test P value is less than a first threshold; wherein, the first threshold is used to characterize the association between genes and phenotypes in the gene-phenotype association pair.
[0060] Step S2230: Draw an association map based on multiple first gene phenotype association pairs.
[0061] The association test results include the p-value for each gene-phenotype association pair. In this embodiment, for each gene-phenotype association pair, a p-value is calculated using an independent association test (e.g., chi-square test, t-test, regression analysis, etc.). This p-value reflects the statistical significance of the association between the gene and the phenotype. The smaller the p-value, the less likely the observed association is to be caused by random chance, i.e., the greater the probability of a genuine association between the gene and the phenotype.
[0062] Further, based on the p-value corresponding to each gene-phenotype association pair, multiple first gene-phenotype association pairs with p-values less than a first threshold are selected from the multiple gene-phenotype association pairs. This step aims to identify a subset of gene-phenotype association pairs with statistically significant associations from all gene-phenotype association pairs. By setting a predefined first threshold, an association is considered noteworthy only when the p-value of a gene-phenotype association pair is below this threshold. This screening mechanism can effectively remove gene-phenotype association pairs with strong randomness and weak associations, thereby focusing on more reliable association information. In addition to directly comparing p-values with the threshold, other screening strategies can be used, such as screening based on effect size, or a comprehensive evaluation combining p-values and effect sizes. The first threshold is used to characterize the association between genes and phenotypes in a gene-phenotype association pair. The first threshold is a preset statistical significance level, such as 0.05. When the p-value of a gene-phenotype association pair is less than the first threshold, it is statistically considered that there is a non-random, meaningful association between the gene and the phenotype.
[0063] Based on this, association maps are drawn according to multiple first gene phenotype association pairs. After screening out the first gene phenotype association pairs, only these association pairs that have been filtered for significance are used to construct the association maps. This method ensures that the association information contained in the maps is statistically validated, thereby improving the quality and reliability of the maps. Association maps can be presented in various forms, such as network diagrams, heatmaps, or scatter plots, with the aim of visually demonstrating the interaction patterns between genes and phenotypes.
[0064] In constructing the association map of gene-phenotype association pairs, this embodiment can effectively filter out statistically insignificant gene-phenotype associations, avoiding the introduction of randomness or weak association information into the subsequent phenotypic fingerprint construction process. This significantly improves the purity and reliability of the association map, making the association direction, effect size, and significance information extracted from the map more accurate. Therefore, the phenotypic fingerprint of each gene constructed based on this can more accurately capture the real relationship between the gene and the disease-related phenotype, thereby providing a more solid and reliable data foundation for the subsequent prediction of pathogenic genes, and thus improving the accuracy and credibility of pathogenic gene prediction.
[0065] In some embodiments of this application, before step S2230, which involves drawing an association map based on multiple first gene phenotype association pairs, the method further includes:
[0066] Step S2221: Update the test p-value based on multiple hypothesis testing;
[0067] Step S2222: Based on the updated test P-value, select multiple second gene phenotype association pairs from multiple first gene phenotype association pairs;
[0068] Step S130, which calculates a gene-centered multiphenotypic score for each candidate gene of the target object based on phenotypic fingerprints and observed phenotypes, includes:
[0069] Step S310: Extract the phenotype from the first gene phenotype association pair corresponding to the candidate gene from the phenotype fingerprint, and calculate the signal term value of the candidate gene based on the phenotype fingerprint and the observed phenotype when the observed phenotype corresponds to the phenotype in the extracted first gene phenotype association pair.
[0070] Step S320: Extract the phenotypes from the first gene phenotype association pair corresponding to the candidate gene and the phenotypes from the second gene phenotype association pair corresponding to the candidate gene from the phenotype fingerprint. If the observed phenotype does not correspond to the phenotypes in the first gene phenotype association pair or the phenotypes in the second gene phenotype association pair, calculate the penalty term value of the candidate gene based on the phenotype fingerprint and the observed phenotype.
[0071] Step S330: Calculate the gene-centered multiphenotype score for each candidate gene based on the signal term value and penalty term value of each candidate gene.
[0072] This embodiment aims to further correct for the increased risk of false positives due to numerous statistical tests using more rigorous statistical methods. Multiple hypothesis testing is a statistical procedure used to control the overall error rate when performing multiple statistical tests simultaneously. Its role is to ensure that the selected second gene phenotype association pairs have higher statistical significance and confidence, thereby providing a more solid foundation for subsequent pathogenic gene prediction. One possible approach is to use the False Discovery Rate (FDR) algorithm for correction. This method balances false positive control and statistical power by controlling the false discovery rate, yielding a new p-value, which is particularly suitable for large-scale genomics data analysis.
[0073] When the observed phenotype corresponds to the phenotype in the first gene-phenotype association pair selected in step S310, the signal term value of the candidate gene is calculated based on the phenotypic fingerprint and the observed phenotype. The purpose is to quantify the strength of evidence observed that matches a known gene phenotype association. The signal term value reflects the degree of matching between the observed phenotype of the target object and the phenotypic fingerprint of the candidate gene, providing positive support for prediction. One possible approach is to multiply the effect size of each corresponding observed phenotype (e.g., regression coefficient or odds ratio from the phenotypic fingerprint) by the confidence weight of that phenotype, and then sum the products of all corresponding phenotypes. Another possible approach is to construct a linear model based on the phenotypic fingerprint, taking the observed phenotype as input and directly outputting the signal term value, where the model parameters reflect the strength of the phenotype-gene association.
[0074] If the observed phenotype does not correspond to the phenotype in the first gene-phenotype association pair extracted in step S320, or to the phenotype in the second gene-phenotype association pair, a penalty term value for the candidate gene is calculated based on the phenotypic fingerprint and the observed phenotype. This penalty term negatively impacts observed phenotypes that do not match the expected phenotypic fingerprint of the candidate gene or are uncertain. The penalty term value aims to reduce the score of candidate genes whose phenotypic expression is inconsistent with the gene association pattern, thereby improving the specificity of the prediction. One possible approach is to multiply each non-corresponding observed phenotype by a preset penalty hyperparameter based on its deviation from the gene phenotypic fingerprint, and then accumulate these penalty values. Another approach is to calculate the penalty term based on the difference between the prevalence of the observed phenotype in the population and the expected phenotypic fingerprint of the gene, where rare phenotypes that do not match the gene may result in a larger penalty.
[0075] Based on the signal and penalty values of each candidate gene, a gene-centric multiphenotypic score is calculated for each candidate gene. The aim is to integrate positive support (signal value) and negative influence (penalty value) into a comprehensive score to fully assess the pathogenicity potential of the candidate gene. This score more accurately reflects the degree of agreement between the overall phenotypic characteristics of the target population and the specific gene's pathogenic pattern. One approach is to directly subtract the penalty value from the signal value to obtain a preliminary comprehensive score. Another approach is to standardize the signal and penalty values before performing a weighted sum or combining them using a nonlinear function to generate the final multiphenotypic score.
[0076] This embodiment introduces multiple hypothesis testing to conduct a second rigorous screening of the initially selected first gene-phenotype association pairs, thereby obtaining second gene-phenotype association pairs with greater statistical significance. This process significantly improves the reliability of gene-phenotype associations, laying a more solid foundation for subsequent pathogenic gene prediction. Based on this, when responding to the pathogenic gene prediction instruction for the target object and determining its observed phenotype, for observed phenotypes that match the first gene-phenotype association pair, this embodiment calculates a positive signal term value to quantify its support for the pathogenicity of the candidate gene; while for those phenotypes that match the first gene-phenotype association pair... For discrepancies, or more strictly, for discrepancies in association with the second gene phenotype selected through multiple hypothesis testing, the system calculates a penalty value. This penalty mechanism effectively reduces the risk of misjudgment due to uncertain or inconsistent phenotypes. Ultimately, by integrating the signal value and the penalty value, this embodiment can calculate a comprehensive, gene-centric multiphenotype score for each candidate gene. This score not only considers evidence of consistency with the gene's expected phenotype fingerprint but also fully considers evidence of discrepancies or absence, thus providing a more comprehensive and accurate risk assessment of pathogenic genes.
[0077] In some embodiments of this application, when the observed phenotype is a continuous phenotype, the process of calculating the signal term value of the candidate gene in step S310 includes:
[0078] (1);
[0079] in, Represents the target object For candidate genes The signal term values of the continuous phenotype, It represents the set of continuous phenotypes in the first gene phenotype association pairs corresponding to the candidate gene (obtained from phenotypic fingerprints). Represents the target object Observational phenotypes Phenotypic values, This indicates the observed phenotype determined based on phenotypic fingerprints. The effect size (obtained from phenotypic fingerprints). Indicates the observed phenotype The confidence weights; where, for the phenotype corresponding to the second gene-phenotype association pair, =1; for the phenotype corresponding to the first gene-phenotype association pair. It is a hyperparameter ranging from 0 to 1, obtained through cross-validation of the derivation queue.
[0080] When the observed phenotype is a binary phenotype, the process of calculating the signal term value of the candidate gene in step S310 includes:
[0081] (2);
[0082] in, Represents the target object For candidate genes The signal term values of the binary phenotype, It represents the set of binary phenotypes (obtained from phenotypic fingerprints) consisting of the first gene phenotype association pairs corresponding to the candidate gene. Represents the target object Observational phenotypes Phenotypic values, This indicates the observed phenotype determined based on phenotypic fingerprints. The effect size (obtained from phenotypic fingerprints). Represents the natural logarithm function. Indicates the observed phenotype The confidence weight.
[0083] Continuous phenotypes are typically obtained through measurement and have a continuous numerical range. They provide more detailed information on individual differences and help capture subtle changes in gene effects. Binary phenotypes are phenotypes that have only two possible states, usually expressed as "yes / no", "present / absent", "affected / not affected", etc.
[0084] Formula (1) provides a mathematical model for quantifying the contribution of continuous observed phenotypes to the pathogenicity of candidate genes. It achieves accurate calculation of signal intensity by combining the observed phenotypic values of the target object with the effect size and confidence weight determined in the gene phenotypic fingerprint. The calculation of Formula (1) can be implemented by a programming language (such as Python) using matrix operations or iterative loops on the set. The summation is performed on all continuous phenotypes. In practical applications, high-performance computing libraries (such as NumPy) can be used for optimization to improve computational efficiency.
[0085] Formula (2) provides a mathematical model for quantifying the contribution of binary observed phenotypes to the pathogenicity of candidate genes. It achieves accurate calculation of signal intensity by combining the binary observed phenotype values of the target object with the effect size and confidence weight determined in the gene phenotype fingerprint. Similar to the calculation of continuous phenotype signal terms, Formula (2) can also be implemented using a programming language.
[0086] This collection includes specific candidate genes. All continuous phenotypes identified as associated in the first gene-phenotype association pair (with a p-value less than the first threshold) are considered to have a potential association with the gene. This limits the range of phenotypes to be considered when calculating the continuous signal term value, ensuring that only continuous phenotypes with statistical association with the gene are included. During data preprocessing, this set can be dynamically constructed based on the type of gene-phenotype association pair and the association test results. Target object In specific continuous observation phenotypes The actual measured values, used as input for signal term calculation, directly reflect the individual phenotypic characteristics of the target object. These values can be numerical data obtained from clinical records, physical examination reports, or biological sample analysis. Quantifying genes and continuous phenotypes The strength and direction of the association between genotypes can be considered; for example, in regression analysis, it might be the regression coefficient of the influence of genotype on phenotypic value. In signal term calculations, the effect size, as a weight, reflects the importance of different phenotypes in contributing to gene pathogenicity. It can be estimated from large-scale population data using statistical methods such as linear regression and generalized linear models. Reflects the observed phenotype The reliability or importance of gene association can be assigned, for example, based on the p-value of the association test, the standard error of the effect size, or expert knowledge. It weights the contributions of different phenotypes in the signal term calculation, giving more reliable or more important phenotypes greater influence. This can be obtained by inverse or exponential transformation of the p-value of the association test; it can also be manually set based on the clinical importance of the phenotype in disease diagnosis; or optimized through machine learning methods such as cross-validation.
[0087] This collection includes specific candidate genes. All binary phenotypes identified as associated in the first gene-phenotype association pair. This limits the range of phenotypes to be considered when calculating binary signal term values, ensuring that only binary phenotypes statistically associated with the gene are included. Similarly, it is built through data preprocessing and filtering operations. Target object In specific binary observation phenotypes The actual state of the signal is usually represented as 0 or 1. As input to the signal term calculation, it directly reflects the individual binary phenotypic characteristics of the target object. It can be a Boolean value obtained from clinical diagnosis, medical history records, or genetic testing results. The Odds Ratio quantifies the relationship between genes and binary phenotypes. The strength of the association between two factors. It represents the ratio of the probability of a specific event (such as developing the disease) occurring in the exposed group (e.g., carrying a pathogenic gene variant) to the probability of the same event occurring in the unexposed group (e.g., not carrying a pathogenic gene variant). In the calculation of the binary phenotype signal term, the natural logarithm of the dominance ratio serves as the weight, reflecting the degree of influence of the gene on the risk of a binary phenotype. It can be estimated from large-scale population data using statistical methods such as logistic regression and chi-square test. It is a constant A logarithmic function with base 0. Its dominance ratio is... Converting to a logarithmic scale makes it mathematically more manageable and allows positive and negative correlations to be represented symmetrically. Reflects the observed phenotype The reliability or importance of gene association, its weighted contribution to the calculation of the signal term for different binary phenotypes, and its confidence weight compared to the continuous phenotype. Similarly, values can be assigned based on the p-value of the correlation test, the standard error of the effect size, or expert knowledge.
[0088] This embodiment provides a precise signal term calculation method for different types of observed phenotypes (such as continuous and binary phenotypes). By combining the observed phenotype values of the target object with the effect size determined in the gene phenotype fingerprint (such as the natural logarithm of the regression coefficient or the odds ratio) and confidence weights, this embodiment can accurately quantify the contribution of each observed phenotype to the pathogenicity of candidate genes. This meticulous differentiation and quantification avoids the errors that may be caused by uniformly processing different types of phenotypes, thereby significantly improving the accuracy and reliability of pathogenic gene prediction and making the prediction results more consistent with biological reality.
[0089] In some embodiments of this application, when the observed phenotype is a continuous phenotype, the process of calculating the penalty term value of the candidate gene includes:
[0090] (3);
[0091] in, Represents the target object For candidate genes The penalty term value for continuous phenotypes, This represents the set of phenotypes from the first gene-phenotype association pairs corresponding to the candidate gene, or the set of phenotypes from the second gene-phenotype association pairs corresponding to the candidate gene. Represents the target object Observational phenotypes Phenotypic values, Indicates the observed phenotype The mean (statistics from the derivation queue). The penalty hyperparameter (0~1, obtained by cross-validation from heap to queue).
[0092] When the observed phenotype is binary, the process of calculating the penalty term value for candidate genes includes:
[0093] (4);
[0094] in, Represents the target object For candidate genes The penalty value of the binary phenotype, Indicates the observed phenotype The mean (statistics from the derivation queue).
[0095] Among them, the penalty term value of continuous phenotype Indicates targeting the target object candidate genes This is the penalty term calculated when the observed phenotype is continuous. This value is used to quantify the penalty term when the observed phenotype is continuous. It does not belong to the set of phenotypes that are significantly associated with this gene. This refers to the negative impact or uncertainty of the phenotype on the prediction of gene pathogenicity risk. The calculation method considers the difference between the observed phenotype value and the mean of that phenotype, and multiplies it by a penalty hyperparameter. Penalty hyperparameter This is an adjustable parameter used to control the weight of the penalty term in the total score. The setting of this parameter can be adjusted according to the actual application scenario, dataset characteristics, and the required sensitivity of the prediction model. For example, it can be set as a fixed constant or optimized through machine learning methods such as cross-validation to achieve the best predictive performance. Phenotypic set This refers to a set of phenotypes from multiple first-gene phenotype association pairs or multiple second-gene phenotype association pairs. This set represents the phenotypes associated with a specific candidate gene. A phenotype that is significantly associated with a gene, when the observed phenotype is not in this set, means that the phenotype is not associated with the gene. The correlation is weak or nonexistent, therefore it needs to be penalized. Observed phenotypic values Represents the target object Observational phenotypes The specific numerical value. For continuous phenotypes, this is usually a real value, such as blood pressure, blood sugar level, etc. Observed phenotype mean. Indicates the observed phenotype The mean value in the reference population or training dataset. By comparing the observed phenotypic values of the target population with the mean value, the degree of deviation of the phenotypic value from the target population can be assessed.
[0096] Penalty values for binary phenotypes Indicates targeting the target object candidate genes This is the penalty term calculated when the observed phenotype is binary. This value is used to quantify the penalty term when the observed binary phenotype is... It does not belong to the set of phenotypes that are significantly associated with this gene. And the target object This binary phenotype (i.e.) When =1), the negative impact of this phenotype on the prediction of gene pathogenicity risk. Observed phenotypic mean. Indicates the observed phenotype The average frequency or prevalence in the reference population or training dataset is typically a probability value between 0 and 1 for a binary phenotype.
[0097] In the above method, to more comprehensively assess the pathogenicity risk of candidate genes, this embodiment further introduces the calculation of a penalty term. When the observed phenotype of the target object does not correspond to the identified associated phenotypes in the gene's phenotypic fingerprint, that is, the observed phenotype does not belong to the set of phenotypes significantly associated with that gene. In this case, a penalty needs to be applied to the score of the gene. For continuous observation phenotypes, the penalty term is calculated based on the observed phenotype values. Compared with the mean of this phenotype in the population The difference between them. The accumulation of this difference, multiplied by the penalty hyperparameter. This constitutes the penalty term value for the continuous phenotype. The principle is that if a continuous phenotype that is not significantly associated with a gene shows a significant deviation from its mean in the target population, it may mean that the pathogenicity of the gene is not manifested through this phenotype, or that the association between the phenotype and the gene is weak. Therefore, an appropriate penalty is needed to avoid its excessive contribution to the signal term. For binary observed phenotypes, the penalty term is calculated focusing on those phenotypes that do not belong to the gene's mean. And the target object actually exhibits a binary phenotype ( =1), at this time, the penalty term By penalizing hyperparameters Multiply It is obtained by accumulating, where This represents the average frequency of this binary phenotype in the population. The principle is that if a binary phenotype not significantly associated with a gene appears in the target population, and this phenotype is itself relatively rare in the population (i.e.,...), then... Smaller, resulting in If the signal is significantly associated with a gene, then its occurrence should be considered noise or nonspecificity, thus incurring a greater penalty for that gene's score. By introducing these two types of penalty terms, the proposed scheme effectively compensates for the shortcomings of relying solely on the signal term. It not only considers phenotypic information significantly associated with the gene but also quantifies observed phenotypes that are not strongly associated with the gene or are absent. This processing mechanism enables the final gene center multiphenotype score to more accurately reflect the true pathogenicity risk of candidate genes, avoiding interference from nonspecific or weakly associated phenotypes, thereby improving the accuracy and reliability of pathogenic gene prediction.
[0098] This embodiment can quantify observed phenotypes that are not strongly associated with candidate genes or do not exist, and incorporate them as penalty terms into the calculation of the multiphenotype score of the gene center. This effectively avoids the evaluation bias that may be caused by relying solely on signal terms, and enables the prediction model to more comprehensively consider the phenotypic characteristics of the target object. Specifically, by distinguishing between continuous and binary phenotypes and using different penalty calculation formulas, this embodiment can more precisely capture the potential impact of unrelated phenotypes on the risk of gene pathogenicity. This mechanism significantly improves the accuracy and robustness of pathogenic gene prediction, making the priority ranking of candidate genes more reliable, thereby providing a more accurate basis for clinical diagnosis and personalized treatment.
[0099] In some embodiments of this application, a gene-centered multiphenotype score is calculated for each candidate gene based on the signal term value and penalty term value of each candidate gene, including:
[0100] (5);
[0101] (6);
[0102] (7);
[0103] (8);
[0104] in, The signal term value represents the continuous phenotype of the target object relative to the candidate genes. This represents the penalty term value for the continuous phenotype of the target object relative to the candidate gene. The signal term value represents the binary phenotype of the target object relative to the candidate gene. This represents the penalty value of the binary phenotype of the target object relative to the candidate gene. This indicates Z-score standardization. Represents the natural base. This represents the candidate gene's multiphenotype score centered on the gene itself.
[0105] Specifically, Formula (5) aims to integrate the signal value and corresponding penalty value of a continuous phenotype to obtain a comprehensive continuous phenotype score. The signal value reflects the strength of evidence linking the observed phenotype to the gene, while the penalty value considers the negative impact of unrelated phenotypes on the score. Through subtraction, positive association evidence and negative non-association evidence can be effectively balanced, making the final score more robust. Similarly, Formula (6) is used to integrate the signal value and penalty value of a binary phenotype to obtain a comprehensive binary phenotype score. A binary phenotype usually represents the presence or absence of a certain trait. The calculation methods of its signal and penalty values are different from those of a continuous phenotype, but the integration purpose is the same, that is, to balance positive association evidence and negative non-association evidence through subtraction to improve the accuracy of the score. For example, direct subtraction can be used, or subtraction can be performed after weighted averaging, or a combination can be performed using a nonlinear function.
[0106] Formula (7) aims to standardize and merge adjusted continuous phenotypic scores and binary phenotypic scores to obtain a unified comprehensive score. Z represents Z-score standardization, a commonly used data standardization method that can transform data with different dimensions and distributions to a uniform scale, eliminating the impact of dimensional differences on score merging. By averaging the two standardized scores, it can be ensured that continuous and binary phenotypic scores have similar weights in the final comprehensive score. For example, in addition to simple averaging, a weighted average of the two scores can be performed according to the actual situation, or other statistical methods such as principal component analysis can be used for fusion. Z-score standardization can convert the original data into standard scores. Through Z-score standardization, data with different distributions can be converted into a standard normal distribution with a mean of 0 and a standard deviation of 1, thereby eliminating dimensional differences and enabling effective comparison and merging of different types of scores.
[0107] Formula (8) converts the overall score into a probability value between 0 and 1, namely the gene-centered multiphenotypic score (GPS). Representing the natural base, it is an important mathematical constant, approximately equal to 2.71828. It converts the comprehensive score into the risk probability of the target object carrying the corresponding pathogenic mutation, making the scoring results more biologically interpretable.
[0108] This embodiment first integrates the signal and penalty values of continuous and binary phenotypes to obtain an adjusted score, ensuring that the net association evidence for each phenotype is accurately reflected. Then, these adjusted scores are Z-score standardized to eliminate differences in dimensions and distributions between different phenotypes, allowing them to be compared and combined on a uniform scale. Next, a comprehensive score is obtained by averaging the standardized scores. This score represents the overall association strength of the gene across all relevant phenotype dimensions. Finally, the comprehensive score is converted into a GPS value between 0 and 1 using the Sigmoid function. This value can be directly interpreted as the risk probability of the target object carrying the corresponding pathogenic variant. This series of steps, organically combined, enables the effective integration of different types of phenotype information and transforms it into a unified, standardized, and biologically meaningful risk prediction indicator, thus solving the problem of how to effectively integrate different types of scoring items and transform them into an interpretable pathogenic gene prediction score.
[0109] This embodiment can effectively integrate different types of phenotypic data (continuous and binary) and their corresponding signal and penalty terms, and generate a unified and interpretable gene-centric multiphenotypic score through standardization and probability transformation. This makes the prediction results of pathogenic genes more biologically meaningful and clinically valuable, and can more accurately assess the risk of target subjects carrying pathogenic variants, thereby providing strong support for disease diagnosis, risk assessment and personalized treatment.
[0110] like Figures 2 to 6 For ease of understanding, this application provides the following embodiments, which take autism spectrum disorder (ASD) as an example. The clinical manifestations of ASD are highly heterogeneous.
[0111] This embodiment can be used to assist in genetic diagnosis and genetic counseling in complex disease scenarios. The core logic of this embodiment is to construct gene-specific phenotypic fingerprints at the population level and transform multidimensional clinical phenotypic information into gene-centric quantitative scores (referred to as GPS (Gene-centric Polyphenotypic Score) in this embodiment) at the object level. This enables the same process to achieve two key outputs: risk assessment of the target object as a carrier of pathogenic variants and priority ranking of candidate genes, and provides interpretable data support for the interpretation of clinical genetic variants and gene localization.
[0112] Step S910, sample screening;
[0113] 142,357 samples with exon sequencing data (ES) were input for ASD. The samples were filtered based on the completeness of the phenotypic data, the quality of the sequencing data, the consistency of phylogenetic relationships, and the confidence of the disease diagnosis. Finally, 131,815 samples were obtained, of which 44,962 were ASD patients and 86,853 were non-ASD patients.
[0114] Phenotypes were derived from the subjects' (children with ASD) clinical scale responses, developmental history records, and comorbidity / complication-related survey information. Phenotype selection was based on the core ASD symptom dimensions emphasized in DSM-5 (social communication and social interaction deficits, repetitive behaviors) and previous research evidence, prioritizing phenotypes that best characterize the core clinical features and common comorbidities of ASD. All phenotypes were standardized and categorized.
[0115] The core characteristics of ASD were assessed using four standardized scales: the Social Communication Questionnaire (SCQ), the Repetitive and Stereotyped Behaviors Scale (RBSR), the DCDQ (Developmental Coordination Disorder Questionnaire), and the ABC (Adaptive Behaviors Scale). The data were obtained as scores for each item on each scale, and the data were standardized into total scale scores and subscale scores. Furthermore, both total scores and subscale scores were included, resulting in 20 scale-related phenotypes used to quantify symptoms and functional performance across different dimensions. In addition to the core scales, this embodiment further included 10 developmental milestones to characterize early developmental trajectories (e.g., whether or not they were reached, or the time point at which they were reached, uniformly coded according to data availability), and five comorbidities were included: mental health (n=19), neurodevelopmental (n=10), perinatal complications (n=5), birth defects (n=6), and growth and developmental abnormalities (n=5). The final result was a standardized phenotype panel containing 75 phenotypes across 7 major categories.
[0116] This embodiment explicitly classifies phenotypes into two data types: continuous phenotypes, including total / subscale scores and quantitative data on developmental milestones; and binary phenotypes, including five categories: mental and psychological, neurodevelopmental, perinatal complications, birth defects, and growth and developmental abnormalities, uniformly coded as "present" or "absent" state variables. Specific phenotypes are shown in the table below:
[0117] Table 1
[0118]
[0119] Step S920, Mutation data processing;
[0120] Step S9210, Quality Control of Mutation Data:
[0121] The genotypes (VCF format) of the 44,962 ASD patients screened were quality controlled using bcftools and PLINK tools according to the following criteria:
[0122] GQ < 30, DP < 7 (for SNP-type mutations);
[0123] DP < 10 (for InDel type mutations);
[0124] AB < 0.15 (for SNP-type variants);
[0125] AB < 0.2 (for InDel type variants);
[0126] callrate < 0.9, p-value of the Hawes equilibrium test < 1× The final output is a VCF format variant file after quality control.
[0127] Step S9220, mutation annotation;
[0128] For the variant data after quality control, ANNOVAR software was used for annotation, and the annotated TSV file was output. Each row represents a variant, and each column represents an annotation result, including: variant type, gene, functional impact, harmfulness score, population frequency, and clinical information.
[0129] Step S9230, mutation screening;
[0130] Variants were filtered based on their annotation information. Those belonging to genes with scores of 1 and S in the SFARIGene database, a population frequency <0.01%, and functional impact and harmfulness classified as LoF or DMis were selected. The final filtered variant data was then saved as a TSV file.
[0131] Step S9240, classification of pathogenicity of variants;
[0132] After screening, each variant was classified into five categories according to the ACMG / AMP variant rating guidelines: pathogenic (P), possibly pathogenic (LP), variant of unknown significance (VUS), possibly benign (LB), and benign (B). Finally, 3726 variants with the pathogenicity category of P or LP were selected and saved as TSV files.
[0133] Step S930, Genotype data construction;
[0134] Step S9310, genotype coding;
[0135] Genotyping was performed on the genes of each sample from 44,962 ASD patients: the gene was genotyped as 1 if it carried any one or more of the 3,726 P / LP variants, and 0 otherwise.
[0136] Step S9320, Genotype screening;
[0137] After filtering out genes with fewer than 5 carriers whose genotypes encode 1, 186 genes were ultimately retained for downstream analysis.
[0138] Step S940: Construction of the association map and phenotypic fingerprint between genes and phenotypes;
[0139] Step S9410: Pairwise pairwise combinations of 186 genes and 75 phenotypes are performed to obtain 14,508 gene-phenotype association pairs.
[0140] Step S9420: Perform independent association tests on each of the 14,508 gene-phenotype association pairs. Gene-phenotype association pairs with binary traits are tested using Firth logistic regression, while gene-phenotype association pairs with continuous traits are tested using linear regression. Each gene-phenotype association pair is adjusted for covariates such as sex, age, genetic ancestry, and sequencing batch. The test results are then subjected to multiple correction using FDR.
[0141] Step S9430: Perform step-by-step screening based on the association test results to draw an association map between genes and phenotypes;
[0142] First, from all 14,508 gene-phenotype association pairs, 1,214 gene-phenotype association pairs with a p-value less than 0.05 were selected and defined as the first gene-phenotype association pair, and their corresponding phenotypes were defined as the nominal association phenotypes.
[0143] Step S9440: After further screening and correction, the 154 association pairs with a P-value less than 0.05 are defined as second gene phenotype association pairs, and their corresponding phenotypes are defined as significantly associated phenotypes. Finally, an association map is drawn based on all 1214 first gene phenotype association pairs (including 154 second gene phenotype association pairs).
[0144] Step S9450: Construct phenotypic fingerprint;
[0145] Based on association maps, the association direction, effect size, and significance information of each gene across all phenotypic dimensions are integrated into a gene phenotypic fingerprint. This gene phenotypic fingerprint reveals the specific patterns by which genes influence different clinical characteristics.
[0146] Step S950, GPS scoring;
[0147] This embodiment proposes a gene-centered multiphenotypic scoring system (GPS) to transform a subject's multidimensional clinical phenotypic information into the probability of carrying a pathogenic variant of a target gene, and thereby achieves two types of outputs:
[0148] 1) Quantitatively assess the risk propensity of individuals to carry pathogenic / potentially pathogenic gene variants (P / LP);
[0149] 2) Prioritize multiple candidate genes within the same target object.
[0150] Step S950 involves two application scenarios:
[0151] Scenario 1: Interpretation of phenotypic driver genes from clinical sequencing data;
[0152] In such scenarios, the target subject has typically been identified with multiple undetermined genetic variants (VUS) through sequencing. The core question facing clinicians is: among the many VUS, which one is most likely to be related to the patient's current complex clinical presentation? Traditional methods rely on experts manually comparing genes and phenotypes, which is time-consuming and subjective. The GPS method in this embodiment uses algorithms to automate and quantify the reverse scoring of candidate genes using patient phenotypes, providing objective evidence for gene interpretation and assisting clinical decision-making.
[0153] Scenario 2: Phenotype-driven pre-sequencing gene prioritization;
[0154] In clinical practice, when faced with a patient suspected of having a genetic disease and exhibiting a complex, non-specific phenotype, doctors often face a strategic choice challenge when prescribing genetic testing: how to quickly screen out the most worthy target genes from among numerous possible pathogenic genes, thereby improving diagnostic efficiency and reasonably controlling testing costs? Another core application of GPS is designed to address this pain point. Based on the patient's detailed clinical phenotype, candidate genes can be quantitatively scored before any genetic testing is performed, automatically generating a data-driven list of priority genes for testing. This not only helps doctors focus on high-probability genes at the initial stage of testing, significantly shortening the analysis cycle, but also accelerates the overall clinical decision-making process by optimizing the testing pathway. This improves diagnostic accuracy while helping to reduce medical resource consumption and improve the patient's treatment experience.
[0155] Step S9510, data preprocessing;
[0156] Of the 44,962 ASD patients obtained earlier, samples with a phenotype loss rate of more than 20% were further excluded, resulting in 26,529 samples, which served as the derivation cohort for GPS.
[0157] Step S9520, GPS calculation;
[0158] 1) Signal term calculation;
[0159] For candidate genes With the target object This embodiment is based on candidate genes. Phenotypic fingerprinting, and extraction of candidate genes from phenotypic fingerprints. The nominal association phenotype is used to calculate the signal term value. The calculation formula for the continuous phenotype is the above formula (1), and the calculation formula for the binary phenotype is the above formula (2).
[0160] 2) Calculation of penalty items;
[0161] Nominal and significant association phenotypes are extracted from fingerprint phenotypes, and non-genetic association phenotypes observed in the target object are analyzed. For nominal and significant association phenotypes, GPS introduces a penalty term to penalize them. Similarly, the penalty term is calculated separately for continuous and binary phenotypes. The calculation formula for continuous phenotypes is formula (3) above, and the calculation formula for binary phenotypes is formula (4) above.
[0162] 3) GPS final score calculation;
[0163] The calculation formulas are formulas (5) to (8) above, which will not be repeated here.
[0164] Step S960, GPS performance evaluation;
[0165] 1) Single-gene level assessment;
[0166] To facilitate clinical interpretation and practical application, this embodiment reports two types of indicators for each gene: the first is the discriminative ability, which uses AUC to measure whether GPS can distinguish between gene P / LP carriers and non-carriers, and reports the median AUC and 95% confidence interval through 1000 bootstrap replicate sampling; the second is the ranking practicality, which sorts the genes within each gene from high to low according to GPS, defines the top 5% of the population as the high priority layer, calculates the enrichment degree of P / LP carriers in the layer relative to the rest of the population, and uses a two-sided Fisher exact test to assess significance.
[0167] like Figure 3 and Figure 4 In the derived data above, GPS performed well overall on 173 genes with nominally associated phenotypes, with a median AUC of 0.795, of which 83 genes achieved an AUC > 0.8. For the 48 FDR-significant genes, GPS performance was further improved: 85.4% (41 / 48) of the FDR-significantly associated genes had an AUC > 0.8, and 33.3% (16 / 48) had an AUC > 0.9; accompanied by significant carrier enrichment, with a median OR of 19.061, and 72.9% (35 / 48) of the genes having an OR > 10. For the 42 nominally significant but not FDR-significant high-performance genes, GPS also maintained strong discriminative and enrichment capabilities: these genes all met the AUC > 0.8 standard, with a median OR of 14.293 for carrier enrichment, reaching a maximum of 114.629. The results show that GPS can not only reliably identify carriers in high-confidence genes, but also extract phenotypic signals with ranking value in some genes that have not yet reached the strict multiple correction threshold, thereby expanding the range of applicable genes.
[0168] 2) Overall assessment;
[0169] Since the distribution scales of the raw GPS scores of different genes may differ, directly using the raw GPS scores to integrate different genes could affect the fairness of cross-gene comparisons. Therefore, this embodiment uses GPS percentiles for integration, merging multiple genes into a unified dataset for global evaluation to eliminate scale differences. The global evaluation consists of two parts: First, calculating the overall discriminative power (AUC, while also reporting APR), and using 1000 bootstrap iterations to provide 95% confidence intervals and plotting ROC curves; Second, evaluating the enrichment efficiency of high-stratified carriers, dividing the subjects into high-priority strata such as Top 20%, 10%, 5%, 1%, and 0.1% according to GPS percentiles, and calculating the P / LP carrier enrichment OR and 95% confidence intervals for each stratum (displayed on logarithmic coordinates).
[0170] like Figure 5 To facilitate clinical use, this embodiment defines two gene sets in the global assessment: a high-confidence set, which consists of genes with significant FDR and a single-gene AUC > 0.8; and an extended set, which adds nominally significant genes with a single-gene AUC > 0.8 to the high-confidence set to cover more available genes. Based on the global assessment results, the high-confidence set AUC was 0.869 (95% CI: 0.853–0.884), and the extended set AUC was 0.865 (95% CI: 0.851–0.877). Furthermore, in the top 20%, 10%, 5%, 1%, and 0.1% strata, the enrichment of P / LP carriers increased with tightening of the strata, indicating that GPS can effectively concentrate carriers at the top of the ranking, consistent with the clinical decision-making logic of prioritizing high-strata assessments.
[0171] 3) Application and verification of independent queues;
[0172] like Figure 6 To apply GPS and verify its generalization ability, this embodiment includes an independent external validation cohort, and preprocessing is performed using the same analytical methods. To simulate real-world applications, all variable parameters of GPS are fixed within the derived cohort. The final results show that the global AUC for the high-confidence gene set is 0.765, and the global AUC for the extended gene set is 0.771. Furthermore, P / LP carriers showed significant enrichment in different high-level stratifications (Top 20%, 10%, 5%, 1%, 0.1%), indicating that GPS can stably concentrate carriers at the top of the ranking even in independent samples. This further verifies that the GPS model has high accuracy and reliability, thus supporting its application in clinical scenarios for phenotype-driven gene prioritization and carrier risk stratification.
[0173] This embodiment first constructs gene-specific phenotypic fingerprints at the population level, and then develops a GPS tool through a systematic technical solution. This tool is used to convert multidimensional clinical phenotypic information into gene-centric quantitative scores at the object level, enabling the localization of phenotypic driver genes and the prioritization of pathogenic genes. This provides important data and tool support for the prediction and ranking of pathogenic variants in complex diseases. Compared with existing technologies, this embodiment has at least the following beneficial effects:
[0174] 1) Reduce single dependencies on existing annotation libraries;
[0175] Existing technologies heavily rely on existing gene and disease / phenotype annotation libraries, and incomplete or inconsistent annotations can easily lead to biases in candidate gene ranking. This embodiment constructs gene-specific phenotypic fingerprints based on a large-scale cohort and performs quantitative scoring at the individual level. Even when faced with atypical or novel phenotypic combinations, it can maintain the systematicity and stability of candidate gene evaluation, thereby reducing omissions and result fluctuations.
[0176] 2) Formation of quantitative mechanisms for phenotypic driving and gene center formation;
[0177] Existing technologies mostly follow an interpretation path from gene to phenotype, with phenotypes often used as posterior filters, making it difficult to directly convert multidimensional phenotypes into comparable gene scores. This embodiment maps individual multidimensional phenotypes to a gene fingerprint space, generating a GPS score for each gene. This enables interpretable ranking of pathogenic genes by inferring their causes from phenotypes, making it suitable for complex disease scenarios with a large number of candidate genes, highly heterogeneous phenotypes, and significant gene-specific effects, significantly improving the efficiency and consistency of clinical screening.
[0178] 3) Integrated output of risk assessment and gene sequencing;
[0179] This embodiment can simultaneously output an individual's carrier tendency (risk stratification) for the target gene P / LP variant and the priority ranking of candidate genes under the same scoring framework, simplifying the workflow and improving reproducibility and feasibility.
[0180] 4) Provide quantifiable quality control indicators;
[0181] This embodiment provides quantitative indicators such as discriminative ability and high stratified enrichment effect at the single gene and global levels, which can directly support stratification decisions on which genes / individuals to prioritize and facilitate consistent verification and continuous quality control in independent data.
[0182] In summary, this embodiment achieves a more systematic, interpretable, and easily implementable priority ranking and carrier risk assessment of pathogenic genes in complex disease scenarios. Compared with existing technologies, it has significant improvements in coverage, stability, integrated output, and clinical robustness, filling the gaps in existing technologies and providing important tools and data support for assessing the pathogenicity of genetic variations.
[0183] Reference Figure 7 Some embodiments of this application provide a pathogenic gene prediction device based on phenotypic fingerprinting, the device comprising:
[0184] The phenotypic fingerprint construction module 1100 is used to construct multiple gene-phenotype association pairs based on multiple phenotypes and multiple genes associated with at least one subject and the target disease, and to determine the association information of each gene in all phenotypic dimensions based on the multiple gene-phenotype association pairs, so as to construct the phenotypic fingerprint of each gene according to the association information; wherein, any gene-phenotype association pair consists of any phenotype and any gene, and any two gene-phenotype association pairs are different; each gene includes at least one pathogenic variant or a potentially pathogenic variant;
[0185] The prediction instruction response module 1200 is used to respond to the pathogenic gene prediction instruction of the target object and determine the observed phenotype of the target object related to the target disease.
[0186] The risk prediction and ranking module 1300 is used to calculate the gene-centered multiphenotypic score of each candidate gene of the target object based on the phenotypic fingerprint and observed phenotype, so as to predict the risk value of the target object carrying the corresponding pathogenic variant and determine the priority of each candidate gene based on the multiphenotypic score of each candidate gene.
[0187] It should be noted that the pathogenic gene prediction device based on phenotypic fingerprints provided in this embodiment is based on the same inventive concept as the pathogenic gene prediction method based on phenotypic fingerprints described above. Therefore, the relevant content of the pathogenic gene prediction method based on phenotypic fingerprints described above also applies to the content of the pathogenic gene prediction device based on phenotypic fingerprints. Therefore, it will not be repeated here.
[0188] like Figure 8 This application also provides an electronic device, which includes:
[0189] The device includes at least one hydrogen fuel cell; at least one memory; at least one processor; and at least one program. The program is stored in the memory, and the processor executes the at least one program to implement the pathogenic gene prediction method based on phenotypic fingerprinting described above. The electronic device can be any smart terminal, including mobile phones, tablets, personal digital assistants (PDAs), and in-vehicle computers. The electronic device according to embodiments of this application will be described in detail below.
[0190] The processor 1600 can be implemented using a general-purpose central processing unit (CPU), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this disclosure.
[0191] The memory 1700 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 1700 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1700 and is called and executed by the processor 1600 to execute a pathogenic gene prediction method based on phenotypic fingerprints according to an embodiment of this disclosure.
[0192] The input / output interface 1800 is used to implement information input and output.
[0193] The communication interface 1900 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0194] Bus 2000 transmits information between various components of the device (e.g., processor 1600, memory 1700, input / output interface 1800, and communication interface 1900);
[0195] The processor 1600, memory 1700, input / output interface 1800 and communication interface 1900 are connected to each other within the device via bus 2000.
[0196] This disclosure also provides a storage medium, which is a computer-readable storage medium storing computer-executable instructions for causing a computer to execute the above-described method for predicting pathogenic genes based on phenotypic fingerprints.
[0197] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0198] The embodiments described in this disclosure are for the purpose of more clearly illustrating the technical solutions of this disclosure and do not constitute a limitation on the technical solutions provided by this disclosure. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by this disclosure are also applicable to similar technical problems.
[0199] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this disclosure, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0200] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0201] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0202] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0203] It should be understood that in this application, "at least one (item)" means one or more, and "more than one" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0204] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or objects may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0205] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0206] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0207] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions to cause an electronic device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0208] The above is a detailed description of the preferred embodiments of this application. However, the embodiments of this application are not limited to the above-described implementation methods. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the embodiments of this application. All such equivalent modifications or substitutions are included within the scope defined by the claims of the embodiments of this application.
Claims
1. A method for pathogenic gene prediction based on a phenotypic fingerprint, characterized by, The method includes: Based on at least one subject's multiple phenotypes and multiple genes associated with the target disease, multiple gene-phenotype association pairs are constructed, and the association information of each gene across all phenotypic dimensions is determined based on the multiple gene-phenotype association pairs to construct a phenotypic fingerprint for each gene according to the association information; wherein any gene-phenotype association pair consists of any phenotype and any gene, and any two gene-phenotype association pairs are different; each gene includes at least one pathogenic variant or potentially pathogenic variant; In response to the pathogenic gene prediction instruction of the target object, determine the observed phenotype of the target object associated with the target disease; Based on the phenotypic fingerprint and the observed phenotype, a gene-centered multiphenotypic score is calculated for each candidate gene of the target object, in order to predict the risk value of the target object carrying the corresponding pathogenic variant and determine the priority of each candidate gene based on the multiphenotypic score of each candidate gene. 2.The pathogenic gene prediction method based on a phenotypic fingerprint according to claim 1, characterized in that, The step of determining the association information of each gene across all phenotypic dimensions based on the multiple gene-phenotype association pairs, and constructing a phenotypic fingerprint for each gene based on the association information, includes: Independent association tests were performed on each of the multiple gene-phenotype association pairs to obtain the association test results. Based on the association test results, an association map of the multiple gene phenotype association pairs is drawn; Based on the association map, determine the association direction, effect size, and significance information of each gene across all phenotypic dimensions; A phenotypic fingerprint of each gene is constructed based on the association direction, effect size, and significance information. 3.The pathogenic gene prediction method based on a phenotypic fingerprint according to claim 2, characterized in that, The association test results include the test p-value corresponding to each of the gene-phenotype association pairs; The step of drawing an association map of the multiple gene phenotype association pairs based on the association test results includes: Based on the test P value corresponding to each gene-phenotype association pair, multiple first gene-phenotype association pairs with test P values less than a first threshold are selected from the multiple gene-phenotype association pairs; wherein, the first threshold is used to characterize the association between the gene and the phenotype in the gene-phenotype association pair. The association map is drawn based on the multiple first gene phenotype association pairs. 4.The pathogenic gene prediction method based on a phenotypic fingerprint according to claim 3, characterized in that, Before drawing the association map based on the plurality of first gene phenotype association pairs, the method further includes: Based on multiple hypothesis testing, update the test p-value; Based on the updated test p-value, multiple second gene phenotype association pairs are selected from the multiple first gene phenotype association pairs; The step of calculating a gene-centered multiphenotypic score for each candidate gene of the target object based on the phenotypic fingerprint and the observed phenotype includes: Extract the phenotype from the first gene phenotype association pair corresponding to the candidate gene from the phenotype fingerprint, and calculate the signal term value of the candidate gene based on the phenotype fingerprint and the observed phenotype when the observed phenotype corresponds to the phenotype in the extracted first gene phenotype association pair. The phenotypes in the first gene phenotype association pair corresponding to the candidate gene and the phenotypes in the second gene phenotype association pair corresponding to the candidate gene are extracted from the phenotypic fingerprint. If the observed phenotype does not correspond to the phenotype in the first gene phenotype association pair or the phenotype in the second gene phenotype association pair, the penalty term value of the candidate gene is calculated based on the phenotypic fingerprint and the observed phenotype. Based on the signal term value and penalty term value of each candidate gene, a gene-centered multiphenotype score is calculated for each candidate gene.
5. The method of claim 4, wherein the method is based on a phenotypic fingerprint. When the observed phenotype is a continuous phenotype, the process of calculating the signal term value of the candidate gene includes: ; in, Represents the target object For candidate genes The signal term values of the continuous phenotype, This represents the set of continuous phenotypes from the first gene phenotype association pairs corresponding to candidate genes. Represents the target object Observational phenotypes Phenotypic values, This indicates the observed phenotype determined based on phenotypic fingerprints. The effect size, Indicates the observed phenotype Confidence weights; When the observed phenotype is a binary phenotype, the process of calculating the signal term value of the candidate gene includes: ; in, Represents the target object For candidate genes The signal term values of the binary phenotype, This represents the set of binary phenotypes from the first gene phenotype association pairs corresponding to the candidate gene. Represents the target object Observational phenotypes Phenotypic values, This indicates the observed phenotype determined based on phenotypic fingerprints. The effect size, Represents the natural logarithm function. Indicates the observed phenotype The confidence weight.
6. The method for predicting pathogenic genes based on phenotypic fingerprints according to claim 4, characterized in that, When the observed phenotype is a continuous phenotype, the process of calculating the penalty term value of the candidate gene includes: ; in, Represents the target object For candidate genes The penalty term value for continuous phenotypes, This represents the set of phenotypes from the first gene-phenotype association pairs corresponding to the candidate gene, or the set of phenotypes from the second gene-phenotype association pairs corresponding to the candidate gene. Represents the target object Observational phenotypes Phenotypic values, Indicates the observed phenotype The mean, To penalize hyperparameters; When the observed phenotype is a binary phenotype, the process of calculating the penalty term value of the candidate gene includes: ; in, Represents the target object For candidate genes The penalty value of the binary phenotype, Indicates the observed phenotype The mean.
7. The method for predicting pathogenic genes based on phenotypic fingerprints according to claim 4, characterized in that, The step of calculating a gene-centered multiphenotype score for each candidate gene based on the signal term value and penalty term value of each candidate gene includes: ; ; ; ; in, The signal term value represents the continuous phenotype of the target object relative to the candidate genes. This represents the penalty term value for the continuous phenotype of the target object relative to the candidate gene. The signal term value represents the binary phenotype of the target object relative to the candidate gene. This represents the penalty value of the binary phenotype of the target object relative to the candidate gene. This indicates Z-score standardization. Represents the natural base. This represents the candidate gene's multiphenotype score centered on the gene itself.
8. A pathogenic gene prediction device based on phenotypic fingerprinting, characterized in that, The device includes: The phenotypic fingerprint construction module is used to construct multiple gene-phenotype association pairs based on multiple phenotypes and multiple genes associated with at least one subject and the target disease, and to determine the association information of each gene across all phenotypic dimensions based on the multiple gene-phenotype association pairs, so as to construct a phenotypic fingerprint for each gene according to the association information; wherein any gene-phenotype association pair consists of any phenotype and any gene, and any two gene-phenotype association pairs are different; each gene includes at least one pathogenic variant or potentially pathogenic variant; The prediction instruction response module is used to respond to the pathogenic gene prediction instruction of the target object and determine the observed phenotype of the target object related to the target disease. The risk prediction and ranking module is used to calculate a gene-centered multiphenotypic score for each candidate gene of the target object based on the phenotypic fingerprint and the observed phenotype, so as to predict the risk value of the target object carrying the corresponding pathogenic variant and determine the priority of each candidate gene based on the multiphenotypic score of each candidate gene.
9. An electronic device, characterized in that: The method includes at least one control processor and a memory for communicatively connecting to the at least one control processor; the memory stores instructions executable by the at least one control processor to enable the at least one control processor to perform a pathogenic gene prediction method based on phenotypic fingerprints as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions for causing a computer to perform a pathogenic gene prediction method based on phenotypic fingerprints as described in any one of claims 1 to 7.