Pathogenic gene prediction method, device and equipment based on phenotypic fingerprints and medium

By constructing gene-specific phenotypic fingerprints and performing multi-dimensional quantitative scoring, the problem of low efficiency in candidate gene screening in complex diseases has been solved, achieving efficient and objective identification of pathogenic genes and reducing reliance on expert experience.

CN121838892AActive Publication Date: 2026-04-10XIANGYA HOSPITAL CENT SOUTH UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIANGYA HOSPITAL CENT SOUTH UNIV
Filing Date
2026-03-16
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

In complex disease scenarios, existing technologies are inefficient in candidate gene screening and sorting, and the consistency of results depends on expert experience. It is difficult to identify the most likely pathogenic genes in a short period of time, especially when there are multiple systems and multiple dimensions of phenotypes.

Method used

The pathogenic gene prediction method based on phenotypic fingerprints constructs gene-specific phenotypic fingerprints and transforms multidimensional clinical phenotypic information into gene-centered quantitative scores at the object level, thereby achieving risk assessment of pathogenic variants and priority ranking of candidate pathogenic genes.

Benefits of technology

It significantly improves the efficiency and objectivity of identifying pathogenic genes in complex diseases, reduces reliance on expert experience, and provides a systematic, quantitative, and efficient pathogenic gene prediction scheme.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121838892A_ABST
    Figure CN121838892A_ABST
Patent Text Reader

Abstract

The invention discloses a pathogenic gene prediction method and device based on phenotypic fingerprints, equipment and a medium. The method is executed by a computer, systematic integration and quantification are carried out on associated information between genes and phenotypes of multiple dimensions, phenotype fingerprints with specific genes are constructed on the group level, and complex effects of the genes on different phenotype dimensions can be captured more comprehensively; a multi-phenotype score value taking genes as the center is calculated through gene phenotype fingerprints, that is, multi-dimensional clinical phenotype information of a target object is converted into quantitative scores taking the genes as the center on the object level, and two types of key output of pathogenic variation carrying risk assessment and candidate gene priority ranking are achieved through observation phenotypes of the target object; and the integrating degree of each candidate gene and the actual phenotype of the target object can be objectively and efficiently evaluated. According to the method, phenotype fingerprints are introduced, a multi-phenotype scoring mechanism is combined, the efficiency and objectivity of complex disease pathogenic gene recognition are remarkably improved, and the method has important clinical application value.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of biomedical informatics, and in particular to a pathogenic gene prediction method and device based on a phenotype fingerprint, equipment and a medium. BACKGROUND

[0002] In recent years, with the rapid development of high-throughput sequencing technology, a large amount of human genetic variation data has been accumulated exponentially. For complex diseases (such as neurodevelopmental disorders), the phenotype spectrum corresponding to a single pathogenic gene often presents incomplete penetrance, variable expression, and significant individual differences. At the same time, different genes have different influence patterns on different phenotype dimensions, making it difficult to directly establish an explanation chain from variation to phenotype.

[0003] Existing clinical workflows usually rely on the following methods to screen and sort candidate genes: 1. screening genes according to existing gene-disease knowledge bases or guidelines; 2. grading candidate variations according to variation pathogenicity evidence; 3. determining possible pathogenic genes and forming a diagnosis conclusion by combining patient phenotypes and comprehensive judgments of previous reports.

[0004] However, in the context of complex diseases, due to incomplete phenotype collection, inconsistent phenotype description, and insufficient knowledge base coverage, the above processes often face problems such as excessive candidate genes, unclear priority, low explanation efficiency, and consistent results relying on expert experience. Especially when a patient presents a combination of multiple systems and multiple dimensions of phenotypes, how to identify the most likely pathogenic gene from a large number of candidate genes in a short time has become a problem to be solved. SUMMARY

[0005] The present application aims to at least solve the technical problems existing in the prior art. To this end, the present application provides a pathogenic gene prediction method and device based on a phenotype fingerprint, which constructs a gene-specific phenotype fingerprint at the population level and converts multi-dimensional clinical phenotype information into a gene-centered quantitative score at the object level, thereby achieving two key outputs of pathogenic variation carrying risk assessment and candidate pathogenic gene priority sorting, significantly improving the efficiency and objectivity of complex disease pathogenic gene identification, reducing the dependence on expert experience, and having important clinical application value.

[0006] To achieve the above purpose, a first aspect of an embodiment of the present application provides a pathogenic gene prediction method based on a phenotype fingerprint, the method comprising: constructing a plurality of gene-phenotype association pairs based on a plurality of phenotypes and a plurality of genes related to a target disease of at least one subject, and determining association information of each of the genes in all phenotype dimensions based on the plurality of gene-phenotype association pairs, to construct a phenotype fingerprint of each of the genes according to the association information; wherein any one of the gene-phenotype association pairs consists of any one of the phenotypes and any one of the pathogenic genes, and any two of the gene-phenotype association pairs are different; each of the genes includes at least one pathogenic variant or a possible pathogenic variant; in response to a pathogenic gene prediction instruction of a target object, determining an observed phenotype of the target object related to the target disease; according to the phenotype fingerprint and the observed phenotype, calculating a multi-phenotype score value of each candidate gene of the target object with the gene as the center, to predict a risk value of the target object carrying a corresponding pathogenic variant and determine a priority of each of the candidate genes according to the multi-phenotype score value of each of the candidate genes.

[0007] The first aspect of the embodiment of the present application provides a pathogenic gene prediction method based on a phenotype fingerprint, which has at least the following beneficial effects: The method can construct a gene-specific phenotype fingerprint at a population level by constructing a phenotype fingerprint of each gene and systematically integrating and quantifying the association information of the gene in multiple phenotype dimensions, which is different from the one-to-one association of the gene and the disease in the prior art. The method can more comprehensively capture the complex effects of the gene in different phenotype dimensions and provide more detailed and rich background information for subsequent individualized prediction. Moreover, the method calculates a multi-phenotype score value with the gene as the center, i.e., converts the multi-dimensional clinical phenotype information of the target object into a quantitative score with the gene as the center at the object level, and realizes two key outputs of pathogenic variant carrying risk assessment and candidate pathogenic gene priority ranking by using the observed phenotype of the target object. The scoring mechanism based on data driving can objectively and efficiently evaluate the degree of fit of each candidate gene and the actual phenotype of the target object. The method provides a systematic, quantitative and efficient pathogenic gene prediction scheme by introducing the concept of phenotype fingerprint and combining the multi-phenotype scoring mechanism, significantly improves the efficiency and objectivity of complex disease pathogenic gene identification, reduces the dependence on expert experience, and has important clinical application value.

[0008] To achieve the above object, the second aspect of the embodiment of the present application provides a pathogenic gene prediction device based on a phenotype fingerprint, which comprises: a phenotype fingerprint constructing module configured to construct a plurality of gene-phenotype association pairs based on a plurality of phenotypes and a plurality of genes associated with a target disease of at least one subject, and determine association information of each of the genes in all phenotype dimensions based on the plurality of gene-phenotype association pairs, so as to construct a phenotype fingerprint of each of the genes according to the association information; wherein any one of the gene-phenotype association pairs consists of any one of the phenotypes and any one of the pathogenic genes, and any two of the gene-phenotype association pairs are different; each of the genes comprises at least one pathogenic variant or a possible pathogenic variant; a prediction instruction response module configured to determine an observed phenotype of a target object associated with the target disease in response to a pathogenic gene prediction instruction of the target object; a risk prediction and ranking module configured to calculate a multi-phenotype score value of each of the candidate genes of the target object based on the phenotype fingerprint and the observed phenotype, so as to predict a risk value of the target object carrying a corresponding pathogenic variant and determine a priority of each of the candidate genes according to the multi-phenotype score value of each of the candidate genes.

[0009] To achieve the above object, a third aspect of the embodiments of the present application provides an electronic device, comprising at least one control processor and a memory connected in communication with the at least one control processor; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor to enable the at least one control processor to execute the pathogenic gene prediction method based on a phenotype fingerprint according to the first aspect.

[0010] To achieve the above object, a fourth aspect of the embodiments of the present application provides a computer readable storage medium, which stores computer executable instructions for causing a computer to execute the pathogenic gene prediction method based on a phenotype fingerprint according to the first aspect. BRIEF DESCRIPTION OF DRAWINGS

[0011] Figure 1 is a flowchart of the pathogenic gene prediction method based on a phenotype fingerprint provided by the embodiments of the present application; Figure 2 is a sample screening flowchart provided by the embodiments of the present application; Figure 3 is a schematic diagram of single high-confidence gene level evaluation provided by the embodiments of the present application; Figure 4 is a schematic diagram of single extended gene level evaluation provided by the embodiments of the present application; Figure 5 is a schematic diagram of multi-gene level evaluation provided by the embodiments of the present application; Figure 6 is a schematic diagram of an independent external verification queue provided by an embodiment of the present application; Figure 7 is a structural schematic diagram of a pathogenic gene prediction device based on a phenotype fingerprint provided by an embodiment of the present application; Figure 8 is a structural schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0012] Embodiments of the present application are described in detail below with reference to examples shown in the accompanying drawings, in which the same or similar numerals represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application and cannot be understood as limiting the present application.

[0013] In the description of the present application, if there is a description of first, second, etc., it is only for the purpose of distinguishing technical features, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features or the sequence of indicated technical features.

[0014] As shown in Figure 1 An embodiment of the present application provides a pathogenic gene prediction method based on a phenotype fingerprint, which comprises the following steps: Step S110, based on a plurality of phenotypes and a plurality of genes related to a target disease of at least one subject, a plurality of gene-phenotype association pairs are constructed, and based on the plurality of gene-phenotype association pairs, association information of each gene in all phenotype dimensions is determined, and a phenotype fingerprint of each gene is constructed according to the association information.

[0015] Step S120, in response to a pathogenic gene prediction instruction of a target object, determining the observed phenotype of the target object related to the target disease.

[0016] Step S130, according to the phenotype fingerprint and the observed phenotype, calculating a multi-phenotype score value of each candidate gene of the target object with the gene as the center, and predicting the risk value of the target object carrying the corresponding pathogenic variation and determining the priority of each candidate gene according to the multi-phenotype score value of each candidate gene.

[0017] For ease of understanding, some key terms in the present embodiment are explained as follows: The subject refers to the data source individual for constructing the phenotype fingerprint, which is usually a patient population with a specific disease phenotype, whose genome and phenotype data have been collected and analyzed. It should be noted that the specific target disease is not limited here, which can be a neurodevelopmental disorder, a mental disease and other multifactorial diseases.

[0018] Phenotype refers to the observable traits of an organism, including morphological, physiological, biochemical, and behavioral characteristics. In this embodiment, phenotype can manifest as disease symptoms, clinical indicators, imaging features, etc.

[0019] Gene refers to a DNA segment that carries genetic information, which is the basic unit of heredity in an organism. Genes affect the traits of an organism by encoding proteins or other functional molecules. Among them, pathogenic variants or potentially pathogenic variants refer to DNA sequence changes in the genome that can cause disease or increase the risk of disease. Pathogenic variants are variants that have been clearly proven to be associated with disease, while potentially pathogenic variants are variants that are highly suspected to be associated with disease based on existing evidence. The classification of pathogenic variants and potentially pathogenic variants in genes is described in subsequent embodiments and will not be described here.

[0020] Target object refers to an individual who needs to be predicted for pathogenic genes, and the observed phenotype of the target object will be used to assess the risk of carrying pathogenic variants. Observed phenotype refers to the phenotype data related to the target disease actually measured or observed from the target object.

[0021] Multi-phenotype score refers to a quantitative score calculated by considering the association information of a candidate gene in multiple phenotype dimensions and the observed phenotype of the target object. Priority refers to the result of ranking multiple candidate genes, reflecting the likelihood of each candidate gene as a pathogenic gene, which helps clinicians quickly focus on the most relevant genes.

[0022] The present embodiment provides a pathogenic gene prediction method based on phenotype fingerprint, which comprises the following main features: The core point of the present application is to construct gene-specific phenotype fingerprints at the population level, and to convert multi-dimensional clinical phenotype information into quantitative scores centered on genes at the object level, thereby achieving two key outputs of pathogenic variant risk assessment and candidate pathogenic gene priority ranking.

[0023] In step S110 of the present embodiment, data is first collected from the subject, including gene data and phenotype data. Taking autism spectrum disorder (ASD) as an example, phenotype data (such as self-feeding with a spoon, Tourette's syndrome, epilepsy, etc.) is collected from multiple ASD patients. It should be noted that phenotype can be classified into binary phenotype (such as Tourette's syndrome, epilepsy) and continuous phenotype (such as self-feeding with a spoon). At the same time, the corresponding gene data is collected. It should be noted that the gene here includes at least one pathogenic variant or potentially pathogenic variant. The classification of gene variants is described in subsequent embodiments and will not be described here.

[0024] In constructing multiple gene-phenotype pairs, one implementation is to manually review literature and expert experience to organize known gene-phenotype relationships to form gene-phenotype pairs. For example, for a genetic disease, the gene mutation information and various clinical symptoms exhibited by patients can be manually collected, and each gene can be paired with each symptom. Another implementation is to extract gene and phenotype data from large biomedical databases and perform preliminary screening to pair each gene with each phenotype to form a set of original gene-phenotype pairs.

[0025] In determining the association information of each gene in all phenotype dimensions and constructing the phenotype fingerprint, one implementation is that for each gene-phenotype pair, a statistical method such as calculating a correlation coefficient or performing a chi-square test can be used to evaluate the strength and direction of the association between the gene and the phenotype, and then these statistical results can be directly used as association information and combined to form the phenotype fingerprint of the gene. Another implementation is to describe the relationship between each gene and all related phenotypes by manual analysis or based on pre-set rules, such as “strong association”, “weak association”, “no association”, etc., and then encode these descriptions as numerical or symbolic values as association information to construct the phenotype fingerprint of the gene.

[0026] In step S120 of the embodiment, in response to the pathogenic gene prediction instruction of the target object, one implementation is that the system receives a manual input instruction from a user or a clinician, which explicitly specifies the target object for which pathogenic gene prediction is required, for example, the doctor selects a patient in the diagnosis system and clicks start prediction. Another implementation is that the system automatically scans the queue of target objects to be processed through a timing task or a batch processing mechanism, and once a new prediction request is detected, the prediction process is automatically triggered, for example, when new patient genomic and phenotype data are uploaded to the database, the system automatically identifies and starts the prediction. In determining the observed phenotypes related to the target object and the target disease, one implementation is to manually enter the phenotypic information in paper or electronic documents such as clinical diagnosis reports and physical examination results of the target object by medical staff or data administrators, for example, the doctor enters the patient's height, weight, blood pressure, and specific symptoms according to the patient's medical record; another implementation is that the system directly obtains the structured phenotype data of the target object from the electronic medical record system or the laboratory information management system through a pre-set data interface, for example, the system can automatically extract the recorded disease symptoms and biochemical indicators of the patient from the patient's electronic health record.

[0027] In an implementation, any gene in step S110 can be a candidate gene, and the candidate gene can also be a gene with multiple unknown significance variations detected by sequencing. In step S130 of calculating the multi-phenotype score value of each candidate gene of the target object, one implementation is that, for each candidate gene, a part corresponding to the observed phenotype of the target object in the phenotype fingerprint of the candidate gene is calculated to obtain a score value. Another implementation is that, the correlation information in the phenotype fingerprint is logically judged with the observed phenotype of the target object, for example, if the gene is associated with a certain phenotype and the target object observes the phenotype, a fixed score is given; otherwise, no score is given, and then the scores of all phenotypes are added to form the multi-phenotype score value.

[0028] In predicting the risk value of the target object carrying the corresponding pathogenic variation according to the multi-phenotype score value of each candidate gene and determining the priority of each candidate gene, one implementation is that the calculated multi-phenotype score value (0-1) is directly taken as the risk value, and the candidate genes are sorted according to the score value from high to low to determine the priority, for example, the higher the score value, the greater the risk value, and the higher the priority. Another implementation is to set a fixed threshold, if the multi-phenotype score value of the candidate gene exceeds the threshold, it is considered that the risk of the target object carrying the corresponding pathogenic variation is high, and it is marked as high priority; otherwise, it is marked as low priority.

[0029] The method can construct a gene-specific phenotype fingerprint at the population level by constructing a phenotype fingerprint of each gene, systematically integrating and quantifying the correlation information of the gene in multiple phenotype dimensions, which is different from the one-to-one association of the gene and the disease in the prior art. The method can more comprehensively capture the complex effects of the gene in different phenotype dimensions, and provide more detailed and rich background information for subsequent individualized prediction. Moreover, the method calculates the multi-phenotype score value centered on the gene, quantitatively matches the observed phenotype of the target object with the pre-constructed phenotype fingerprint of the gene, that is, converts the multi-dimensional clinical phenotype information of the target object into a quantitative score centered on the gene at the object level, realizes the two key outputs of pathogenic variation carrying risk assessment and candidate pathogenic gene priority sorting. This data-driven scoring mechanism avoids excessive dependence on single expert experience, and can objectively and efficiently evaluate the fitness of each candidate gene and the actual phenotype of the target object. The method provides a systematic, quantitative and efficient pathogenic gene prediction scheme by introducing the concept of phenotype fingerprint and combining the multi-phenotype scoring mechanism, significantly improves the efficiency and objectivity of complex disease pathogenic gene identification, reduces the dependence on expert experience, and has important clinical application value.

[0030] In some embodiments of the present application, the determination of the association information of each gene in all phenotype dimensions based on the plurality of gene-phenotype association pairs in step S110 comprises: Step S210, performing independent association test on each of the plurality of gene-phenotype association pairs to obtain an association test result.

[0031] Step S220, drawing an association map of the plurality of gene-phenotype association pairs according to the association test result.

[0032] Step S230, determining the association direction, effect size and significance information of each gene in all phenotype dimensions according to the association map.

[0033] Step S240, constructing a phenotype fingerprint of each gene according to the association direction, effect size and significance information.

[0034] In the present embodiment, independent association test is performed on each of the plurality of gene-phenotype association pairs, aiming to evaluate whether there is a statistically significant association between each gene and a specific phenotype. By performing separate statistical analysis on each gene-phenotype association pair, the strength and reliability of the correlation between them can be quantified. For example, chi-square test can be used to evaluate the association between categorical variables (such as genes and phenotypes), or logistic regression or linear regression models can be used to analyze the association between gene variants and phenotypes according to the type of phenotype (binary or continuous), and control potential confounding factors.

[0035] According to the association test result, an association map of the plurality of gene-phenotype association pairs is drawn, which is a visualization tool for intuitively showing the association patterns between genes and multiple phenotypes. It can help researchers quickly identify which genes are significantly associated with which phenotypes, as well as the strength and direction of these associations. For example, the association test results (such as P value, effect size, etc.) can be presented in the form of a heat map, where rows represent genes and columns represent phenotypes, and color intensity represents the strength of association; or a network graph can be constructed, with nodes representing genes and phenotypes, and edges representing the association between them, and the thickness or color of the edges reflecting the strength of the association.

[0036] According to the association map, the association direction, effect size and significance information of each gene on all phenotype dimensions are determined. This step extracts the key information required to construct the phenotype fingerprint from the association map. The association direction indicates the trend of the gene's influence on the phenotype (e.g., increasing or decreasing a certain phenotype characteristic), the effect size quantifies the size of this influence, and the significance information indicates the statistical reliability of this influence. For example, for a heat map, significant associations can be identified by color gradients and pre-set thresholds, and the corresponding effect size (e.g., odds ratio, regression coefficient) and P value are extracted from the original test results as significance information. For a network graph, the connection strength and direction between nodes can be analyzed, and the effect size and significance P value of each association can be directly obtained from the statistical report of the association test.

[0037] According to the association direction, effect size and significance information, the phenotype fingerprint of each gene is constructed. The phenotype fingerprint is a multi-dimensional phenotype association feature vector specific to each gene, which integrates the association direction, effect size and significance between the gene and all related phenotypes, forming a unique signature for subsequent pathogenic gene prediction. For example, the association direction, effect size and significance information (e.g., P value or corrected P value) of each gene on all phenotype dimensions can be combined into a vector or matrix as the phenotype fingerprint of the gene. Alternatively, the effect size and significance information can be weighted or transformed, e.g., multiplying the effect size by the inverse or logarithm of the significance level to highlight stronger and more reliable associations, forming a more discriminative phenotype fingerprint.

[0038] The present embodiment can accurately extract biologically meaningful association information from complex gene and phenotype data. By performing independent association tests, the present embodiment ensures the statistical reliability of each gene-phenotype association, avoiding false positive associations that interfere with subsequent prediction. The association map provides researchers with an intuitive and comprehensive view to quickly understand the complex relationships between genes and multiple phenotypes. Finally, the phenotype fingerprint constructed by integrating the association direction, effect size and significance information can more accurately capture the unique biological function of each gene and its role in disease development, significantly improving the accuracy and reliability of subsequent pathogenic gene prediction.

[0039] In some embodiments of the present application, the association test results include a test P value corresponding to each gene-phenotype association pair; In step S220, the association map of multiple gene-phenotype association pairs is drawn according to the association test results, including: In step S2210, a plurality of first gene-phenotype association pairs with a test P value less than a first threshold are selected from the plurality of gene-phenotype association pairs according to the test P value corresponding to each gene-phenotype association pair. The first threshold is used to represent the existence of an association between the gene and the phenotype in the gene-phenotype association pair.

[0040] Step S2230, drawing an association map according to the plurality of first gene-phenotype association pairs.

[0041] wherein the association test results include a test P-value corresponding to each gene-phenotype association pair. In this embodiment, for each gene-phenotype association pair, a P-value is calculated by an independent association test (e.g., chi-square test, t-test, regression analysis, etc.), which reflects the statistical significance of the association between the gene and the phenotype. The smaller the P-value, the less likely that the observed association is caused by random chance, i.e., the greater the likelihood that there is a true association between the gene and the phenotype.

[0042] Further, according to the test P-value corresponding to each gene-phenotype association pair, a plurality of first gene-phenotype association pairs with test P-values less than a first threshold are screened from the plurality of gene-phenotype association pairs. This step aims to identify a subset of statistically significant associations from all gene-phenotype association pairs. By setting a predefined first threshold, only when the test P-value of a gene-phenotype association pair is lower than the threshold, the association is considered to be worth attention. This screening mechanism can effectively remove gene-phenotype pairs with strong randomness and weak association, thereby focusing on more reliable association information. In addition to directly comparing P-values with thresholds, other screening strategies can also be used, such as screening based on effect size, or comprehensive evaluation combining P-values and effect size. The first threshold is used to characterize the association between the gene and the phenotype in the gene-phenotype association pair. The first threshold is a pre-set statistical significance level, for example, 0.05. When the test P-value of a gene-phenotype association pair is less than the first threshold, it is statistically considered that there is a non-random, meaningful association between the gene and the phenotype.

[0043] On this basis, an association map is drawn according to the plurality of first gene-phenotype association pairs. After screening the first gene-phenotype association pairs, only these significantly filtered association pairs are used to construct the association map, which ensures that the association information contained in the map is statistically verified, thereby improving the quality and reliability of the map. The association map can be presented in various forms, such as network graph, heat map, or scatter plot, etc., which aims to visually demonstrate the interaction pattern between genes and phenotypes.

[0044] The embodiment can effectively filter out statistically insignificant gene-phenotype associations when constructing the association map of the gene-phenotype association pairs, avoid introducing random or weak association information into the subsequent phenotype fingerprint construction process, significantly improve the purity and reliability of the association map, and make the extracted association direction, effect size and significance information more accurate. Therefore, the phenotype fingerprint of each gene constructed based on this can more accurately capture the real relationship between the gene and the disease-related phenotype, thereby providing a more solid and reliable data foundation for subsequent prediction of pathogenic genes, and further improving the accuracy and reliability of the prediction of pathogenic genes.

[0045] In some embodiments of the present application, before the association map is drawn according to the plurality of first gene-phenotype association pairs in step S2230, the method further comprises: Step S2221, updating the test P value based on multiple hypothesis testing; Step S2222, screening a plurality of second gene-phenotype association pairs from the plurality of first gene-phenotype association pairs according to the updated test P value; The step S130 of calculating the multi-phenotype score value of each candidate gene of the target object with the candidate gene as the center according to the phenotype fingerprint and the observed phenotype comprises: Step S310, extracting the phenotype in the first gene-phenotype association pair corresponding to the candidate gene from the phenotype fingerprint, and calculating the signal term value of the candidate gene according to the phenotype fingerprint and the observed phenotype when the observed phenotype corresponds to the extracted phenotype in the first gene-phenotype association pair; Step S320, extracting the phenotype in the first gene-phenotype association pair corresponding to the candidate gene from the phenotype fingerprint, and extracting the phenotype in the second gene-phenotype association pair corresponding to the candidate gene, and calculating the penalty term value of the candidate gene according to the phenotype fingerprint and the observed phenotype when the observed phenotype does not correspond to the extracted phenotype in the first gene-phenotype association pair or does not correspond to the extracted phenotype in the second gene-phenotype association pair; Step S330, calculating the multi-phenotype score value of each candidate gene with the candidate gene as the center according to the signal term value and the penalty term value of each candidate gene.

[0046] In this embodiment, the aim is to further correct the false positive risk due to multiple statistical tests by a more rigorous statistical method. Multiple hypothesis testing is a statistical procedure to control the overall error rate when multiple statistical tests are performed simultaneously, which serves to ensure that the screened second gene-phenotype associations have higher statistical significance and credibility, thus providing a more solid foundation for subsequent pathogenic gene prediction. One way to achieve this is to use the FDR (False Discovery Rate) algorithm to correct, which balances false positive control and statistical power by controlling the false discovery rate, resulting in a new P value, especially suitable for large-scale genomic data analysis.

[0047] In the case where the observed phenotype corresponds to the phenotype in the first gene-phenotype association pair selected in step S310, the signal term value of the candidate gene is calculated according to the phenotype fingerprint and the observed phenotype, which aims to quantify the observed evidence strength consistent with the known gene-phenotype association. The signal term value reflects the matching degree between the observed phenotype of the target object and the phenotype fingerprint of the candidate gene, providing positive support for prediction. One way to achieve this is to multiply the effect size of each corresponding observed phenotype (e.g., the regression coefficient or odds ratio from the phenotype fingerprint) by the confidence weight of that phenotype, and then accumulate the products of all corresponding phenotypes; another way is to construct a linear model based on the phenotype fingerprint, taking the observed phenotype as input and directly outputting the signal term value, where the parameters of the model reflect the strength of the phenotype-gene association.

[0048] In the case where the observed phenotype does not correspond to the phenotype in the first gene-phenotype association pair extracted in step S320 or does not correspond to the phenotype in the second gene-phenotype association pair, the penalty term value of the candidate gene is calculated according to the phenotype fingerprint and the observed phenotype, which aims to negatively affect the observed phenotypes that do not match or are uncertain about the expected phenotype fingerprint of the candidate gene. The penalty term value aims to reduce the score of candidate genes whose phenotype performance does not match the gene association pattern, thus improving the specificity of the prediction. One way to achieve this is to multiply a pre-set penalty hyperparameter by the degree of deviation of each non-corresponding observed phenotype from the phenotype fingerprint of the gene, and then accumulate these penalty values; another way is to calculate the penalty term according to the difference between the universality of the observed phenotype in the population and the expected phenotype fingerprint of the gene, where rare but inconsistent phenotypes with the gene may result in greater penalties.

[0049] Based on the signal and penalty values ​​of each candidate gene, a gene-centric multiphenotypic score is calculated for each candidate gene. The aim is to integrate positive support (signal value) and negative influence (penalty value) into a comprehensive score to fully assess the pathogenicity potential of the candidate gene. This score more accurately reflects the degree of agreement between the overall phenotypic characteristics of the target population and the specific gene's pathogenic pattern. One approach is to directly subtract the penalty value from the signal value to obtain a preliminary comprehensive score. Another approach is to standardize the signal and penalty values ​​before performing a weighted sum or combining them using a nonlinear function to generate the final multiphenotypic score.

[0050] This embodiment introduces multiple hypothesis testing to conduct a second rigorous screening of the initially selected first gene-phenotype association pairs, thereby obtaining second gene-phenotype association pairs with greater statistical significance. This process significantly improves the reliability of gene-phenotype associations, laying a more solid foundation for subsequent pathogenic gene prediction. Based on this, when responding to the pathogenic gene prediction instruction for the target object and determining its observed phenotype, for observed phenotypes that match the first gene-phenotype association pair, this embodiment calculates a positive signal term value to quantify its support for the pathogenicity of the candidate gene; while for those phenotypes that match the first gene-phenotype association pair... For discrepancies, or more strictly, for discrepancies in association with the second gene phenotype selected through multiple hypothesis testing, the system calculates a penalty value. This penalty mechanism effectively reduces the risk of misjudgment due to uncertain or inconsistent phenotypes. Ultimately, by integrating the signal value and the penalty value, this embodiment can calculate a comprehensive, gene-centric multiphenotype score for each candidate gene. This score not only considers evidence of consistency with the gene's expected phenotype fingerprint but also fully considers evidence of discrepancies or absence, thus providing a more comprehensive and accurate risk assessment of pathogenic genes.

[0051] In some embodiments of this application, when the observed phenotype is a continuous phenotype, the process of calculating the signal term value of the candidate gene in step S310 includes: (1); in, Represents the target object For candidate genes The signal term values ​​of the continuous phenotype, It represents the set of continuous phenotypes in the first gene phenotype association pairs corresponding to the candidate gene (obtained from phenotypic fingerprints). Represents the target object Observational phenotypes Phenotypic values, This indicates the observed phenotype determined based on phenotypic fingerprints. the effect size (obtained from the phenotype fingerprint) of the observed phenotype, the confidence weight of the observed phenotype ; wherein, for the phenotype corresponding to the second gene-phenotype association pair, = 1; for the phenotype corresponding to the first gene-phenotype association pair, is a hyper-parameter between 0 and 1, which is obtained by deriving a queue cross-validation.

[0052] In the case of binary phenotype, the process of calculating the signal term value of the candidate gene in step S310 includes: (2); wherein, the observed phenotype of the target object the signal term value of the binary phenotype of the candidate gene , wherein, represents a set composed of the binary phenotype in the first gene-phenotype association pair corresponding to the candidate gene (obtained from the phenotype fingerprint), the observed phenotype of the target object the phenotype value of the observed phenotype , wherein, the effect size (obtained from the phenotype fingerprint) of the observed phenotype , wherein, represents a natural logarithm function, the confidence weight of the observed phenotype .

[0053] wherein, the continuous phenotype is usually obtained by measurement, has a continuous numerical range, and can provide more fine individual difference information, which helps to capture the small changes of gene effect. Binary phenotype refers to the phenotype whose value has only two possible states, usually represented as "yes / no", "existence / absence", "disease / health", etc.

[0054] Formula (1) provides a mathematical model for quantifying the contribution of continuous observed phenotype to the pathogenicity of candidate genes. It realizes the accurate calculation of signal strength by combining the observed phenotype value of the target object with the effect size and confidence weight determined in the gene phenotype fingerprint. The calculation of formula (1) can be realized by programming language (such as Python), using matrix operation or loop iteration to sum all continuous phenotypes in the set In practical application, high-performance computing library (such as NumPy) can be used for optimization to improve the calculation efficiency.

[0055] Formula (2) provides a mathematical model for quantifying the contribution of binary observed phenotypes to the pathogenicity of candidate genes. It achieves accurate calculation of signal intensity by combining the binary observed phenotype values ​​of the target object with the effect size and confidence weight determined in the gene phenotype fingerprint. Similar to the calculation of continuous phenotype signal terms, Formula (2) can also be implemented using a programming language.

[0056] This collection includes specific candidate genes. All continuous phenotypes identified as associated in the first gene-phenotype association pair (with a p-value less than the first threshold) are considered to have a potential association with the gene. This limits the range of phenotypes to be considered when calculating the continuous signal term value, ensuring that only continuous phenotypes with statistical association with the gene are included. During data preprocessing, this set can be dynamically constructed based on the type of gene-phenotype association pair and the association test results. Target object In specific continuous observation phenotypes The actual measured values, used as input for signal term calculation, directly reflect the individual phenotypic characteristics of the target object. These values ​​can be numerical data obtained from clinical records, physical examination reports, or biological sample analysis. Quantifying genes and continuous phenotypes The strength and direction of the association between genotypes can be considered; for example, in regression analysis, it might be the regression coefficient of the influence of genotype on phenotypic value. In signal term calculations, the effect size, as a weight, reflects the importance of different phenotypes in contributing to gene pathogenicity. It can be estimated from large-scale population data using statistical methods such as linear regression and generalized linear models. Reflects the observed phenotype The reliability or importance of gene association can be assigned, for example, based on the p-value of the association test, the standard error of the effect size, or expert knowledge. It weights the contributions of different phenotypes in the signal term calculation, giving more reliable or more important phenotypes greater influence. This can be obtained by inverse or exponential transformation of the p-value of the association test; it can also be manually set based on the clinical importance of the phenotype in disease diagnosis; or optimized through machine learning methods such as cross-validation.

[0057] This collection includes specific candidate genes. All binary phenotypes identified as associated in the first gene-phenotype association pair. This limits the range of phenotypes to be considered when calculating binary signal term values, ensuring that only binary phenotypes statistically associated with the gene are included. Similarly, it is built through data preprocessing and filtering operations. Target object In specific binary observation phenotypes The actual state of the signal is usually represented as 0 or 1. As input to the signal term calculation, it directly reflects the individual binary phenotypic characteristics of the target object. It can be a Boolean value obtained from clinical diagnosis, medical history records, or genetic testing results. The Odds Ratio quantifies the relationship between genes and binary phenotypes. The strength of the association between two factors. It represents the ratio of the probability of a specific event (such as developing the disease) occurring in the exposed group (e.g., carrying a pathogenic gene variant) to the probability of the same event occurring in the unexposed group (e.g., not carrying a pathogenic gene variant). In the calculation of the binary phenotype signal term, the natural logarithm of the dominance ratio serves as the weight, reflecting the degree of influence of the gene on the risk of a binary phenotype. It can be estimated from large-scale population data using statistical methods such as logistic regression and chi-square test. It is a constant A logarithmic function with base 0. Its dominance ratio is... Converting to a logarithmic scale makes it mathematically more manageable and allows positive and negative correlations to be represented symmetrically. Reflects the observed phenotype The reliability or importance of gene association, its weighted contribution to the calculation of the signal term for different binary phenotypes, and its confidence weight compared to the continuous phenotype. Similarly, values ​​can be assigned based on the p-value of the correlation test, the standard error of the effect size, or expert knowledge.

[0058] This embodiment provides a precise signal term calculation method for different types of observed phenotypes (such as continuous and binary phenotypes). By combining the observed phenotype values ​​of the target object with the effect size determined in the gene phenotype fingerprint (such as the natural logarithm of the regression coefficient or the odds ratio) and confidence weights, this embodiment can accurately quantify the contribution of each observed phenotype to the pathogenicity of candidate genes. This meticulous differentiation and quantification avoids the errors that may be caused by uniformly processing different types of phenotypes, thereby significantly improving the accuracy and reliability of pathogenic gene prediction and making the prediction results more consistent with biological reality.

[0059] In some embodiments of this application, when the observed phenotype is a continuous phenotype, the process of calculating the penalty term value of the candidate gene includes: (3); in, Represents the target object For candidate genes The penalty term value for continuous phenotypes, This represents the set of phenotypes from the first gene-phenotype association pairs corresponding to the candidate gene, or the set of phenotypes from the second gene-phenotype association pairs corresponding to the candidate gene. the observed phenotype of the target object the observed phenotype of the target object the observed phenotype of the target object the observed phenotype of the target object the observed phenotype of the target object penalty hyperparameter (0~1, obtained by stack-to-queue cross-validation).

[0060] In the case of binary phenotype, the process of calculating the penalty term value of the candidate gene includes: (4); wherein, the observed phenotype of the target object the observed phenotype of the target object the observed phenotype of the target object the observed phenotype of the target object the observed phenotype of the target object

[0061] wherein, the penalty term value of the continuous phenotype the observed phenotype of the target object the observed phenotype of the target object the observed phenotype of the target object does not belong to the phenotype set significantly associated with the gene the observed phenotype of the target object penalty hyperparameter (0~1, obtained by stack-to-queue cross-validation). penalty hyperparameter (0~1, obtained by stack-to-queue cross-validation). the observed phenotype of the target object the observed phenotype of the target object the observed phenotype of the target object the observed phenotype of the target object the observed phenotype of the target object the observed phenotype of the target object the observed phenotype of the target object The mean value in the reference population or training dataset. By comparing the observed phenotype value of the target subject with the mean value, the deviation of the phenotype in the target subject can be assessed.

[0062] The penalty term value for a binary phenotype The observed phenotype of the target subject The candidate gene The penalty term value calculated when the observed phenotype is binary. This value is used to quantify the negative impact of the phenotype on the gene’s pathogenicity risk prediction when the observed binary phenotype does not belong to the set of phenotypes significantly associated with the gene , and the target subject exhibits the binary phenotype (i.e. = 1). The observed phenotype mean represents the average occurrence frequency or prevalence of the observed phenotype in the reference population or training dataset, which is usually a probability value between 0 and 1 for binary phenotypes.

[0063] In the above method, to more comprehensively assess the pathogenicity risk of the candidate gene, this embodiment further introduces the calculation of the penalty term. When the observed phenotype of the target subject does not correspond to the associated phenotype that has been determined in the phenotype fingerprint of the gene, i.e., the observed phenotype does not belong to the set of phenotypes significantly associated with the gene , a penalty needs to be imposed on the score of the gene. For continuous observed phenotypes, the calculation of the penalty term is based on the difference between the observed phenotype value and the mean value of the phenotype in the population. The accumulation of this difference, multiplied by the penalty hyperparameter , constitutes the penalty term value for the continuous phenotype. The principle is that if a continuous phenotype not significantly associated with the gene exhibits a significant deviation from its mean value in the target subject, it may mean that the pathogenicity of the gene is not reflected through this phenotype, or the association of the phenotype with the gene is weak, so appropriate punishment is needed to avoid its excessive contribution to the signal term. For binary observed phenotypes, the calculation of the penalty term focuses on those that do not belong to and the binary phenotype actually exhibited by the target subject ( = 1), at this time, the penalty term is obtained by accumulating multiplied by the penalty hyperparameter , where is the average occurrence frequency of the binary phenotype in the population. The principle is that if a binary phenotype not significantly associated with the gene appears in the target subject, and the phenotype itself is relatively rare in the population (i.e. is small, leading to If the number of observed phenotypes is large, then this occurrence can be more likely to be considered as noise or non-specific manifestation, thus imposing a greater penalty on the score of the gene. By introducing these two types of penalty terms, the scheme of the present application can effectively compensate for the deficiency of relying only on the signal term. It not only considers the phenotypic information significantly associated with the gene, but also quantitatively processes those observed phenotypes that are weakly or not associated with the gene. This processing mechanism enables the final gene-centered multi-phenotype score value to more accurately reflect the true pathogenic risk of the candidate gene, avoiding the interference of non-specific phenotypes or weakly associated phenotypes on the score, thereby improving the accuracy and reliability of pathogenic gene prediction.

[0064] The present embodiment can quantitatively process those observed phenotypes that are weakly or not associated with the candidate gene and include them as penalty terms in the calculation of the gene-centered multi-phenotype score value, which effectively avoids the evaluation bias that can be caused by relying only on the signal term, enabling the prediction model to more comprehensively consider the phenotypic characteristics of the target object. Specifically, by distinguishing between continuous and binary phenotypes and using different penalty calculation formulas, the present embodiment can more finely capture the potential impact of non-associated phenotypes on the pathogenic risk of the gene, which significantly improves the accuracy and robustness of pathogenic gene prediction, making the prioritization of candidate genes more reliable, thereby providing more accurate basis for clinical diagnosis and personalized treatment.

[0065] In some embodiments of the present application, the multi-phenotype score value of each candidate gene is calculated based on the signal term value and the penalty term value of each candidate gene, including: (5); (6); (7); (8); wherein, represents the signal term value of the continuous phenotype of the target object to the candidate gene, represents the penalty term value of the continuous phenotype of the target object to the candidate gene, represents the signal term value of the binary phenotype of the target object to the candidate gene, represents the penalty term value of the binary phenotype of the target object to the candidate gene, represents Z-score standardization, represents natural base, represents the gene-centered multi-phenotype score value of the candidate gene.

[0066] Specifically, formula (5) aims to integrate the signal term value of continuous phenotype and the corresponding penalty term value to obtain a comprehensive continuous phenotype score. The signal term value reflects the strength of evidence of the observed phenotype associated with the gene, while the penalty term value considers the negative impact of unassociated phenotypes on the score. Through subtraction operation, the positive association evidence and negative non-association evidence can be effectively balanced, making the final score more robust. Similarly, formula (6) is used to integrate the signal term value and the penalty term value of the binary phenotype to obtain a comprehensive binary phenotype score. Binary phenotype usually represents the presence or absence of a certain trait, and the calculation method of its signal term and penalty term is different from that of continuous phenotype, but the integration purpose is the same, that is, through subtraction operation, the positive association evidence and negative non-association evidence are balanced to improve the accuracy of the score, for example, direct subtraction can be used, or weighted average is subtracted after weighted average, or nonlinear function is combined.

[0067] Formula (7) aims to standardize and combine the adjusted continuous phenotype score and binary phenotype score to obtain a unified comprehensive score. Z represents Z-score standardization, which is a commonly used data standardization method that can convert data of different dimensions and distributions to a unified scale, eliminating the influence of dimension difference on score combination. By averaging the two standardized scores, it can ensure that continuous and binary phenotypes have similar weights in the final comprehensive score, for example, in addition to simple average, weighted average or other statistical methods such as principal component analysis can be used to combine the two scores. Z-score standardization can convert raw data into standard scores. Through Z-score standardization, data of different distributions can be converted to a standard normal distribution with a mean of 0 and a standard deviation of 1, thereby eliminating dimension and dimension difference, so that different types of scores can be effectively compared and combined.

[0068] Formula (8) converts the comprehensive score Score into a probability value between 0 and 1, that is, the gene-centered multi-phenotype score value (GPS). e represents the natural base, which is an important mathematical constant, approximately equal to 2.71828. It converts the comprehensive score Score into the risk probability of the target object carrying the corresponding pathogenic variation, making the score result more biologically interpretable.

[0069] The embodiment first integrates the signal term value and the penalty term value of the continuous type and the binary type phenotype to obtain an adjusted score, thereby ensuring that the net association evidence of each phenotype type is accurately reflected, then the adjusted scores are Z-score standardized to eliminate the differences in dimensions and distributions between different phenotype types, so that they can be compared and combined on a unified scale, then the comprehensive score representing the overall association strength of the gene in all related phenotype dimensions is obtained by averaging the standardized scores, finally, the Sigmoid function is used to convert the comprehensive score into a GPS value between 0 and 1, which can be directly explained as the risk probability of the target object carrying the corresponding pathogenic variant. The combination of these steps enables the effective integration of different types of phenotype information and the conversion into a unified, standardized and biologically meaningful risk prediction index, thereby solving the problem of how to effectively integrate different types of score items and convert them into an interpretable pathogenic gene prediction score.

[0070] The embodiment can effectively integrate different types of phenotype data (continuous and binary) and their corresponding signal terms and penalty terms, and generate a unified, interpretable gene-centric multi-phenotype score value through standardization and probability conversion, which makes the prediction result of pathogenic genes more biologically meaningful and clinically valuable, and can more accurately assess the risk of the target object carrying pathogenic variants, thereby providing strong support for disease diagnosis, risk assessment and personalized treatment.

[0071] As Figures 2 to 6 , for ease of understanding, the present application provides the following embodiments, which take autism spectrum disorder (ASD) as an example, and ASD has highly heterogeneous clinical manifestations.

[0072] The embodiment can be used to assist genetic diagnosis and genetic counseling in complex disease scenarios. The core logic of the embodiment is to construct a gene-specific phenotype fingerprint at the population level, and to convert multi-dimensional clinical phenotype information into a gene-centric quantitative score (referred to as GPS (Gene-centric Polyphenotypic Score) in the embodiment) at the object level, thereby realizing the risk assessment of the target object as a pathogenic variant carrier and the priority ranking of candidate genes in the same process, and providing interpretable data support for clinical genetic variant interpretation and gene positioning.

[0073] Step S910, sample screening; Input 142357 ASD samples with exome sequencing data (ES), filtered according to sample phenotype data integrity, sequencing data quality, consistency of kinship, and confidence of disease diagnosis, and finally obtained 131815 samples, of which 44962 were ASD patients and 86853 were non-ASD patients.

[0074] The clinical scale answer information of the subjects (ASD children), the development history record and the comorbidity / concomitant disease related investigation information are taken as phenotypes. In the selection of phenotypes, the ASD core symptom dimensions (social communication and social interaction defects, repetitive behavior) emphasized by DSM-5 and the evidence of previous research are combined, and the ASD clinical core characteristics and common comorbidities are preferentially included, and all phenotypes are standardized, sorted and classified.

[0075] The ASD core characteristics are evaluated by four standardized scales, namely SCQ (social communication scale), RBSR (repetitive stereotyped behavior scale), DCDQ (developmental coordination disorder scale) and ABC (adaptive behavior scale), and the obtained data are the scores of each scale item, which are standardized to the total score of the scale and the score of each sub-scale; and at the same time, the total score and the sub-score are included, and a total of 20 scale related phenotypes are formed, which are used to quantify the symptoms and functional performance of different dimensions. In addition to the core scale, 10 developmental milestones are further included in this embodiment to characterize the early development trajectory (for example, whether to reach or reach the time point, unified coding according to data availability), and five types of comorbidities are included, including mental and psychological (n=19), neural development related (n=10), perinatal complications (n=5), birth defects (n=6) and abnormal growth and development (n=5). Finally, a standardized phenotype panel containing 7 categories and 75 phenotypes is formed.

[0076] In this embodiment, the phenotypes are clearly divided into two types of data: continuous phenotypes, including scale total score / sub-scale score and quantitative data of developmental milestones; binary phenotypes, including five types of phenotypes such as mental and psychological, neural development related, perinatal complications, birth defects and abnormal growth and development, which are uniformly coded as "yes / no" two types of state variables. The specific phenotypes are as follows: Table 1

[0077] Step S920, variant data processing; Step S9210, variant data quality control: The genotypes (VCF format) of the screened 44962 ASD patients are quality controlled using bcftools and PLINK tools according to the following standards: GQ<30, DP<7 (for SNP type variants); DP < 10 (for InDel type of variation); AB < 0.15 (for SNP type of variation); AB < 0.2 (for InDel type of variation); callrate < 0.9, Hardy-Weinberg equilibrium test p-value < 1 x The final output is a VCF format of the variation file after quality control.

[0078] Step S9220, variation annotation; For the variation data after quality control, ANNOVAR software is used for annotation, and a TSV file after annotation is output. Each row represents a variation, and each column represents an annotation result, including: variation type, belonging gene, functional impact, harmfulness score, population frequency, clinical information, etc.

[0079] Step S9230, variation screening; Based on the variation annotation information, the variation is screened. The genes belonging to SFARIGene database with a score of 1 and S, population frequency < 0.01%, functional impact and harmfulness classification as LoF or DMis type are screened out. Finally, the screened variation data is obtained and saved as a TSV file.

[0080] Step S9240, variation pathogenicity classification; For the screened variation, according to ACMG / AMP variation rating guidelines, each variation is classified into five categories: pathogenic (P), possibly pathogenic (LP), unknown significance (VUS), possibly benign (LB), and benign (B). Finally, 3726 variations with pathogenicity classification of P or LP are selected and saved as a TSV file.

[0081] Step S930, genotype data construction; Step S9310, genotype coding; The genotype of each sample in the 44962 ASD patients is coded: if the gene carries any one or more of the 3726 P / LP variations, the genotype of the gene is coded as 1, otherwise as 0.

[0082] Step S9320, genotype screening; The genes with genotype coding of 1 and fewer than 5 carriers are screened out, and finally 186 genes are retained for downstream analysis.

[0083] Step S940, construction of gene-phenotype association map and phenotype fingerprint; Step S9410: 186 genes and 75 phenotypes are combined two by two, thereby obtaining 14508 gene-phenotype association pairs.

[0084] Step S9420: Perform independent association test for each of the 14508 gene-phenotype association pairs. For gene-phenotype association pairs with binary traits, use Firth logistic regression for test. For gene-phenotype association pairs with continuous traits, use linear regression for test. For each gene-phenotype association pair, adjust for covariates such as gender, age, genetic ancestry, sequencing batch, etc. Use FDR for multiple correction of test results.

[0085] Step S9430: Perform stepwise screening based on the results of the association test to draw the association map of genes and phenotypes. First, from all 14508 gene-phenotype association pairs, screen out 1214 gene-phenotype association pairs with P value less than 0.05, defined as the first gene-phenotype association pairs, and the corresponding phenotypes are defined as the nominal associated phenotypes.

[0086] Step S9440: Further screen out 154 association pairs with P value less than 0.05 after correction, defined as the second gene-phenotype association pairs, and the corresponding phenotypes are defined as the significantly associated phenotypes. Finally, based on all 1214 first gene-phenotype association pairs (including 154 second gene-phenotype association pairs), draw the association map.

[0087] Step S9450: Construct phenotype fingerprints. Based on the association map, integrate the association direction, effect size and significance information of each gene in all phenotype dimensions into the phenotype fingerprint of the gene. The phenotype fingerprint of the gene reveals the specific pattern of the gene affecting different clinical characteristics.

[0088] Step S950, GPS score. The present embodiment proposes a gene-centered multi-phenotype score value GPS for converting the multi-dimensional clinical phenotype information of the subject into the possibility of carrying pathogenic variants of the target gene, and thereby realizing two types of output: 1) Quantitative assessment of the risk tendency of the subject carrying gene pathogenic / likely pathogenic variants (P / LP); 2) Priority ranking of multiple candidate genes in the same target subject.

[0089] Step S950 involves two application scenarios: Scenario one: phenotype-driven gene interpretation of clinical sequencing data; In such scenarios, the target subject usually has multiple variants of unknown significance (VUS) detected by sequencing, and the core problem faced by the clinician is: among the numerous VUS, which one is most likely to be associated with the patient's current complex clinical presentation? The traditional method relies on experts to manually compare genes and phenotypes, which is time-consuming and subjective. The GPS method of the present embodiment automatically and quantitatively scores candidate genes with patient phenotypes, providing an objective basis for gene interpretation and assisting clinical decision-making.

[0090] Scenario two: phenotype-driven pre-sequencing gene prioritization In clinical practice, when facing a patient with suspected genetic disease who has complex and non-specific phenotypes, doctors often face a strategic selection dilemma when ordering genetic testing: how to quickly screen out the target genes most worthy of priority detection from among the numerous possible pathogenic genes, thereby improving diagnostic efficiency and reasonably controlling testing costs? Another core application of GPS is designed to solve this pain point. Based on the patient's detailed clinical phenotype, a quantitative score of candidate genes can be generated before any genetic testing is performed, automatically generating a data-driven list of priority detection genes. This not only helps doctors focus on high-likelihood genes at the early stage of testing, significantly shortening the analysis period, but also accelerates the overall clinical decision-making process by optimizing the detection path, improving diagnostic accuracy while reducing medical resource consumption and improving patient diagnosis experience.

[0091] Step S9510, data preprocessing; Among the 44962 ASD patients obtained in the foregoing, samples with a phenotype absence rate of more than 20% were further excluded, and finally 26529 samples were obtained as the derivation cohort of GPS.

[0092] Step S9520, GPS calculation; 1) Signal term calculation; For the candidate gene and the target subject , the present embodiment calculates the signal term value of the candidate gene based on the phenotype fingerprint of the candidate gene . Among them, the calculation formula of the continuous phenotype is formula (1) above, and the calculation formula of the binary phenotype is formula (2) above.

[0093] 2) Penalty term calculation; The nominal associated phenotype and the significantly associated phenotype are extracted from the fingerprint phenotype, and the non-genetic The nominal association phenotype and the significant association phenotype are penalized by the GPS by introducing a penalty term. Similarly, the penalty term is calculated separately for continuous and binary phenotypes. The formula for the continuous phenotype is formula (3) above, and the formula for the binary phenotype is formula (4) above.

[0094] 3) GPS final score calculation; The formula is formula (5) to formula (8) above, which are not repeated here.

[0095] Step S960, GPS performance evaluation; 1) Single gene level evaluation; To facilitate clinical interpretation and application, the embodiment reports two types of indicators for each gene: one is the discrimination ability, which measures whether the GPS can distinguish between gene P / LP carriers and non-carriers by AUC, and reports the median AUC and 95% confidence interval through 1000 bootstrap repeated sampling; the other is the ranking utility, which ranks the highest 5% of the population according to the GPS from high to low, defines the high priority layer, calculates the enrichment degree of the P / LP carriers relative to the rest of the population, and uses the two-sided Fisher's exact test to evaluate the significance.

[0096] As Figure 3 and Figure 4 In the above derived data, the GPS performs well on the 173 genes with nominal association phenotypes, with a median AUC of 0.795, of which 83 genes have an AUC>0.8. For the 48 FDR significant genes, the GPS performs further improvement: 85.4% (41 / 48) of the FDR significant association genes have an AUC>0.8, and 33.3% (16 / 48) of the genes have an AUC>0.9; accompanied by significant carrier enrichment, with a median OR of 19.061, and 72.9% (35 / 48) of the genes have an OR>10. For the 42 nominally significant but not FDR high-performance genes, the GPS also maintains strong discrimination and enrichment ability: these genes all satisfy AUC>0.8, with a median OR of 14.293 in carrier enrichment, and the highest can reach 114.629. The above results show that the GPS can not only stably identify carriers in high-confidence genes, but also extract phenotype signals with ranking value in some genes that have not yet reached the strict multiple correction threshold, thereby expanding the range of applicable genes.

[0097] 2) Overall level evaluation; Since the distribution scale of GPS raw scores of different genes may be different, directly using GPS raw scores to integrate different genes will affect the fairness of cross-gene comparison, therefore, the embodiment uses the percentile of GPS for integration, and combines multiple genes into a unified data set for global evaluation to eliminate the scale difference. The global evaluation includes two parts: first, the overall discrimination ability (AUC, reporting APR at the same time) is calculated, and the 95% confidence interval is given by 1000 times bootstrap, and the ROC curve is drawn; second, the high stratification carrier enrichment efficiency is evaluated, the objects are divided into Top20%, 10%, 5%, 1%, 0.1% and other high priority layers according to the GPS percentile, and the P / LP carrier enrichment OR and 95% confidence interval (logarithmic coordinate display) of each layer are calculated respectively.

[0098] As Figure 5 In order to facilitate clinical use, the embodiment defines two sets of gene sets in global evaluation: the high confidence set is the gene set with FDR significant and single gene AUC>0.8; the extended set adds the genes with nominal significant but single gene AUC>0.8 to the high confidence set to cover more available genes. Based on the global evaluation results, the AUC of the high confidence set is 0.869 (95% CI: 0.853-0.884), and the AUC of the extended set is 0.865 (95% CI: 0.851-0.877); at the same time, in the high stratification of Top20%, 10%, 5%, 1%, 0.1%, the enrichment degree of P / LP carriers increases with the tightening of stratification, indicating that GPS can effectively concentrate carriers in the top of the ranking, which meets the decision logic of clinical priority to high stratification.

[0099] 3) Independent queue application and verification; As Figure 6 In order to apply GPS and verify its generalization ability, the embodiment includes an independent external verification queue, which is preprocessed by using the same analysis method. In order to simulate the actual application, all the variable parameters of GPS are fixed from the derivation queue. The final results show that the global AUC of the high confidence gene set is 0.765, and the global AUC of the extended gene set is 0.771. At the same time, in the different high stratification of Top20%, 10%, 5%, 1%, 0.1%, P / LP carriers show significant enrichment, indicating that GPS can still stably concentrate carriers in the top of the ranking in independent samples, further verifying that GPS model has high accuracy and reliability, thereby supporting its application in clinical scenarios for phenotype-driven gene priority ranking and carrier risk stratification.

[0100] The embodiment first constructs gene-specific phenotype fingerprints at the population level, and then develops a GPS tool through a systematic technical solution to convert multi-dimensional clinical phenotype information into gene-centered quantitative scores at the subject level, so as to realize phenotype-driven gene positioning and priority ranking of pathogenic genes, and provide important data and tool support for pathogenic variation prediction and ranking of complex diseases. Compared with the prior art, the embodiment has at least the following beneficial effects: 1) Reducing the single dependence on the existing annotation library; The prior art highly depends on the existing gene and disease / phenotype annotation library, and the incomplete annotation or inconsistent update may easily lead to deviation in the ranking of candidate genes. The embodiment constructs gene-specific phenotype fingerprints based on large-scale cohorts, and quantitatively scores at the individual level. Even in the face of atypical or novel phenotype combinations, the systematization and stability of candidate gene evaluation can be maintained, thereby reducing omissions and result fluctuations.

[0101] 2) Forming a phenotype-driven and gene-centered quantitative mechanism; The prior art is mostly an explanatory path from genes to phenotypes, and the phenotype is often used as a posteriori filtering, and it is difficult to directly convert multi-dimensional phenotypes into comparable gene scores. The embodiment maps individual multi-dimensional phenotypes to gene fingerprint space to generate GPS scores for each gene, realizes the explainable ranking of pathogenic genes from phenotypes, is suitable for complex disease scenarios with a large number of candidate genes, highly heterogeneous phenotypes, and gene-specific significant gene effects, and significantly improves the efficiency and consistency of clinical screening.

[0102] 3) Integrated output of risk assessment and gene ranking; The embodiment can output the carrying tendency (risk stratification) of an individual to the target gene P / LP variation and the priority ranking result of the candidate gene under the same scoring framework, simplify the workflow, and improve the repeatability and landability.

[0103] 4) Providing quantifiable quality control indicators; The embodiment provides quantifiable indicators such as discrimination ability and high stratification enrichment effect at the single gene and global levels, which can directly support stratification decisions on which genes / individuals to prioritize, and facilitate consistent verification and continuous quality control in independent data.

[0104] In summary, the embodiment realizes more systematic, more explainable and more easily landed priority ranking of pathogenic genes and carrying risk assessment in complex disease scenarios, which has obvious improvements in coverage, stability, integrated output and clinical robustness compared with the prior art, fills the gap of the prior art, and provides important tool and data support for pathogenicity assessment of genetic variations.

[0105] Reference Figure 7Some embodiments of the present application provide a phenotypic fingerprint-based pathogenic gene prediction device, the device comprising: The phenotypic fingerprint construction module 1100 is configured to construct a plurality of gene-phenotype association pairs based on a plurality of phenotypes and a plurality of genes of at least one subject related to a target disease, determine association information of each gene on all phenotype dimensions based on the plurality of gene-phenotype association pairs, and construct a phenotypic fingerprint of each gene according to the association information; wherein any one gene-phenotype association pair is composed of any one phenotype and any one gene, and any two gene-phenotype association pairs are different; each gene comprises at least one pathogenic variant or a possible pathogenic variant; The prediction instruction response module 1200 is configured to determine observed phenotypes of a target object related to a target disease in response to a pathogenic gene prediction instruction of the target object. The risk prediction and ranking module 1300 is configured to calculate a multi-phenotype score value of each candidate gene of the target object with the candidate gene as the center according to the phenotypic fingerprint and the observed phenotypes, and predict a risk value of the target object carrying a corresponding pathogenic variant and determine a priority of each candidate gene according to the multi-phenotype score value of each candidate gene.

[0106] It should be noted that the phenotypic fingerprint-based pathogenic gene prediction device provided in the present embodiment is based on the same inventive concept as the above-mentioned phenotypic fingerprint-based pathogenic gene prediction method, and therefore the related content of the above-mentioned phenotypic fingerprint-based pathogenic gene prediction method is also applicable to the content of the phenotypic fingerprint-based pathogenic gene prediction device, and therefore, the details are not repeated here.

[0107] As Figure 8 The present application also provides an electronic device, which comprises: at least one hydrogen fuel cell, at least one memory, at least one processor, and at least one program; the program is stored in the memory, and the processor executes the at least one program to implement the phenotypic fingerprint-based pathogenic gene prediction method provided in the present disclosure. The electronic device can be any intelligent terminal, such as a mobile phone, a tablet computer, a personal digital assistant (PDA), a vehicle-mounted computer, etc. The electronic device of the present application will be described in detail below.

[0108] The processor 1600 can be implemented in the form of a general-purpose central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is configured to execute related programs to implement the technical solutions provided in the present disclosure. The memory 1700 can be implemented in the form of a Read Only Memory (ROM), a static storage device, a dynamic storage device, or a Random Access Memory (RAM), etc. The memory 1700 can store an operating system and other application programs. When the technical solutions provided by the embodiments of the present disclosure are implemented by software or firmware, the related program codes are stored in the memory 1700 and are invoked and executed by the processor 1600 to perform a phenotypic fingerprint-based pathogenic gene prediction method.

[0109] The input / output interface 1800 is configured to realize information input and output. The communication interface 1900 is configured to realize the communication interaction between the device and other devices. The communication can be realized in a wired manner (for example, a USB, a network cable, etc.) or in a wireless manner (for example, a mobile network, WIFI, Bluetooth, etc.). The bus 2000 is configured to transmit information between various components (for example, the processor 1600, the memory 1700, the input / output interface 1800, and the communication interface 1900) of the device. The processor 1600, the memory 1700, the input / output interface 1800, and the communication interface 1900 are connected to each other through the bus 2000 to realize the communication connection between them in the device.

[0110] The present disclosure also provides a storage medium. The storage medium is a computer-readable storage medium, and the computer-readable storage medium stores computer-executable instructions. The computer-executable instructions are configured to enable a computer to perform the phenotypic fingerprint-based pathogenic gene prediction method.

[0111] The memory is a non-transitory computer-readable storage medium, which can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include a high-speed random access memory and can also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state memory device. In some embodiments, the memory can optionally include a memory remotely arranged relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0112] The embodiments described in the present disclosure are used to more clearly illustrate the technical solutions of the present disclosure and do not constitute a limitation on the technical solutions provided by the present disclosure. Those skilled in the art can know that, as technology evolves and new application scenarios appear, the technical solutions provided by the present disclosure are also applicable to similar technical problems.

[0113] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation to the embodiments of the present disclosure, and can include more or fewer steps than the figures, or combine certain steps, or different steps.

[0114] The apparatus embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separate, that is, can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiments.

[0115] Those skilled in the art can understand that all or some steps in the above disclosed method, functions of the modules / units in the system and the device can be implemented as software, firmware, hardware and appropriate combinations thereof.

[0116] The terms "first", "second", "third", "fourth" and the like in the description of the application and the above-mentioned figures (if any) are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not necessarily limit to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0117] It should be understood that in the present application, "at least one" means one or more, and "multiple" means two or more. "And / or" is used to describe the association between the associated objects, which means that there can be three relationships, for example, "A and / or B" can mean that there are three cases of only A, only B and A and B at the same time, where A and B can be singular or plural. The character " / " generally represents that the associated objects before and after are in an "or" relationship. "At least one of the following" or similar expressions means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b or c can mean a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0118] In several embodiments provided in the present application, it should be understood that the disclosed apparatus and method can be implemented in other manners. For example, the described apparatus embodiments are merely schematic. The division of the units is merely logical function division. There can be another division manner for the actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.

[0119] The units described as separated components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purposes of the embodiments.

[0120] In addition, each functional unit in the embodiments of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be implemented in the form of hardware, or in the form of software functional units.

[0121] If the integrated unit is implemented in the form of software functional units and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on such an understanding, the technical solutions of the present application essentially or the part making contributions to the prior art, or all or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes multiple instructions used to cause an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the embodiments of the present application. The foregoing storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, and various other media that can store programs.

[0122] The above is a specific description of the preferred implementation of the embodiments of the present application, but the embodiments of the present application are not limited to the above implementation. Those skilled in the art can make various equivalent modifications or replacements without departing from the spirit of the embodiments of the present application, and these equivalent modifications or replacements are all included in the scope defined by the claims of the embodiments of the present application.

Claims

1. A method for predicting pathogenic genes based on phenotypic fingerprinting, characterized in that, The method includes: Based on at least one subject's multiple phenotypes and multiple genes associated with the target disease, multiple gene-phenotype association pairs are constructed, and the association information of each gene across all phenotypic dimensions is determined based on the multiple gene-phenotype association pairs to construct a phenotypic fingerprint for each gene according to the association information; wherein any gene-phenotype association pair consists of any phenotype and any gene, and any two gene-phenotype association pairs are different; each gene includes at least one pathogenic variant or potentially pathogenic variant; In response to the pathogenic gene prediction instruction of the target object, determine the observed phenotype of the target object associated with the target disease; Based on the phenotypic fingerprint and the observed phenotype, a gene-centered multiphenotypic score is calculated for each candidate gene of the target object, in order to predict the risk value of the target object carrying the corresponding pathogenic variant and determine the priority of each candidate gene based on the multiphenotypic score of each candidate gene.

2. The method for predicting pathogenic genes based on phenotypic fingerprints according to claim 1, characterized in that, The step of determining the association information of each gene across all phenotypic dimensions based on the multiple gene-phenotype association pairs, and constructing a phenotypic fingerprint for each gene based on the association information, includes: Independent association tests were performed on each of the multiple gene-phenotype association pairs to obtain the association test results. Based on the association test results, an association map of the multiple gene phenotype association pairs is drawn; Based on the association map, determine the association direction, effect size, and significance information of each gene across all phenotypic dimensions; A phenotypic fingerprint of each gene is constructed based on the association direction, effect size, and significance information.

3. The method for predicting pathogenic genes based on phenotypic fingerprints according to claim 2, characterized in that, The association test results include the test p-value corresponding to each of the gene-phenotype association pairs; The step of drawing an association map of the multiple gene phenotype association pairs based on the association test results includes: Based on the test P value corresponding to each gene-phenotype association pair, multiple first gene-phenotype association pairs with test P values ​​less than a first threshold are selected from the multiple gene-phenotype association pairs; wherein, the first threshold is used to characterize the association between the gene and the phenotype in the gene-phenotype association pair. The association map is drawn based on the multiple first gene phenotype association pairs.

4. The method for predicting pathogenic genes based on phenotypic fingerprints according to claim 3, characterized in that, Before drawing the association map based on the plurality of first gene phenotype association pairs, the method further includes: Based on multiple hypothesis testing, update the test p-value; Based on the updated test p-value, multiple second gene phenotype association pairs are selected from the multiple first gene phenotype association pairs; The step of calculating a gene-centered multiphenotypic score for each candidate gene of the target object based on the phenotypic fingerprint and the observed phenotype includes: Extract the phenotype from the first gene phenotype association pair corresponding to the candidate gene from the phenotype fingerprint, and calculate the signal term value of the candidate gene based on the phenotype fingerprint and the observed phenotype when the observed phenotype corresponds to the phenotype in the extracted first gene phenotype association pair. The phenotypes in the first gene phenotype association pair corresponding to the candidate gene and the phenotypes in the second gene phenotype association pair corresponding to the candidate gene are extracted from the phenotypic fingerprint. If the observed phenotype does not correspond to the phenotype in the first gene phenotype association pair or the phenotype in the second gene phenotype association pair, the penalty term value of the candidate gene is calculated based on the phenotypic fingerprint and the observed phenotype. Based on the signal term value and penalty term value of each candidate gene, a gene-centered multiphenotype score is calculated for each candidate gene.

5. The method for predicting pathogenic genes based on phenotypic fingerprints according to claim 4, characterized in that, When the observed phenotype is a continuous phenotype, the process of calculating the signal term value of the candidate gene includes: ; in, Represents the target object For candidate genes The signal term values ​​of the continuous phenotype, This represents the set of continuous phenotypes from the first gene phenotype association pairs corresponding to candidate genes. Represents the target object Observational phenotypes Phenotypic values, This indicates the observed phenotype determined based on phenotypic fingerprints. The effect size, Indicates the observed phenotype Confidence weights; When the observed phenotype is a binary phenotype, the process of calculating the signal term value of the candidate gene includes: ; in, Represents the target object For candidate genes The signal term values ​​of the binary phenotype, This represents the set of binary phenotypes from the first gene phenotype association pairs corresponding to the candidate gene. Represents the target object Observational phenotypes Phenotypic values, This indicates the observed phenotype determined based on phenotypic fingerprints. The effect size, Represents the natural logarithm function. Indicates the observed phenotype The confidence weight.

6. The method for predicting pathogenic genes based on phenotypic fingerprints according to claim 4, characterized in that, When the observed phenotype is a continuous phenotype, the process of calculating the penalty term value of the candidate gene includes: ; in, Represents the target object For candidate genes The penalty term value for continuous phenotypes, This represents the set of phenotypes from the first gene-phenotype association pairs corresponding to the candidate gene, or the set of phenotypes from the second gene-phenotype association pairs corresponding to the candidate gene. Represents the target object Observational phenotypes Phenotypic values, Indicates the observed phenotype The mean, To penalize hyperparameters; When the observed phenotype is a binary phenotype, the process of calculating the penalty term value of the candidate gene includes: ; in, Represents the target object For candidate genes The penalty term value of the binary phenotype, Indicates the observed phenotype The mean.

7. The method for predicting pathogenic genes based on phenotypic fingerprints according to claim 4, characterized in that, The step of calculating a gene-centered multiphenotype score for each candidate gene based on the signal term value and penalty term value of each candidate gene includes: ; ; ; ; in, The signal term value represents the continuous phenotype of the target object relative to the candidate genes. This represents the penalty term value for the continuous phenotype of the target object relative to the candidate gene. The signal term value represents the binary phenotype of the target object relative to the candidate gene. This represents the penalty value of the binary phenotype of the target object relative to the candidate gene. This indicates Z-score standardization. Represents the natural base. This represents the candidate gene's multiphenotype score centered on the gene itself.

8. A pathogenic gene prediction device based on phenotypic fingerprinting, characterized in that, The device includes: The phenotypic fingerprint construction module is used to construct multiple gene-phenotype association pairs based on multiple phenotypes and multiple genes associated with at least one subject and the target disease, and to determine the association information of each gene across all phenotypic dimensions based on the multiple gene-phenotype association pairs, so as to construct a phenotypic fingerprint for each gene according to the association information; wherein any gene-phenotype association pair consists of any phenotype and any gene, and any two gene-phenotype association pairs are different; each gene includes at least one pathogenic variant or potentially pathogenic variant; The prediction instruction response module is used to respond to the pathogenic gene prediction instruction of the target object and determine the observed phenotype of the target object related to the target disease. The risk prediction and ranking module is used to calculate a gene-centered multiphenotypic score for each candidate gene of the target object based on the phenotypic fingerprint and the observed phenotype, so as to predict the risk value of the target object carrying the corresponding pathogenic variant and determine the priority of each candidate gene based on the multiphenotypic score of each candidate gene.

9. An electronic device, characterized in that: The method includes at least one control processor and a memory for communicatively connecting to the at least one control processor; the memory stores instructions executable by the at least one control processor to enable the at least one control processor to perform a pathogenic gene prediction method based on phenotypic fingerprints as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions for causing a computer to perform a pathogenic gene prediction method based on phenotypic fingerprints as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method for screening candidate genes by integrating various neurodevelopmental diseases

    CN111540407A

  • Rare variation driven Alzheimer disease new gene identification and function evaluation method

    CN119049545A

  • Gene detection and phenotype verification health detection method and system

    CN120375919A

  • Breast cancer lung metastasis related biomarker screening system

    CN121601045A