A pathogenic gene identification method, device, equipment and storage medium
By using phenotypic feature matching and similarity calculation, the problem of relying on standardized descriptions for pathogenic gene identification in existing technologies has been solved, improving the efficiency and accuracy of rare disease diagnosis, especially the ability to handle atypical phenotypes.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PEKING UNION MEDICAL COLLEGE HOSPITAL
- Filing Date
- 2025-12-16
- Publication Date
- 2026-05-29
Smart Images

Figure CN122117012A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical information processing technology, and in particular to a method, apparatus, device, and storage medium for identifying pathogenic genes. Background Technology
[0002] Although rare diseases have low incidence rates for individual illnesses, they are diverse and have a significant impact on population health as a whole. More than half of rare diseases are caused by single-gene defects. Current advancements in high-throughput sequencing technology have made it possible to identify pathogenic mutations. However, each patient sample typically contains tens of thousands of genetic variations, and identifying the pathogenic variant that leads to the patient's phenotype remains a major challenge.
[0003] In existing technologies, the identification of pathogenic genes mainly relies on screening based on the biochemical characteristics, population frequency, and conservation of variations. For example, most gene / variable prioritization tools depend on standardized descriptions of patient phenotypes using HPO (Human Phenotype Ontology) terminology or codes, performing independent matching of case phenotypes. When the description of a patient's phenotype is not accurately mapped to an HPO, or when unrecorded terms are used, these traditional methods cannot effectively utilize this information. Furthermore, rare genetic diseases often exhibit atypical or complex phenotypic presentations, with significant differences between patients. HPO-based matching methods show a significant performance decline when encountering unconventional phenotypic combinations. Currently, these shortcomings in existing technologies severely impact the efficiency and accuracy of rare disease diagnosis. Summary of the Invention
[0004] This invention provides a method, apparatus, device, and storage medium for identifying pathogenic genes, aiming to improve the information utilization rate and robustness to atypical phenotypes in the identification of pathogenic genes for rare diseases, and to achieve high-efficiency and high-accuracy identification of pathogenic genes.
[0005] In a first aspect, embodiments of the present invention provide a method for identifying pathogenic genes, comprising:
[0006] Phenotypic features were matched between the target cases and candidate reference cases to obtain phenotypic matching results;
[0007] Based on the phenotypic matching results, a target reference case for the target case to be diagnosed is obtained from the candidate reference cases.
[0008] Based on the pathogenic genes of the target reference case, the pathogenic gene identification result of the target case to be diagnosed is generated.
[0009] Optionally, the step of matching the phenotypic features of the target case to be diagnosed with the candidate reference cases to obtain the phenotypic matching results includes:
[0010] Obtain the phenotypic features of the target case;
[0011] The feature similarity between the phenotypic features to be diagnosed and the reference phenotypic features of the candidate reference cases is calculated as the phenotypic matching result.
[0012] Optionally, obtaining the phenotypic features of the target case to be diagnosed includes:
[0013] Obtain the phenotypic description text of the target case to be diagnosed;
[0014] The phenotypic description text is encoded into a phenotypic feature vector using a preset vector model, which serves as the phenotypic feature to be diagnosed.
[0015] Optionally, obtaining the target reference case of the target case for diagnosis from the candidate reference cases based on the phenotypic matching result includes:
[0016] Based on the feature similarity, the candidate reference cases are sorted in descending order of similarity to obtain a candidate case sequence;
[0017] The candidate reference cases in the candidate case sequence are sequentially determined as the target reference cases by a predetermined number of candidate reference cases.
[0018] Optionally, generating the pathogenic gene identification result of the target case based on the pathogenic gene of the target reference case includes:
[0019] The gene suspicion of the pathogenic gene in each case is calculated based on the feature similarity between the reference phenotypic features and the phenotypic features to be diagnosed in each target reference case.
[0020] Based on the gene suspicion level, the pathogenic genes of the cases are sorted in descending order of suspicion level to obtain candidate pathogenic gene sequences, which are then used as the identification results of the pathogenic genes.
[0021] Optionally, before performing phenotypic feature matching between the target case and candidate reference cases to obtain the phenotypic matching result, the method further includes:
[0022] Collect confirmed case data;
[0023] Extract phenotypic description text and pathogenic genes from the confirmed case data;
[0024] In a pre-defined case database, the phenotypic description text and the pathogenic gene of the case are stored as candidate reference cases.
[0025] Optionally, generating the pathogenic gene identification result of the target case based on the pathogenic gene of the target reference case includes:
[0026] In the preset case database, obtain the phenotypic description text of the target reference case and the pathogenic gene of the case;
[0027] Based on the phenotypic description text and the pathogenic gene of the case, the pathogenic gene identification result is generated.
[0028] Secondly, embodiments of the present invention provide a pathogenic gene identification device, comprising:
[0029] The phenotypic matching module is used to match the phenotypic features of the target case to be diagnosed with the candidate reference cases to obtain the phenotypic matching results.
[0030] The reference acquisition module is used to acquire the target reference case of the target case for diagnosis from the candidate reference cases based on the phenotypic matching results.
[0031] The result generation module is used to generate the pathogenic gene identification result of the target case to be diagnosed based on the pathogenic gene of the target reference case.
[0032] Thirdly, embodiments of the present invention provide a pathogenic gene identification device, comprising:
[0033] One or more processors;
[0034] Memory, used to store one or more programs;
[0035] When the one or more programs are executed by the one or more processors, the one or more processors implement the pathogenic gene identification method provided in any embodiment of the present invention.
[0036] Fourthly, embodiments of the present invention provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the pathogenic gene identification method provided in any embodiment of the present invention.
[0037] This invention provides a method, apparatus, device, and storage medium for identifying pathogenic genes. By matching the phenotypic features of a target case to be diagnosed with candidate reference cases, a phenotypic matching result is obtained. Based on the phenotypic matching result, a target reference case of the target case to be diagnosed is obtained from the candidate reference cases. Thus, based on the pathogenic gene of the target reference case, a pathogenic gene identification result of the target case to be diagnosed is generated. This solves the problems of information loss and insufficient performance caused by the reliance on standardized descriptions and independent matching of patient phenotypes in the prior art. It improves the information utilization rate and robustness to atypical phenotypes in the identification of pathogenic genes in rare diseases, and achieves high-efficiency and high-accuracy pathogenic gene identification. Attached Figure Description
[0038] Figure 1 This is a flowchart of a pathogenic gene identification method provided in Embodiment 1 of the present invention;
[0039] Figure 2 This is a flowchart of a pathogenic gene identification method provided in Embodiment 2 of the present invention;
[0040] Figure 3 This is a flowchart of a pathogenic gene identification method provided in Embodiment 3 of the present invention;
[0041] Figure 4 This is a schematic diagram of the structure of a pathogenic gene recognition device provided in Embodiment 4 of the present invention;
[0042] Figure 5 This is a schematic diagram of the structure of a pathogenic gene identification device provided in Embodiment 5 of the present invention. Detailed Implementation
[0043] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present invention, and not all of the structures.
[0044] Example 1
[0045] Figure 1 This is a flowchart of a pathogenic gene identification method provided in Embodiment 1 of the present invention. This embodiment is applicable to the identification of pathogenic genes in patients with undiagnosed rare diseases. The method can be executed by a pathogenic gene identification device, which can be implemented by hardware and / or software and is generally integrated into an electronic device, such as a computer device. The method specifically includes:
[0046] Step 110: Perform phenotypic feature matching between the target case and the candidate reference case to obtain the phenotypic matching results.
[0047] The target case can be a patient with an undiagnosed rare disease. The candidate reference case can be a patient whose rare disease is confirmed to be caused by a specific gene defect. Phenotypic matching is an operation that determines the degree of phenotypic similarity between the target case and the candidate reference case; the resulting phenotypic matching data describes the degree of phenotypic similarity between the target case and the candidate reference case.
[0048] Specifically, target cases for diagnosis typically present with the clinical phenotype of a rare disease, but the specific rare disease and the underlying genetic defect causing it have not yet been determined. Therefore, the known information available for target cases includes at least their clinical phenotype. For candidate reference cases, the known information is usually more comprehensive, also including the patient's clinical phenotype, as well as their diagnosed rare disease and the causative gene.
[0049] It should be noted that the data type used to describe the patient's phenotype is not limited in this embodiment. For example, it may include, but is not limited to, texts described by medical staff in natural language after clinical diagnosis, structured phenotype lists filled in, indicator data obtained by medical testing methods, and part or all of the patient's self-report text.
[0050] Furthermore, based on the descriptions of the patient clinical phenotypes obtained from the target case and the candidate reference cases, phenotypic feature matching can be performed on the target case and the candidate reference cases to obtain phenotypic matching results. The degree of similarity between the patient phenotypes of the target case and the candidate reference cases reflected by the phenotypic matching results can be used as a basis for judging whether the diagnostic status of the candidate reference cases can provide a reliable reference in the diagnosis of the target case.
[0051] Step 120: Based on the phenotypic matching results, obtain the target reference case for the target case from the candidate reference cases.
[0052] The target reference case can be a candidate reference case whose clinical phenotype is sufficiently similar to the target case based on the description of its phenotypic matching results, so as to provide a reliable reference in the diagnosis of the target case.
[0053] Specifically, there can be one or more candidate reference cases. Preferably, the number of candidate reference cases is usually large enough that the degree of similarity between the clinical phenotype of each candidate reference case and the target case can correspond to a phenotype matching result. In this way, the target reference case with the most similar clinical phenotype to the target case can be obtained from a sufficient number of candidate reference cases.
[0054] Optionally, the number of target reference cases can be one or more. Typically, since the probability of obtaining a target reference case whose clinical phenotype is completely identical to the target case is extremely low, and even if the patients' clinical phenotypes are very similar, considering the complex characteristics of rare diseases, it is impossible to diagnose the target case based on the diagnosis of a single target reference case. Therefore, the preferred number of target reference cases is multiple. Specifically, a specific number of target reference cases can be selected from the candidate reference cases by setting screening criteria for phenotype matching results.
[0055] Step 130: Based on the pathogenic genes of the target reference case, generate the pathogenic gene identification results of the target case to be diagnosed.
[0056] In this context, the pathogenic gene in the case can be a gene in the target reference case that has been diagnosed as causing the rare disease in the patient. The pathogenic gene identification result can be data describing the current results of identifying the pathogenic genes causing the rare disease in the target case to be diagnosed.
[0057] Specifically, since the target reference cases are already diagnosed cases, the pathogenic genes of those diagnosed cases can be obtained. Based on the pathogenic genes of one or more target reference cases of the target case to be diagnosed, a reasonable suspected diagnosis of the pathogenic gene of the target case to be diagnosed can be made, thereby generating the diagnosis as a pathogenic gene identification result.
[0058] The technical solution of this embodiment obtains phenotypic matching results by matching the phenotypic features of the target case to be diagnosed with candidate reference cases. Based on the phenotypic matching results, the target reference case of the target case to be diagnosed is obtained from the candidate reference cases. Thus, based on the pathogenic gene of the target reference case, the pathogenic gene identification result of the target case to be diagnosed is generated. This solves the problems of information loss and insufficient performance caused by the reliance on standardized description and independent matching of patient phenotypes in the prior art. It improves the information utilization rate and robustness to atypical phenotypes in the identification of pathogenic genes of rare diseases, and achieves high-efficiency and high-accuracy pathogenic gene identification.
[0059] Example 2
[0060] Figure 2 This is a flowchart of a pathogenic gene identification method provided in Embodiment 2 of the present invention. This embodiment further refines the above technical solution, and may involve matching the phenotypic features of a target case to be diagnosed with candidate reference cases to obtain a phenotypic matching result. This includes: obtaining the phenotypic features of the target case to be diagnosed; calculating the feature similarity between the phenotypic features of the target case to be diagnosed and the reference phenotypic features of the candidate reference cases, as the phenotypic matching result. Specifically, this method includes:
[0061] Step 210: Obtain the phenotypic characteristics of the target case.
[0062] Among them, the phenotypic features to be diagnosed can be data that describes the clinical phenotype of a patient in a specific form.
[0063] Specifically, based on the raw data describing the clinical phenotype of the target case, its diagnostic phenotypic features can be obtained. These features can describe all the characteristics of the target case's clinical phenotype related to the identification and diagnosis of rare disease-causing genes, and can be presented in a pre-defined specific form. Optionally, the diagnostic phenotypic features can be a feature vector representation of the raw data of the target case's clinical phenotype.
[0064] In one optional implementation, obtaining the phenotypic features of the target case may include: obtaining the phenotypic description text of the target case; and encoding the phenotypic description text into a phenotypic feature vector using a preset vector model, which serves as the phenotypic features to be diagnosed.
[0065] The phenotypic description text can be a written presentation of the clinical phenotype of a target case for diagnosis. The pre-built vector model can be a pre-built and trained artificial intelligence model that encodes the input phenotypic description text into phenotypic feature vectors. The phenotypic feature vectors can describe the semantic features of the phenotypic description text in a vector space.
[0066] Specifically, the target case for diagnosis may include phenotypic description text generated by the patient during clinical diagnosis and treatment. The phenotypic description text may include free text case records or structured phenotypic lists. By inputting the phenotypic description text of the target case for diagnosis into a preset vector model, the corresponding phenotypic feature vector output by the preset vector model can be obtained as the phenotypic features to be diagnosed.
[0067] In an optional implementation, when the phenotypic description text of the target case contains multiple independent patient clinical phenotypic records, the preset vector model can encode the phenotypic description text of each independent patient clinical phenotypic record separately and then take a simple average or a weighted average. The weights used in the weighted average can be pre-configured based on the clinician's prior knowledge of the importance of the phenotype. The resulting phenotypic feature vector can comprehensively describe the phenotype of the target case, making the phenotypic features of the target case comprehensive phenotypic features at the case / patient level, thereby improving the robustness of phenotypic feature matching in the face of complex phenotypes.
[0068] In an optional implementation, the preset vector model can be the SBERT model (SentenceBidirectional Encoder Representations from Transformers).
[0069] Step 220: Calculate the feature similarity between the phenotypic features of the patient to be diagnosed and the reference phenotypic features of the candidate reference cases, and use this as the phenotypic matching result.
[0070] The reference phenotypic features can be data describing the characteristics of a patient's clinical phenotype in a specific form, and the form and the described features are consistent with the phenotypic features to be diagnosed. Feature similarity can be data describing the degree of similarity between the phenotypic features to be diagnosed and the reference phenotypic features; the higher the feature similarity, the more similar the described phenotypic features to be diagnosed are to the reference phenotypic features.
[0071] Specifically, using the same data processing methods as those used to obtain the phenotypic features of the target cases, reference phenotypic features of candidate cases can be obtained. Therefore, the similarity between the phenotypic features of the target cases and the reference phenotypic features can be used to determine the similarity between the clinical phenotypes of the target cases and the candidate reference cases. This can be achieved by calculating the feature similarity between the phenotypic features of the target cases and each reference phenotypic feature, thus obtaining the phenotypic matching result. When there are multiple candidate reference cases, each candidate reference case has a corresponding reference phenotypic feature, and a corresponding feature similarity is obtained, which serves as the phenotypic matching result corresponding to each candidate reference case.
[0072] In an optional implementation, when the phenotypic features to be diagnosed and the reference phenotypic features are in the form of feature vectors, the feature similarity can be the cosine similarity between the phenotypic features to be diagnosed and the reference phenotypic features.
[0073] In an optional implementation, the reference phenotypic features of candidate reference cases can be pre-acquired and stored in a specific storage space. When a target case for diagnosis appears, the phenotypic features of the target case for diagnosis can be acquired in a specific way, and feature similarity calculations can be performed between them and each pre-stored reference phenotypic feature.
[0074] Step 230: Based on the phenotypic matching results, obtain the target reference case for the target case from the candidate reference cases.
[0075] In an optional implementation, obtaining target reference cases for the target case from the candidate reference cases based on the phenotypic matching results may include: sorting the candidate reference cases in descending order of similarity based on feature similarity to obtain a candidate case sequence; and sequentially determining a preset number of candidate reference cases in the candidate case sequence as target reference cases.
[0076] The descending similarity sorting can be an operation that sorts candidate reference cases from high to low according to their respective feature similarity, resulting in a candidate case sequence that is all candidate reference cases sorted from high to low feature similarity. The preset reference number can be a number of target reference cases set in advance as needed, not exceeding the total number of candidate reference cases.
[0077] Specifically, the higher the feature similarity of the candidate reference cases, the greater the similarity between the clinical phenotypes of the patients in those cases and the clinical phenotypes of the target cases. Therefore, the candidate reference cases are more suitable as target reference cases. Thus, by sorting the candidate reference cases in descending order of similarity, a candidate case sequence is obtained. Then, by sequentially selecting a predetermined number of candidate reference cases from this sequence as target reference cases, the predetermined number of candidate reference cases with the highest feature similarity can be obtained and used as the target reference cases.
[0078] Step 240: Based on the pathogenic genes of the target reference case, generate the pathogenic gene identification results of the target case to be diagnosed.
[0079] In an optional implementation, generating pathogenic gene identification results for target cases based on the pathogenic genes of target reference cases may include: calculating the gene suspicion of each pathogenic gene based on the feature similarity between the reference phenotypic features and the phenotypic features of each target reference case; and sorting the pathogenic genes of the cases in descending order of suspicion based on the gene suspicion to obtain candidate pathogenic gene sequences as pathogenic gene identification results.
[0080] In this context, gene suspicion level describes the likelihood that the pathogenic gene in a case is the pathogenic gene in the target case that causes the patient to suffer from a rare disease. The higher the gene suspicion level, the greater the likelihood that the described pathogenic gene in the case is the pathogenic gene in the target case that causes the patient to suffer from a rare disease. Descending suspicion ranking involves sorting the pathogenic genes in the cases according to their respective gene suspicion levels from high to low. The resulting candidate pathogenic gene sequences can be the pathogenic genes of all target reference cases, ranked from high to low gene suspicion level.
[0081] Specifically, considering that different target reference cases may share the same pathogenic gene, it is necessary to calculate the gene suspiciousness of the pathogenic gene for each case. The more target reference cases corresponding to a pathogenic gene in any case, and the higher the feature similarity between the corresponding target reference cases, the higher the gene suspiciousness of the pathogenic gene in that case.
[0082] Furthermore, by sorting the pathogenic genes of the cases in descending order of suspicion, the resulting candidate pathogenic gene sequences reflect the ranking of the probability of pathogenic genes in the target cases and can be used as the results of pathogenic gene identification.
[0083] The technical solution of this embodiment provides an effective means of utilizing free text phenotypic descriptions by acquiring the phenotypic features of the target case to be diagnosed and the reference phenotypic features of the candidate reference cases, and calculating the feature similarity between the phenotypic features, and further eliminates the dependence on predefined standardized terms such as HPO.
[0084] Example 3
[0085] Figure 3 This is a flowchart of a pathogenic gene identification method provided in Embodiment 3 of the present invention. This embodiment further refines the above technical solution by including, before performing phenotypic feature matching between the target case and candidate reference cases to obtain the phenotypic matching result, the following steps are included: collecting confirmed case data; obtaining the case diagnosis conclusion, phenotypic description text, and the pathogenic gene of the case from the confirmed case data; and storing the case diagnosis conclusion, the phenotypic description text, and the pathogenic gene of the case as corresponding candidate reference cases in a preset case database. The method specifically includes:
[0086] Step 310: Collect data on confirmed cases.
[0087] Among them, confirmed case data can be related case data of patients who have been diagnosed with a specific rare disease caused by a specific gene.
[0088] Specifically, confirmed case data can be collected from any reliable database, such as hospital electronic medical record systems and publicly available medical literature databases. Confirmed case data can be generated during the patient's treatment process or recorded in literature, and should at least include a description of each patient's clinical phenotype and a record of the diagnosed pathogenic gene.
[0089] In one alternative implementation, collecting confirmed case data may include collecting case data from a large number of confirmed patients with related rare diseases using a large language model from publicly available literature.
[0090] Step 320: Obtain phenotypic description text and pathogenic genes from the confirmed case data.
[0091] Specifically, through data processing such as data preprocessing and semantic understanding of confirmed case data, phenotypic description text and pathogenic genes can be extracted from the confirmed case data. Optionally, data preprocessing of confirmed case data can be achieved by standardizing and cleaning the data, such as correcting spelling errors and standardizing medical terminology, thereby extracting phenotypic description text and pathogenic genes from the preprocessed confirmed case data.
[0092] In an optional implementation, the case diagnosis conclusion can also be obtained from the confirmed case data. The case diagnosis conclusion can be the name of a rare disease diagnosed in the patient.
[0093] Step 330: In the preset case database, store the phenotypic description text and the corresponding pathogenic gene of the case as candidate reference cases.
[0094] The preset case library can be a pre-defined storage space for storing candidate reference cases. By storing phenotypic description texts and corresponding pathogenic genes of cases as candidate reference cases in the preset case library, the phenotypic description texts of each candidate reference case can be sequentially retrieved from the preset case library for phenotypic feature matching with the target case to be diagnosed.
[0095] In an optional implementation, the candidate reference cases may also include case diagnosis conclusions stored in relation to phenotypic description text and the pathogenic gene of the case.
[0096] In one optional implementation, the preset case library can be a pre-built knowledge base, in which candidate reference cases can be stored in the form of triples. Specifically, in the preset case library, storing the phenotypic description text and the pathogenic gene of the case as corresponding candidate reference cases can include: constructing a triple of "case ID-phenotypic description text-pathogenic gene of the case" and storing it in the case knowledge base.
[0097] In an optional implementation, after storing the phenotypic description text and the corresponding pathogenic gene of the case as candidate reference cases in a preset case database, the implementation may further include: generating reference phenotypic features based on the phenotypic description text of each candidate reference case; wherein, the reference phenotypic features are preferably phenotypic feature vectors; and storing the association between each candidate reference case and each reference phenotypic feature in the preset case database.
[0098] Step 340: Perform phenotypic feature matching between the target case and the candidate reference case to obtain the phenotypic matching results.
[0099] Step 350: Based on the phenotypic matching results, obtain the target reference case for the target case from the candidate reference cases.
[0100] Step 360: Based on the pathogenic genes of the target reference case, generate the pathogenic gene identification results of the target case to be diagnosed.
[0101] In one optional implementation, generating pathogenic gene identification results for the target case based on the pathogenic gene of the target reference case may include:
[0102] Obtain the phenotypic description text and pathogenic genes of the target reference cases from the preset case database;
[0103] Based on the phenotypic description text and the pathogenic gene of the case, generate pathogenic gene identification results.
[0104] Specifically, when using a pre-set case database to store candidate reference cases, after identifying the target reference case, the corresponding phenotypic description text and the pathogenic gene of the case can be obtained from the pre-set case database to jointly generate the pathogenic gene identification result. Specifically, the pathogenic gene and phenotypic description text of the case can be added to the pathogenic gene identification result, so that the pathogenic gene identification result not only includes the pathogenic gene of the identified target case, but also includes the specific phenotypic information of the target reference case it references, so as to provide interpretable and robust diagnostic support.
[0105] In an optional implementation, the case diagnosis conclusions of the target reference cases and / or the collection sources of the corresponding confirmed case data can also be obtained from the preset case database, and these data can also be added to the pathogenic gene identification results along with the pathogenic genes of the cases, further enriching the information that the pathogenic gene identification results can provide.
[0106] The technical solution of this embodiment provides a reliable data foundation for the final pathogenic gene identification results by extensively collecting confirmed case data and extracting candidate reference cases based on the confirmed case data.
[0107] In a typical example of an embodiment of the present invention, the pathogenic gene identification method may specifically include four stages: constructing a case knowledge base, generating phenotype embeddings, encoding patients to be diagnosed, and similarity retrieval and candidate gene ranking.
[0108] The case knowledge base construction phase utilizes large language models, such as GPT-4, from publicly available literature to collect a large amount of case data from patients diagnosed with Mendelian diseases. This includes clinical phenotype descriptions and the diagnosed pathogenic genes for each patient. The collected phenotypic text information is then standardized and cleaned, such as correcting spelling errors and standardizing medical terminology. After processing, each case phenotype and its corresponding gene are combined into a "case-phenotype-gene" triple and stored in the knowledge base. For example, a case reported in a literature: "The child presented with congenital blindness and hearing loss, diagnosed with Usher syndrome, with the pathogenic gene being MYO7A," would be recorded in the knowledge base as: "Case ID 123 - Phenotypic Description: Congenital blindness, hearing loss - Gene: MYO7A." Through the accumulation of a large amount of such case data, a rich case knowledge base is formed, laying the foundation for subsequent similarity retrieval.
[0109] The phenotypic embedding generation stage can employ the SBERT model, encoding the phenotypic text of each patient into a high-dimensional phenotypic embedding vector. Specifically, the phenotypic description of a case, including an entire text or a list of multiple phenotypic traits, is input into a pre-trained SBERT model, and the output sentence vector is extracted as the phenotypic feature vector for that case. For cases containing multiple independent phenotypic records, each record can be encoded separately and then a weighted average can be taken to obtain a comprehensive phenotypic vector at the patient level. In this way, each case in the knowledge base is associated with a numerical vector representation, which captures the semantic features of the case's phenotypic phenotype in the vector space. Compared to traditional ontology one-hot encoding or handcrafted features, this vector representation can more finely measure the semantic similarity between phenotypic traits of different cases.
[0110] In the patient coding stage, for patients requiring genetic identification, their clinical phenotype descriptions can be obtained. These descriptions can be free text medical records or structured phenotype lists. Using the same SBERT model as in the phenotype embedding generation stage, the patient's phenotype text is encoded into a phenotype feature vector for diagnosis. If a patient has multiple phenotype records, they can be processed and fused using the same method as for cases in the knowledge base to obtain a comprehensive phenotype vector for the patient. Because the SBERT model is trained on a large-scale corpus, it can map different expressions with similar meanings to similar vectors. This allows the present invention to handle diverse expressions in clinical descriptions without being limited to fixed terminology.
[0111] The similarity retrieval and candidate gene ranking stage calculates the cosine similarity between the phenotypic feature vector of the patient to be diagnosed and the phenotypic vector of each confirmed patient in the case knowledge base. Then, based on the similarity scores, the N reference cases with the highest similarity are retrieved from highest to lowest. These reference cases are a group of known cases that are phenotypically most similar to the patient to be diagnosed. Next, the pathogenic genes corresponding to each of these reference cases are extracted, and their similarity scores are weighted and aggregated to generate a list of candidate pathogenic genes ranked in descending order of score. Specifically, the similarity of each reference case can be considered as the weight of the pathogenic gene in that case. When multiple cases share the same gene, their weights can be accumulated to improve the ranking of that gene. For example, if three of the ten most similar cases share the same pathogenic gene, GeneA, and their similarity scores with the patient to be diagnosed are 0.95, 0.90, and 0.85, respectively, then the overall score of GeneA can be increased accordingly. Finally, the system outputs candidate gene sequences such as "GeneA, GeneB, GeneC...", arranged from highest to lowest probability, and provides reference case information associated with each gene, such as case ID, disease name, literature source, etc., for clinicians to refer to and verify.
[0112] Example 4
[0113] Figure 4 This is a schematic diagram of the structure of a pathogenic gene recognition device provided in Embodiment 4 of the present invention, as shown below. Figure 4 As shown, the pathogenic gene identification device includes: a phenotype matching module 410, a reference acquisition module 420, and a result generation module 430, wherein,
[0114] Phenotypic matching module 410 is used to perform phenotypic feature matching between the target case to be diagnosed and the candidate reference case to obtain phenotypic matching results.
[0115] Reference acquisition module 420 is used to acquire a target reference case of the target case of diagnosis from the candidate reference cases based on the phenotypic matching result;
[0116] The result generation module 430 is used to generate the pathogenic gene identification result of the target case to be diagnosed based on the pathogenic gene of the target reference case.
[0117] The technical solution of this embodiment obtains phenotypic matching results by matching the phenotypic features of the target case to be diagnosed with candidate reference cases. Based on the phenotypic matching results, the target reference case of the target case to be diagnosed is obtained from the candidate reference cases. Thus, based on the pathogenic gene of the target reference case, the pathogenic gene identification result of the target case to be diagnosed is generated. This solves the problems of information loss and insufficient performance caused by the reliance on standardized description and independent matching of patient phenotypes in the prior art. It improves the information utilization rate and robustness to atypical phenotypes in the identification of pathogenic genes of rare diseases, and achieves high-efficiency and high-accuracy pathogenic gene identification.
[0118] In an optional implementation, the phenotypic matching module 410 may include: a feature acquisition unit for acquiring the phenotypic features of the target case to be diagnosed; and a feature matching unit for calculating the feature similarity between the phenotypic features to be diagnosed and the reference phenotypic features of the candidate reference case, as the phenotypic matching result.
[0119] In an optional implementation, the feature acquisition unit may be specifically used to: acquire the phenotypic description text of the target case to be diagnosed; and encode the phenotypic description text into a phenotypic feature vector through a preset vector model, as the phenotypic feature to be diagnosed.
[0120] In an optional implementation, the reference acquisition module 420 may include: a candidate sorting unit, configured to sort the candidate reference cases in descending order of similarity according to the feature similarity to obtain a candidate case sequence; and a reference determination unit, configured to sequentially determine a preset number of candidate reference cases in the candidate case sequence as the target reference cases.
[0121] In an optional implementation, the result generation module 430 may include: a gene acquisition unit, configured to calculate the gene suspicion of each pathogenic gene in each of the target reference cases based on the feature similarity between the reference phenotypic features and the phenotypic features to be diagnosed; and a gene sorting unit, configured to sort the pathogenic genes in the cases in descending order of suspicion based on the gene suspicion, obtain candidate pathogenic gene sequences, and use them as the pathogenic gene identification results.
[0122] In an optional implementation, the pathogenic gene identification device may further include: a case database construction module for collecting confirmed case data; obtaining phenotypic description text and the pathogenic gene of the case from the confirmed case data; and storing the phenotypic description text and the pathogenic gene of the case as candidate reference cases in a preset case database.
[0123] In an optional implementation, the result generation module 430 may be specifically used to: obtain the phenotypic description text and the pathogenic gene of the target reference case from the preset case database; and generate the pathogenic gene identification result based on the phenotypic description text and the pathogenic gene.
[0124] The pathogenic gene identification device provided in the embodiments of the present invention can execute the pathogenic gene identification method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method.
[0125] Example 5
[0126] Figure 5 This is a schematic diagram of the structure of a pathogenic gene identification device provided in Embodiment 5 of the present invention, as shown below. Figure 5 As shown, the pathogenic gene identification device includes a processor 510, a memory 520, an input device 530, and an output device 540; the number of processors 510 in the pathogenic gene identification device can be one or more. Figure 5 Taking a processor 510 as an example; the processor 510, memory 520, input device 530, and output device 540 in the pathogenic gene recognition device can be connected via a bus or other means. Figure 5 Taking the example of a connection between China and Israel via a bus.
[0127] The memory 520, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the pathogenic gene identification method in this embodiment of the invention (e.g., the phenotype matching module 410, the reference acquisition module 420, and the result generation module 430 in the pathogenic gene identification device). The processor 510 executes various functional applications and data processing of the pathogenic gene identification device by running the software programs, instructions, and modules stored in the memory 520, thereby realizing the aforementioned pathogenic gene identification method.
[0128] The memory 520 may primarily include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a given function; the data storage area may store data created based on terminal usage. Furthermore, the memory 520 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some instances, the memory 520 may further include memory remotely located relative to the processor 510, which can be connected to a pathogenic gene identification device via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0129] The input device 530 can be used to receive input digital or character information, and to generate key signal inputs related to user settings and function control of the pathogenic gene identification device. The output device 540 may include a display device such as a screen.
[0130] Example 6
[0131] Embodiment 6 of the present invention also provides a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform a method for identifying pathogenic genes, including:
[0132] Phenotypic features were matched between the target cases and candidate reference cases to obtain phenotypic matching results;
[0133] Based on the phenotypic matching results, a target reference case for the target case to be diagnosed is obtained from the candidate reference cases.
[0134] Based on the pathogenic genes of the target reference case, the pathogenic gene identification result of the target case to be diagnosed is generated.
[0135] Of course, the computer-executable instructions provided in the embodiments of the present invention are not limited to the method operations described above, but can also perform related operations in the pathogenic gene identification method provided in any embodiment of the present invention.
[0136] Based on the above description of the implementation methods, those skilled in the art can clearly understand that the present invention can be implemented using software and necessary general-purpose hardware, and of course, it can also be implemented using hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk, or optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0137] It is worth noting that in the embodiments of the above-mentioned pathogenic gene identification device, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the scope of protection of the present invention.
[0138] Although the present invention has been described in detail above with general descriptions, specific embodiments, and experiments, modifications or improvements can be made to it, which will be obvious to those skilled in the art. Therefore, all such modifications or improvements made without departing from the spirit of the present invention fall within the scope of protection claimed by the present invention.
Claims
1. A method for identifying pathogenic genes, characterized in that, include: Phenotypic features were matched between the target cases and candidate reference cases to obtain phenotypic matching results; Based on the phenotypic matching results, a target reference case for the target case to be diagnosed is obtained from the candidate reference cases. Based on the pathogenic genes of the target reference case, the pathogenic gene identification result of the target case to be diagnosed is generated.
2. The method according to claim 1, characterized in that, The phenotypic feature matching between the target case and the candidate reference cases to obtain the phenotypic matching results includes: Obtain the phenotypic features of the target case; The feature similarity between the phenotypic features to be diagnosed and the reference phenotypic features of the candidate reference cases is calculated as the phenotypic matching result.
3. The method according to claim 2, characterized in that, The process of obtaining the phenotypic features of the target case includes: Obtain the phenotypic description text of the target case to be diagnosed; The phenotypic description text is encoded into a phenotypic feature vector using a preset vector model, which serves as the phenotypic feature to be diagnosed.
4. The method according to claim 2, characterized in that, The step of obtaining the target reference case for the target case of diagnosis from the candidate reference cases based on the phenotypic matching results includes: Based on the feature similarity, the candidate reference cases are sorted in descending order of similarity to obtain a candidate case sequence; The candidate reference cases in the candidate case sequence are sequentially determined as the target reference cases by a predetermined number of candidate reference cases.
5. The method according to claim 4, characterized in that, The step of generating the pathogenic gene identification result of the target case based on the pathogenic gene of the target reference case includes: The gene suspicion of the pathogenic gene in each case is calculated based on the feature similarity between the reference phenotypic features and the phenotypic features to be diagnosed in each target reference case. Based on the gene suspicion level, the pathogenic genes of the cases are sorted in descending order of suspicion level to obtain candidate pathogenic gene sequences, which are then used as the identification results of the pathogenic genes.
6. The method according to claim 1, characterized in that, Before performing phenotypic feature matching between the target case and candidate reference cases to obtain the phenotypic matching results, the procedure also includes: Collect confirmed case data; Extract phenotypic description text and pathogenic genes from the confirmed case data; In a pre-defined case database, the phenotypic description text and the pathogenic gene of the case are stored as candidate reference cases.
7. The method according to claim 6, characterized in that, The step of generating the pathogenic gene identification result of the target case based on the pathogenic gene of the target reference case includes: In the preset case database, obtain the phenotypic description text of the target reference case and the pathogenic gene of the case; Based on the phenotypic description text and the pathogenic gene of the case, the pathogenic gene identification result is generated.
8. A pathogenic gene recognition device, characterized in that, include: The phenotypic matching module is used to match the phenotypic features of the target case to be diagnosed with the candidate reference cases to obtain the phenotypic matching results. The reference acquisition module is used to acquire the target reference case of the target case for diagnosis from the candidate reference cases based on the phenotypic matching results. The result generation module is used to generate the pathogenic gene identification result of the target case to be diagnosed based on the pathogenic gene of the target reference case.
9. A pathogenic gene identification device, characterized in that, include: One or more processors; Memory, used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the pathogenic gene identification method as described in any one of claims 1-7.
10. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to perform the pathogenic gene identification method as described in any one of claims 1-7.