Artificial intelligence-based mutation point site prediction screening system and method
By constructing a medical knowledge graph and dynamically adjusting the weights of evaluation indicators, the problems of insufficient integration of multi-source evidence and insufficient dynamic adaptability in gene testing are solved, enabling efficient and accurate screening of variant sites and supporting disease risk prediction and personalized medicine.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- DONGHUA MEDICAL TECH CO LTD
- Filing Date
- 2025-08-13
- Publication Date
- 2026-05-29
Smart Images

Figure CN121122388B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the fields of bioinformatics, gene testing and medicine, and in particular to an artificial intelligence-based variant site prediction and screening system and method. Background Technology
[0002] In the field of gene testing, mutation site screening is a core component of disease risk prediction and personalized medicine. Current technology faces three major bottlenecks:
[0003] Inefficient literature screening: 90% of variant locus studies rely on manual literature searches (PubMed, etc.), requiring an average of 42 reviews per locus, taking more than 3 hours per locus. Inconsistent evaluation criteria: Pathogenicity rating conflicts between different databases (such as ClinVar, dbSNP) reach 23%, increasing the risk of clinical misdiagnosis. Lagging dynamic updates: Traditional techniques have an update cycle of ≥6 months, failing to respond to newly published research findings. Summary of the Invention
[0004] This disclosure provides at least one artificial intelligence-based mutation site prediction and screening system and method to solve at least one of the above-mentioned technical problems.
[0005] According to one aspect of this disclosure, an artificial intelligence-based variant site prediction and screening system is provided, including a literature retrieval module, an evaluation index extraction module, a scoring module, a variant site screening module, and a user interaction module.
[0006] The literature retrieval module is used to capture data from multiple sources and construct a medical knowledge graph for predicting variant locations based on the captured data. Specifically, a trained graph construction model is used to perform entity recognition and relation extraction on the captured data, and the medical knowledge graph is constructed based on the recognized entities and extracted relations.
[0007] The evaluation index extraction module is used to extract evaluation indicators using the medical knowledge graph and determine the dynamic weight of each evaluation indicator based on scene tags and time decay factors.
[0008] The scoring module is used to determine the risk score value of each site in the gene sequence to be detected based on each evaluation index and the dynamic weight corresponding to each evaluation index.
[0009] The variant site screening module is used to screen out abnormal sites from each site based on the risk score value;
[0010] The user interaction module is used to display the abnormal locations and show the text boxes for modifying the dynamic weights corresponding to each evaluation indicator, so that the user can input the correction values for the dynamic weights corresponding to each evaluation indicator.
[0011] In one possible implementation, the graph construction model performs entity recognition on the captured data based on part-of-speech tagging; and uses a multi-head attention mechanism for relation extraction.
[0012] In one possible implementation, the evaluation index extraction module, when determining the dynamic weights corresponding to each evaluation index based on scene labels and time decay factors, is specifically used for:
[0013] The evaluation indicators were normalized.
[0014] Based on the publication time of the literature, determine the time decay factor, and determine the scene label;
[0015] The normalized evaluation indicators, time decay factor and scene label are input into the trained weights to dynamically adjust the model.
[0016] Obtain the dynamic weights of each evaluation index output by the weight dynamic adjustment model.
[0017] In one possible implementation, the evaluation metrics include: functional evidence, clinical relevance, population frequency, evolutionary conservation, literature support, and study confidence.
[0018] In one possible implementation, when the scoring module determines the risk score value of each locus in the gene sequence to be detected based on each evaluation index and the dynamic weight corresponding to each evaluation index, it is specifically used for:
[0019] The features of each site in the gene sequence to be detected, each evaluation index, and the dynamic weights corresponding to each evaluation index are mapped to the target space.
[0020] Determine the risk score of the source domain and the sufficiency of the target domain data;
[0021] By utilizing the source domain risk score and target domain data sufficiency, as well as the characteristics of each location in the target space, each evaluation index, and the dynamic weights corresponding to each evaluation index, the risk score value of each location in the gene sequence to be detected is determined.
[0022] In one possible implementation, the risk score is calculated using the following formula:
[0023] Risk score = ∑(evaluation index i × dynamic weight i) + migration compensation item;
[0024] Migration compensation item = Source domain risk score × (1 - Target domain data sufficiency).
[0025] In one possible implementation, when the variant site screening module selects anomalous sites from each site based on the risk score, it is specifically used for:
[0026] For each location, the clinical value is determined using the corresponding risk score and disease burden weight; and the technical cost is determined using experimental validation cost and clinical validation cost; an optimization model is constructed using the clinical value and technical cost, and the optimization model is solved to obtain the outlier score for that location.
[0027] Points with abnormal scores higher than the preset value are designated as abnormal points.
[0028] In one possible implementation, the user interaction module is further configured to:
[0029] It displays a medical knowledge graph, various evaluation indicators, and the dynamic weights corresponding to each evaluation indicator.
[0030] In one possible implementation, the user interaction module is further configured to:
[0031] A heat map is generated based on the anomaly scores of each location;
[0032] The heatmap is shown.
[0033] In one possible implementation, the multi-source evidence includes data from literature databases, data from public databases, clinical data, and preprints.
[0034] According to another aspect of this disclosure, an artificial intelligence-based method for predicting and screening variant sites is provided, comprising:
[0035] The literature retrieval module is used to capture data from multiple sources, and a medical knowledge graph for predicting variant locations is constructed based on the captured data. Specifically, a trained graph construction model is used to perform entity recognition and relation extraction on the captured data, and the medical knowledge graph is constructed based on the recognized entities and extracted relations.
[0036] The evaluation index extraction module extracts evaluation indicators based on the medical knowledge graph, and determines the dynamic weight of each evaluation index based on scene tags and time decay factors.
[0037] The scoring module is used to determine the risk score value of each site in the gene sequence to be tested based on each evaluation indicator and the dynamic weight of each evaluation indicator.
[0038] The abnormal sites are selected from each site based on the risk score using the variant site screening module.
[0039] The abnormal points are displayed using a user interaction module, and a text box for modifying the dynamic weights of each evaluation indicator is shown, so that the user can input the correction value for the dynamic weights of each evaluation indicator.
[0040] This disclosed AI-based variant site prediction and screening system includes a literature retrieval module, an evaluation index extraction module, a scoring module, a variant site screening module, and a user interaction module. The literature retrieval module is used to capture data from multiple sources and construct a medical knowledge graph for variant site prediction based on the captured data. Specifically, a trained knowledge graph construction model is used to perform entity recognition and relation extraction on the captured data, and the medical knowledge graph is constructed based on the recognized entities and extracted relations. The evaluation index extraction module is used to extract evaluation indicators from the medical knowledge graph and determine the dynamic weights corresponding to each evaluation indicator based on scene tags and time decay factors. The scoring module is used to determine the risk score value of each site in the gene sequence to be detected based on each evaluation indicator and its corresponding dynamic weight. The variant site screening module is used to screen out abnormal sites from each site based on the risk score value. The user interaction module displays the abnormal sites and shows a text box for modifying the dynamic weights corresponding to each evaluation indicator, allowing the user to input correction values for the dynamic weights of each evaluation indicator. The disclosed solution enables automated functions such as literature retrieval, evaluation index extraction, site scoring, and variant site screening. It can not only effectively improve the efficiency, timeliness, and accuracy of variant site screening, but also effectively screen out high-quality variant sites, so as to provide important support for disease risk prediction, personalized medicine, and clinical gene testing.
[0041] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0042] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0043] Figure 1 This is one of the structural schematic diagrams of the AI-based mutation site prediction and screening system in this embodiment;
[0044] Figure 2 This is the second schematic diagram of the structure of the AI-based mutation site prediction and screening system in this embodiment;
[0045] Figure 3 This is a flowchart of the AI-based mutation site prediction and screening method in this embodiment. Detailed Implementation
[0046] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0047] Currently, the following technical solutions are used for anomaly detection:
[0048] 1. Static database screening (ClinVar), which is based on expert-reviewed pathogenicity classification (Pathogenic / Benign).
[0049] The technology has the following drawbacks: delayed updates result in an 8-month lag in gene locus rating; and the lack of a quantitative scoring mechanism makes it impossible to distinguish the differences in clinical value among loci of the same grade.
[0050] 2. GWAS (Generation-wide Association Study), which uses a p-value < 5 × 10⁻ 8 Screening for disease-associated loci.
[0051] This technology has the following limitations: it is only applicable to common variants (MAF > 1%), and the false negative rate for rare disease sites is > 67%; it ignores functional annotation evidence (such as changes in protein structure).
[0052] 3. The Blockchain Gene Bank (HGBC) is a distributed storage system for genetic data that ensures security.
[0053] The technology has the following drawbacks: it does not integrate an AI analysis engine and relies on manual annotation; the accuracy of phenotypic-genotypic association analysis is only 58%.
[0054] In summary, existing technologies have the following shortcomings: Lack of multi-source evidence integration: Current technologies cannot simultaneously process multimodal evidence chains such as literature, databases, and functional annotations (e.g., the Alzheimer's disease risk of rs429358 requires integration of 27 types of evidence). Insufficient dynamic adaptability: Assessment of new variant sites in sudden situations is delayed by more than 4 months. Weak clinical interpretability: Physicians find it difficult to understand the site selection logic of AI models, resulting in a clinical adoption rate of <40%.
[0055] To address the aforementioned shortcomings, this disclosure proposes an artificial intelligence-based variant site prediction and screening system and method. The scheme of this disclosure realizes automated functions such as literature retrieval, evaluation index extraction, site scoring, and variant site screening. It can not only effectively improve the efficiency, timeliness, and accuracy of variant site screening, but also effectively screen out high-quality variant sites, so as to provide important support for disease risk prediction, personalized medicine, and clinical gene testing.
[0056] The technical solution of this disclosure will be described below through specific embodiments.
[0057] like Figure 1 The diagram shown is a structural schematic of the AI-based variant site prediction and screening system of this embodiment. Specifically, the system of this embodiment includes: a literature retrieval module 110, an evaluation index extraction module 120, a scoring module 130, a variant site screening module 140, and a user interaction module 150.
[0058] The literature retrieval module 110 is used to capture data from multiple sources of evidence and construct a medical knowledge graph for predicting variant locations based on the captured data. Specifically, a trained graph construction model is used to perform entity recognition and relation extraction on the captured data, and the medical knowledge graph is constructed based on the recognized entities and extracted relations.
[0059] The evaluation index extraction module 120 is used to extract evaluation indicators using the medical knowledge graph and determine the dynamic weights of each evaluation indicator based on scene tags and time decay factors.
[0060] The scoring module 130 is used to determine the risk score value of each site in the gene sequence to be detected based on each evaluation index and the dynamic weight corresponding to each evaluation index.
[0061] The variant site screening module 140 is used to screen out abnormal sites from each site based on the risk score.
[0062] The user interaction module 150 is used to display the abnormal points and show the text boxes for modifying the dynamic weights corresponding to each evaluation indicator, so that the user can input the correction value of the dynamic weights corresponding to each evaluation indicator.
[0063] In some instances, graph construction models perform entity recognition on the crawled data based on part-of-speech tagging; and multi-head attention mechanisms are used for relation extraction.
[0064] The literature retrieval module disclosed herein can dynamically retrieve data from PubMed / EMBASE and other sources, and uses the BERT-BiLSTM model to construct a medical knowledge graph.
[0065] Multi-source evidence can include the following:
[0066] Literature databases: PubMed / EMBASE (batch retrieval of full-text studies related to variant sites from 2020 to 2025 via API).
[0067] Public databases: ClinVar, dbSNP, COSMIC (structured variation data), UniProt (protein function annotation).
[0068] Clinical data: TCGA, ICGC (cancer genomics data), EHR (electronic health record, which needs to be anonymized).
[0069] Preprints: bioRxiv / medRxiv (latest unpublished research, reliability level needs to be indicated).
[0070] like Figure 2 As shown, multi-source evidence can also include existing international data collection standards, domestic industry data collection standards, and parts for which there are no international or domestic standards.
[0071] After the data is captured, it needs to be cleaned and standardized. Data cleaning is used to remove low-quality literature (such as unverified conclusions in preprints) and conflicting or variant annotations (which are resolved through a voting mechanism).
[0072] Standardization includes: Gene nomenclature: unifying to HGNC symbols (e.g., TP53 instead of p53). Variance description: following HGVS rules (e.g., c.524G>A instead of rs121913254). Entity alignment: linking to a unified ID using a Biomedical Entity Linker (e.g., the improved version GNormPlus in 2025).
[0073] The extracted entities can include genes, variants, diseases, drugs, etc.; the extracted relationships can include variant-cause-disease, drug-target-variant, literature-support-association, etc.
[0074] The specific applications of the BERT-BiLSTM joint model include the following:
[0075] 1) Text embedding
[0076] Use BioBERT-2025 (fine-tuned based on the latest PubMed literature) to extract contextual embeddings from literature abstracts.
[0077] Input: Sentence-level text (e.g., “BRAF V600E mutation leads to melanoma”).
[0078] Output: 768-dimensional vector (CLS token for classification tasks).
[0079] 2) Sequence labeling
[0080] BiLSTM-CRF layer identifies entities and relationships:
[0081] Input: BioBERT embedding + part-of-speech tagging (Spacy 4.0 Medical Edition).
[0082] Output: Entities annotated with BIOES (e.g.) <variant> BRAF V600E< / variant> ).
[0083] Relation extraction: Employ a multi-head attention mechanism (e.g., F1 ≥ 0.92 for "cause" relations).
[0084] 3) Map filling:
[0085] Integrate model output with structured data such as ClinVar pathogenicity labels.
[0086] 4) Model optimization:
[0087] Adversarial training: Adding gradient penalty (WGAN-GP) improves the generalization of small sample relationships (such as rare mutations).
[0088] Dynamic negative sampling: Adjusting the weights of the loss function for long-tailed entities (such as non-coding region mutations).
[0089] In addition, TF-IDF keyword matching can be used to crawl data.
[0090] In some embodiments, when determining the dynamic weights corresponding to each evaluation indicator based on scene labels and time decay factors, the evaluation indicator extraction module 120 is specifically used for:
[0091] Each evaluation index is normalized; based on the publication time of the literature, the time decay factor is determined, and the scene label is determined; the normalized evaluation index, the time decay factor, and the scene label are input into the trained weight dynamic adjustment model; the dynamic weights of each evaluation index output by the weight dynamic adjustment model are obtained.
[0092] The steps above map different dimensional indicators to the [0,1] interval. The time decay factor is used to represent the importance of time to the indicator, and the scenario labels include cancer screening, rare disease diagnosis, etc.
[0093] In this disclosure, the dynamic weights can be dynamically adjusted by the gradient boosting tree.
[0094] In addition, the present disclosure can also utilize genetic algorithms to optimize dynamic weights.
[0095] In some instances, the evaluation metrics include: functional evidence, clinical relevance, population frequency, evolutionary conservation, literature support, and study confidence.
[0096] In some instances, when determining the risk score value for each locus in the gene sequence to be detected based on each evaluation indicator and its corresponding dynamic weight, the scoring module is specifically used for:
[0097] The characteristics, evaluation indicators, and dynamic weights of each site in the gene sequence to be tested are mapped to the target space; the source domain risk score and the target domain data sufficiency are determined; using the source domain risk score and the target domain data sufficiency, as well as the characteristics, evaluation indicators, and dynamic weights of each site in the target space, the risk score value of each site in the gene sequence to be tested is determined.
[0098] The risk score calculation formula mentioned above is as follows:
[0099] Risk score = ∑(evaluation index i × dynamic weight i) + migration compensation item;
[0100] Migration compensation item = Source domain risk score × (1 - Target domain data sufficiency).
[0101] The source domain risk score can be determined based on population genetics data or using a trained machine learning model. The sufficiency of the target domain data can be determined based on multimodal data coverage.
[0102] In this disclosure, the scoring module can incorporate transfer learning to achieve scoring, and the risk score ranges from 0 to 10.
[0103] In some instances, when the variant site screening module selects anomalous sites from various sites based on the risk score, it is specifically used for:
[0104] For each location, the clinical value is determined using the corresponding risk score and disease burden weight; and the technical cost is determined using experimental validation cost and clinical validation cost; an optimization model is constructed using the clinical value and technical cost, and the optimization model is solved to obtain the outlier score for that location; locations with outlier scores higher than the preset value are identified as outlier locations.
[0105] The above screening steps are based on maximizing clinical value versus minimizing the cost of technical validation.
[0106] In some instances, the user interaction module is also used to: display the medical knowledge graph, various evaluation indicators, and the dynamic weights corresponding to each evaluation indicator; generate a heat map based on the abnormal scores of each location; and display the heat map.
[0107] In addition, the user interaction module is also used to visualize the decision tree to display the site scoring criteria and support clinicians to provide feedback and correct dynamic weights.
[0108] The embodiments disclosed herein break through the limitations of traditional single-dimensional rating, achieving real-time fusion analysis of multimodal evidence chains including literature, databases, clinical data, and functional annotations. A six-dimensional dynamic weighting system is established: a quantitative evaluation framework is built upon functional evidence (E_func), clinical relevance (E_clin), population frequency (E_freq), evolutionary conservation (E_evol), literature support (E_pub), and study confidence (E_conf). The weighting coefficients are dynamically optimized through transfer learning (source domain: millions of known sites in gnomAD → target domain: novel variants). BERT-BiLSTM literature mining: high-speed parsing of PubMed / EMBASE literature generates a medical knowledge graph, solving the problems of low efficiency and lagging terminology updates in manual retrieval.
[0109] Furthermore, this disclosure implements the Pareto frontier screening algorithm for variant sites, balancing the maximization of clinical value with the minimization of technical validation costs, and resolving the contradiction of "high recall accompanied by high false positives" in existing technologies. A visual decision tree feedback mechanism allows clinicians to adjust weight parameters, driving the model's adaptive iteration.
[0110] Through a full-chain innovation of "dynamic integration of multi-source evidence → Pareto cutting-edge screening → dual-chain evidence storage and clinical application", it solves three major industry pain points in gene mutation site screening: evidence fragmentation, data silos, and difficulty in clinical translation. It provides a highly reliable site library for disease risk prediction, breaks through the "last mile" from site screening to clinical application, and improves the diagnostic adoption rate.
[0111] The embodiments disclosed herein achieve the following beneficial effects:
[0112] 1) Increased efficiency: Site screening speed increased by 300 times (3 hours → 36 seconds / site)
[0113] The completeness of the literature analysis increased from 68% to 99% (test set: 1000 BRCA1 sites).
[0114] 2) Precision breakthrough
[0115] index Traditional solution This disclosed solution Rare disease recall rate 29% 60% Clinical misjudgment rate 18% 5%
[0116] 3) Ecosystem compatibility: Supports hybrid deployment of Kubernetes / Spark, improving resource utilization by 70%.
[0117] Based on the same inventive concept, this disclosure provides an artificial intelligence-based method for predicting and screening mutation sites. The steps performed by this method are the same as or similar to the components in the aforementioned system; therefore, similar details will not be repeated. Figure 3 As shown, the AI-based mutation site prediction and screening method in this embodiment includes:
[0118] S310. Use the literature retrieval module to capture data from multiple sources, and construct a medical knowledge graph for predicting variant locations based on the captured data; wherein, use the trained graph construction model to perform entity recognition and relation extraction on the captured data, and construct the medical knowledge graph based on the recognized entities and extracted relations.
[0119] S320. The evaluation index extraction module extracts evaluation indicators based on the medical knowledge graph, and determines the dynamic weight of each evaluation index based on the scene label and time decay factor.
[0120] S330. Using the scoring module, the risk score value of each site in the gene sequence to be detected is determined based on each evaluation indicator and the dynamic weight corresponding to each evaluation indicator.
[0121] S340. Using the mutation site screening module, abnormal sites are selected from each site based on the risk score value.
[0122] S350. The abnormal points are displayed using the user interaction module, and a text box for modifying the dynamic weights corresponding to each evaluation indicator is displayed so that the user can input the correction value for the dynamic weights corresponding to each evaluation indicator.
[0123] The various embodiments of the techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0124] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0125] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0126] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0127] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0128] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0129] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0130] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A mutation site prediction and screening system based on artificial intelligence, characterized in that, It includes a literature retrieval module, an evaluation index extraction module, a scoring module, a variant site screening module, and a user interaction module; The literature retrieval module is used to capture data from multiple sources and construct a medical knowledge graph for predicting variant locations based on the captured data. Specifically, a trained graph construction model is used to perform entity recognition and relation extraction on the captured data, and the medical knowledge graph is constructed based on the recognized entities and extracted relations. The evaluation index extraction module is used to extract evaluation indicators using the medical knowledge graph and determine the dynamic weight of each evaluation indicator based on scene tags and time decay factors. The scoring module is used to determine the risk score value of each site in the gene sequence to be detected based on each evaluation index and the dynamic weight corresponding to each evaluation index. The variant site screening module is used to screen out abnormal sites from each site based on the risk score value; The user interaction module is used to display the abnormal points and show the text boxes for modifying the dynamic weights of each evaluation indicator, so that the user can input the correction value of the dynamic weights of each evaluation indicator. The scoring module, when determining the risk score value of each locus in the gene sequence to be detected based on each evaluation indicator and its corresponding dynamic weight, is specifically used for: The features of each site in the gene sequence to be detected, each evaluation index, and the dynamic weights corresponding to each evaluation index are mapped to the target space. Determine the risk score of the source domain and the sufficiency of the target domain data; By utilizing source domain risk scores and target domain data sufficiency, as well as the characteristics of each location in the target space, each evaluation index, and the dynamic weights corresponding to each evaluation index, the risk score value of each location in the gene sequence to be detected is determined. The target domain is the newly emerging variant; the data sufficiency of the target domain is used to characterize the coverage of multimodal data.
2. The artificial intelligence-based mutation site prediction and screening system according to claim 1, characterized in that, The graph construction model performs entity recognition on the captured data based on part-of-speech tagging; and uses a multi-head attention mechanism for relation extraction.
3. The artificial intelligence-based mutation site prediction and screening system according to claim 1, characterized in that, The evaluation index extraction module, when determining the dynamic weights of each evaluation index based on scene labels and time decay factors, is specifically used for: The evaluation indicators were normalized. Based on the publication time of the literature, determine the time decay factor, and determine the scene label; The normalized evaluation indicators, time decay factor and scene label are input into the trained weights to dynamically adjust the model. Obtain the dynamic weights of each evaluation index output by the weight dynamic adjustment model.
4. The artificial intelligence-based mutation site prediction and screening system according to claim 3, characterized in that, The evaluation indicators include: functional evidence, clinical relevance, population frequency, evolutionary conservation, literature support, and study confidence.
5. The artificial intelligence-based mutation site prediction and screening system according to claim 1, characterized in that, The risk score is calculated using the following formula: Risk score = ∑(evaluation index i × dynamic weight i) + migration compensation item; Migration compensation item = Source domain risk score × (1 - Target domain data sufficiency).
6. The artificial intelligence-based mutation site prediction and screening system according to claim 1, characterized in that, When the mutation site screening module selects abnormal sites from each site based on the risk score, it is specifically used for: For each location, the clinical value is determined using the corresponding risk score and disease burden weight; and the technical cost is determined using experimental validation cost and clinical validation cost; an optimization model is constructed using the clinical value and technical cost, and the optimization model is solved to obtain the outlier score for that location. Points with abnormal scores higher than the preset value are designated as abnormal points.
7. The artificial intelligence-based mutation site prediction and screening system according to claim 1, characterized in that, The user interaction module is also used for: It displays a medical knowledge graph, various evaluation indicators, and the dynamic weights corresponding to each evaluation indicator.
8. The artificial intelligence-based mutation site prediction and screening system according to claim 6, characterized in that, The user interaction module is also used for: A heat map is generated based on the anomaly scores of each location; The heatmap is shown.
9. A method for predicting and screening mutation sites based on artificial intelligence, characterized in that, include: The literature retrieval module is used to capture data from multiple sources, and a medical knowledge graph for predicting variant locations is constructed based on the captured data. Specifically, a trained graph construction model is used to perform entity recognition and relation extraction on the captured data, and the medical knowledge graph is constructed based on the recognized entities and extracted relations. The evaluation index extraction module extracts evaluation indicators based on the medical knowledge graph, and determines the dynamic weight of each evaluation index based on scene tags and time decay factors. The scoring module is used to determine the risk score value of each site in the gene sequence to be tested based on each evaluation indicator and the dynamic weight of each evaluation indicator. The abnormal sites are selected from each site based on the risk score using the variant site screening module. The abnormal points are displayed using the user interaction module, and a text box for modifying the dynamic weights of each evaluation indicator is shown, so that the user can input the correction value of the dynamic weights of each evaluation indicator. Specifically, when the scoring module determines the risk score value for each locus in the gene sequence to be detected based on each evaluation indicator and its corresponding dynamic weight, The features of each site in the gene sequence to be detected, each evaluation index, and the dynamic weights corresponding to each evaluation index are mapped to the target space. Determine the risk score of the source domain and the sufficiency of the target domain data; By utilizing source domain risk scores and target domain data sufficiency, as well as the characteristics of each location in the target space, each evaluation index, and the dynamic weights corresponding to each evaluation index, the risk score value of each location in the gene sequence to be detected is determined. The target domain is the newly emerging variant; the data sufficiency of the target domain is used to characterize the coverage of multimodal data.
Citation Information
Patent Citations
Risk prediction variation site screening evaluation system and method based on artificial intelligence
CN119007825A
Disease science popularization error correction method and system based on artificial intelligence
CN120108774A