A Construction Method, Device, Equipment and Medium for a Spatial Knowledge Graph Database of Stroke

By using the Tongyi Qianwen 2.1 model and bioinformatics tool to construct a stroke spatial knowledge map database, the problems of low data integration efficiency, insufficient labeling accuracy and insufficient dynamic correlation in stroke research were solved, and efficient and accurate gene spatial distribution analysis was achieved.

CN119626575BActive Publication Date: 2025-07-22INNOVATION CENTER OF YANGTZE RIVER DELTA ZHEJIANG UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510162109.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-07-22
Estimated Expiration
2045-02-14

AI Technical Summary

Technical Problem

There are problems in existing stroke studies with low data integration efficiency, insufficient labeling accuracy and insufficient dynamic correlation. Especially when analyzing the association of genes and spatial regions, traditional methods take time, high cost, and are difficult to ensure data consistency, and lack automated interaction and dynamic analysis capabilities.

Method used

The Tongyi Qianwen 2.1 model was used to automatically extract gene information from the literature, combine the STRING database to build a protein interaction network, and use tools such as Metascape and KEGG to perform pathway enrichment analysis to construct a dynamically updated stroke spatial knowledge map database.

Benefits of technology

It has achieved efficient integration of multi-source heterogeneous data, accurately labeled the spatial distribution of genes, and built a dynamic updated knowledge map, which has improved the data integration efficiency and labeling accuracy of spatial heterogeneity research in stroke, and supported dynamic correlation analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119626575B_ABST
    Figure CN119626575B_ABST
Patent Text Reader

Abstract

The present application discloses a method, apparatus, device and medium for constructing a spatial knowledge graph database of stroke, relating to the technical field of big data analysis. The present application utilizes the Tongyi Qianwen 2.1 model to deeply mine and analyze publicly available papers, identify and label the spatial distribution of genes, physiological regions, and pathological regions related to ischemic stroke, and display them in the form of a spatial knowledge graph database of stroke. The present application utilizes the natural language processing ability of the Tongyi Qianwen 2.1 model to automatically extract genes and their spatial distribution information from the literature and constructs a dynamically updated knowledge graph, solving the problems of low data integration efficiency, insufficient annotation accuracy, and insufficient dynamic association in the study of stroke spatial heterogeneity in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of big data analysis technology, and in particular to a method, device, equipment and medium for constructing a stroke spatial knowledge graph database. Background Art

[0002] With the acceleration of global population aging, stroke has become one of the major public health challenges facing the world. It is not only the main cause of disability, but also the leading cause of death in the world. Although a large number of studies have made significant progress in the understanding of stroke in the past few decades, there is still a lack of in-depth understanding of its complex disease mechanisms, especially the spatial heterogeneity shown in the development of stroke. This spatial heterogeneity is reflected in the pathological changes in different stages of stroke, different brain regions and their microenvironments, which is of great significance for the diagnosis, treatment and prognosis of stroke.

[0003] In existing stroke research, spatial heterogeneity analysis faces the following technical bottlenecks.

[0004] (1) Low data integration efficiency: Gene-region association information is scattered in massive literature. Traditional manual extraction methods are time-consuming, costly, and difficult to ensure data consistency.

[0005] (2) Insufficient labeling accuracy: The labeling of physiological areas (such as hippocampus and cortex) and pathological areas (such as the core area of infarction) relies on expert experience and is easily affected by subjective factors.

[0006] (3) Lack of dynamic analysis capabilities: Existing databases lack automated interaction with bioinformatics tools (such as STRING and KEGG), and are unable to construct protein interaction networks and pathway enrichment analysis in real time, which limits the in-depth exploration of spatial heterogeneity mechanisms.

[0007] The existing technologies for studying spatial heterogeneity of stroke have low data integration efficiency, insufficient annotation accuracy, and insufficient dynamic association, which has become a technical problem that needs to be solved urgently. Summary of the invention

[0008] The purpose of this application is to provide a method, device, equipment and medium for constructing a stroke spatial knowledge graph database, so as to achieve efficient integration of multi-source heterogeneous data through automated tools, accurately annotate the spatial distribution of genes, and construct a dynamically updated knowledge graph, so as to solve the problems of low data integration efficiency, insufficient annotation accuracy and insufficient dynamic association in the existing technology of stroke spatial heterogeneity research.

[0009] To achieve the above objectives, this application provides the following solutions.

[0010] In a first aspect, the present application provides a method for constructing a spatio - knowledge graph database of stroke, including the following steps.

[0011] Obtain papers related to ischemic stroke.

[0012] Use the Tongyi Qianwen 2.1 model to identify the genes related to ischemic stroke and the description information of the genes recorded in each paper, and output them in the form of structured data; the structured data includes gene names, description information, and sources.

[0013] Use the Tongyi Qianwen 2.1 model to analyze the papers recording each structured data, perform physiological region annotation and pathological region annotation on each structured data, and add the physiological region annotation and pathological region annotation to the structured data respectively to construct a structured data set.

[0014] Clean the structured data set, delete the invalid and duplicate structured data in the structured data set, obtain the cleaned structured data set, and output the cleaned structured data set in the form of an Excel table to obtain Excel table data.

[0015] According to the Excel table data, use the STRING database to construct a protein - protein interaction network.

[0016] According to the Excel table data, use online bioinformatics tools to analyze the distribution differences of genes in different physiological regions and pathological regions, and determine the functional annotations and pathway enrichments of each gene in the Excel table data.

[0017] According to the Excel table data, the protein - protein interaction network, and the functional annotations and pathway enrichments of each gene in the Excel table data, construct a spatio - knowledge graph database of stroke.

[0018] In a second aspect, the present application provides a device for constructing a spatio - knowledge graph database of stroke. The device for constructing a spatio - knowledge graph database of stroke applies the above - mentioned method for constructing a spatio - knowledge graph database of stroke. The device for constructing a spatio - knowledge graph database of stroke includes the following modules.

[0019] A paper acquisition module, configured to obtain papers related to ischemic stroke.

[0020] A gene recognition module, configured to use the Tongyi Qianwen 2.1 model to identify the genes related to ischemic stroke and the description information of the genes recorded in each paper, and output them in the form of structured data; the structured data includes gene names, description information, and sources.

[0021] The region annotation module is used to analyze the papers recording each structured data by using Tongyi Qianwen 2.1 model, perform physiological region annotation and pathological region annotation on each structured data, and add the physiological region annotation and pathological region annotation to the structured data respectively to construct a structured data set.

[0022] The cleaning module is used to clean the structured data set, delete the invalid and duplicate structured data in the structured data set, obtain the cleaned structured data set, and output the cleaned structured data set in the form of an Excel table to obtain Excel table data.

[0023] The protein-protein interaction network construction module is used to construct a protein-protein interaction network according to the Excel table data by using the STRING database.

[0024] The distribution difference analysis module is used to analyze the distribution differences of genes in different physiological regions and pathological regions according to the Excel table data by using online bioinformatics tools, and determine the functional annotations and pathway enrichments of each gene in the Excel table data.

[0025] The database construction module is used to construct a stroke spatial knowledge graph database according to the Excel table data, the protein-protein interaction network, and the functional annotations and pathway enrichments of each gene in the Excel table data.

[0026] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the above-mentioned method for constructing a stroke spatial knowledge graph database.

[0027] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above-mentioned method for constructing a stroke spatial knowledge graph database is implemented.

[0028] According to the specific embodiments provided by the present application, the present application has the following technical effects.

[0029] This application provides a method, device, equipment, and medium for constructing a spatial knowledge graph database of stroke. This application uses the Tongyi Qianwen 2.1 model to deeply mine and analyze publicly available papers, identify and label the spatial distribution of genes, physiological regions, and pathological regions related to ischemic stroke (the spatial distribution refers to the specific localization of genes in the physiological regions (such as the cortex, hippocampus) and pathological regions (such as the infarct core, ischemic penumbra) where genes play a role in stroke diseases. Through annotation and dynamic analysis, the functional differences of genes in different spatial environments are revealed), and it is presented in the form of a spatial knowledge graph database of stroke. This application uses the natural language processing ability of the Tongyi Qianwen 2.1 model to automatically extract genes and their spatial distribution information from the literature and constructs a dynamically updated knowledge graph, solving the problems of low data integration efficiency, insufficient annotation accuracy, and lack of dynamic association in the study of stroke spatial heterogeneity in the prior art.

[0030] This application uses the natural language processing ability of the Tongyi Qianwen 2.1 model to automatically extract genes and their spatial distribution information from the literature and generate a structured data set (to solve the problem of data dispersion).

[0031] This application identifies physiological / pathological regions through a pre-trained deep neural network model and improves the annotation accuracy through multiple rounds of data cleaning (such as keyword matching, standard gene set verification) (to solve the problem of annotation complexity).

[0032] This application constructs a protein-protein interaction network based on the STRING database, combines tools such as Metascape and KEGG to achieve pathway enrichment analysis, and forms a dynamically updated knowledge graph (to solve the lack of dynamic analysis ability). BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0034] Figure 1 It is a schematic flowchart of a method for constructing a spatial knowledge graph database of stroke provided by an embodiment of this application.

[0035] Figure 2 It is a data collection and mining flowchart provided by an embodiment of this application.

[0036] Figure 3 It is a schematic diagram of the data proofreading result provided by an embodiment of this application.

[0037] Figure 4Schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0038] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0039] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0040] In an exemplary embodiment, as Figure 1 shown, a method for constructing a spatio-temporal knowledge graph database for stroke is provided, including the following steps 101 to 107.

[0041] Step 101, obtain papers related to ischemic stroke.

[0042] Step 102, use the Tongyi Qianwen 2.1 model to identify the genes related to ischemic stroke and the description information of the genes recorded in each paper, and output them in the form of structured data; the structured data includes gene names, description information, and sources.

[0043] Step 103, use the Tongyi Qianwen 2.1 model to analyze the papers recording each structured data, perform physiological region annotation and pathological region annotation on each structured data, and add the physiological region annotation and pathological region annotation to the structured data respectively to construct a structured data set.

[0044] Step 104, clean the structured data set, delete the invalid and duplicate structured data in the structured data set, obtain the cleaned structured data set, and output the cleaned structured data set in the form of an Excel table to obtain Excel table data.

[0045] Step 105, according to the Excel table data, use the STRING database to construct a protein-protein interaction network.

[0046] Step 106, according to the Excel table data, use an online bioinformatics tool to analyze the distribution differences of genes in different physiological regions and pathological regions, and determine the functional annotations and pathway enrichments of each gene in the Excel table data.

[0047] Step 107: Construct a spatial knowledge graph database for stroke based on Excel table data, protein-protein interaction networks, and functional annotations and pathway enrichments of each gene in the Excel table data.

[0048] Implementing the above Steps 101 to 107 can systematically sort out and summarize the knowledge related to spatial regions in stroke diseases.

[0049] This application aims to construct a comprehensive, experimentally verified, and traceable knowledge graph for ischemic stroke diseases.

[0050] In another exemplary embodiment of this application, the keywords related to stroke are determined as "ischemiastroke", "cerebral ischemia", "cerebral infarct", and "MCAO". Based on the keywords "ischemiastroke", "cerebral ischemia", "cerebral infarct", and "MCAO", papers published in the past 20 years from January 2003 to December 2023 related to ischemic stroke are retrieved from the PubMed database. The full PDFs of the open access (OA) papers are downloaded and uniformly named as "Stroke + serial number" for subsequent convenient invocation by large language models and in-depth text mining analysis.

[0051] Among them, the PubMed database is the most widely used free MEDLINE retrieval tool on the Internet. Among them, the MEDLINE retrieval tool is a WEB-based biomedical information retrieval system developed by the National Center for Biotechnology Information affiliated with the National Library of Medicine of the United States in April 2000. The PubMed database contains more than 32 million biomedical literature and abstracts.

[0052] Exemplarily, in the embodiments of the present application, 139,267 academic papers on ischemic stroke published during the period from January 2003 to December 2023 were collected from the PubMed database. During the detailed review of the downloaded papers, a total of 92,610 papers were attached with Digital Object Identifiers (DOIs). The papers with DOIs included OA papers and papers obtained by other means. Among them, 56,903 were open access papers, and the full text could be directly downloaded for further analysis. Using the knowledge extraction technology of Tongyi Qianwen 2.1 model, gene information related to ischemic stroke and response space information were identified from these OA papers, including 1125 human genes, 1094 mouse genes, etc. Among them, there were 55 genes related to the infarct core area, 89 genes related to the ischemic penumbra, 36 genes related to the cortex, and 9 genes related to the hippocampus. However, there were fewer independent studies on other brain regions such as the striatum and thalamus. Finally, important information such as their gene names, species, references, detailed descriptions, gene locations, and synonyms were accurately recorded, as Figure 2 shown, Figure 2 where A in Figure 2 is a schematic diagram of the information mining process of the Tongyi Qianwen 2.1 model; Figure 2 where B in

[0053] is all the literature information, divided into the total number of retrieved literatures (ALL column), the number of literatures with DOI numbers (DOI column), and the number of OA papers (OA column);

[0054] In another exemplary embodiment, in step 102 above, based on the PDF document named "Stroke + serial number" obtained in step 101, the downloaded literature was deeply analyzed using the Tongyi Qianwen 2.1 model to identify genes related to ischemic stroke mentioned in the literature, specifically including the following steps 201 - step 204.

[0055] Step 201, data preprocessing.

[0056] Unify the downloaded PDF documents named "Stroke + serial number" into text format so that the Tongyi Qianwen 2.1 model can process them. During the conversion process, it is necessary to ensure the integrity and accuracy of the text, retain the chart titles, annotations, etc. in the document that may contain gene information, and minimize text garbling or information loss caused by format conversion.

[0057] Step 202, model input and task setting.

[0058] Input the converted text into the Tongyi Qianwen 2.1 model piece by piece. When inputting, clearly inform the Tongyi Qianwen 2.1 model that the task is to identify gene information related to ischemic stroke, including gene names, gene mutations (if mentioned in the text), and descriptions of the association between genes and stroke. For example, an instruction like "Please extract all genes related to ischemic stroke and their descriptive information from the paper" can be given to the Tongyi Qianwen 2.1 model.

[0059] Step 203, model identification.

[0060] The Tongyi Qianwen 2.1 model performs gene identification based on its deep neural network architecture and large-scale pre-trained knowledge system. During the pre-training process, the model has learned a large number of text patterns and knowledge in the medical and biological fields, and can identify specific vocabulary, terms related to genes, and their context relationships in the text.

[0061] When processing the input text, the Tongyi Qianwen 2.1 model uses natural language processing functions to perform natural language processing operations such as word segmentation and part-of-speech tagging on the text. Then, through the multi-layer calculation of the neural network, the deep neural network gene identification function matches the vocabulary in the text with the gene-related patterns learned during pre-training. For example, the Tongyi Qianwen 2.1 model can identify common gene name abbreviations (such as TNF-α representing tumor necrosis factor-α), full gene names (such as brain-derived neurotrophic factor representing brain-derived neurotrophic factor), and specific expressions when describing the relationship between gene functions, mechanisms of action, and ischemic stroke (such as gene X is upregulated in the ischemic stroke model and is related to neuron protection).

[0062] For the identified possible gene information, the Tongyi Qianwen 2.1 model will further analyze its context to confirm the relevance of the gene to ischemic stroke and extract relevant descriptive information. For example, if it is mentioned in a text that "in the brain tissue samples of ischemic stroke patients, the mutation frequency of gene Y is found to be relatively high and is related to the abnormal activation of the apoptosis pathway", the Tongyi Qianwen 2.1 model can accurately extract the name of gene Y, the mutation situation, and its association with the apoptosis pathway and ischemic stroke.

[0063] Step 204, Result Output and Arrangement.

[0064] The Tongyi Qianwen 2.1 model outputs the gene information related to ischemic stroke in a structured form. The output results may include a list of gene names, the corresponding relevant descriptions for each gene (such as its role in ischemic stroke, expression changes, etc.), and the source of the gene in the literature (such as the title, page number, etc. of the literature where it is located for subsequent verification).

[0065] Arrange and preliminarily screen the results output by the Tongyi Qianwen 2.1 model to remove possible misidentifications or irrelevant information. For example, exclude some gene information mentioned in other unrelated diseases or normal physiological processes but not directly related to ischemic stroke, and at the same time merge and deduplicate the repeatedly identified gene information to ensure the accuracy and conciseness of the gene list and provide reliable basic data for subsequent analysis and research.

[0066] The specific methods for arranging and preliminarily screening the results output by the Tongyi Qianwen 2.1 model include the following steps A1 - A3.

[0067] A1. Filtering of Misidentifications.

[0068] Match the preset stroke keywords (such as 'ischemia stroke', 'neuronal apoptosis', 'infarct core area') with the description information output by the model, and delete the entries that do not contain keywords or are semantically irrelevant;

[0069] Use natural language processing techniques (such as semantic similarity calculation) to identify entries with vague descriptions or contradictions with the context (such as 'Gene X is highly expressed in normal brain tissue'), and mark them as data to be reviewed.

[0070] A2. Merging of Duplicate Data.

[0071] Adopt a string matching algorithm based on edit distance (such as the Levenshtein distance algorithm) to compare the gene names with the NCBI standard gene set and merge different expressions of the same gene (such as 'BDNF' and 'Brain - derived neurotrophic factor').

[0072] For duplicate gene entries (same gene name, same description, same source), only keep the record of the first occurrence, and append the DOI numbers of other literatures in the 'Source' field.

[0073] A3. Manual Review and Verification.

[0074] The data screened and selected independently by two researchers was focused on verifying the accuracy of gene names and descriptions, the completeness of sources, and inconsistent entries were resolved through cross-comparison.

[0075] The data that passed the review was incorporated into the cleaned structured dataset, while the data that did not pass was returned to the model for reprocessing or directly excluded.

[0076] The Tongyi Qianwen 2.1 model normalizes diverse expressions of the same gene in different literatures into standard terms through semantic analysis and term mapping, while retaining the differential information of the original descriptions, forming a structured "main description - additional description - multiple sources" data framework.

[0077] In another exemplary embodiment, in step 103 above, based on the genes related to ischemic stroke diseases obtained from the above steps, using the Tongyi Qianwen 2.1 model, the model further analyzes the context information in the literature to identify and label whether each gene is experimentally studied in relevant brain regions, such as the hippocampus, cortex, striatum, etc. In addition to the annotation of physiological regions, it also focuses on identifying the pathological partitions of the genes mentioned in the literature, such as the infarct core area and the ischemic penumbra. This step is crucial for understanding how genes function in different physiological / pathological brain region environments and also reveals the specific regions that may be affected by stroke diseases.

[0078] Among them, the annotation bases for physiological regions and pathological regions include:

[0079] (1) Anatomical terms in medical knowledge bases (such as Allen Brain Atlas, AHA stroke classification).

[0080] (2) Experimental descriptions in the literature context (such as "hippocampal slice analysis", "gene expression in the infarct core area").

[0081] (3) The pathological feature library pre-trained by the model.

[0082] In another exemplary embodiment, the above-mentioned structured data includes gene names, descriptions, species, physiological regions, pathological regions, literature sources, and functional annotations, where the physiological / pathological region fields are directly associated with brain region spatial distribution information (such as the cortex, hippocampus, infarct core area).

[0083] This structured data not only records the association information between genes and brain regions, but also realizes in-depth analysis through the following steps B1 - B3.

[0084] B1. Construct a protein - protein interaction network based on the STRING database to analyze the functional synergy of genes in specific brain regions.

[0085] B2. Use Metascape and KEGG tools for pathway enrichment to clarify the regulatory mechanism of genes in pathological regions.

[0086] B3. Dynamically associate multi-dimensional data (such as gene expression levels and experimental verification results) through a knowledge graph to support spatial heterogeneity research.

[0087] In another exemplary embodiment, step 104 cleans the structured data set obtained in the above steps, matches it with the standard gene set of a certain biotechnology information center, removes irrelevant items, and deletes duplicate items, ensuring the accuracy and consistency of the data. Subsequently, to ensure the convenience of subsequent analysis and display, the cleaned data is formatted and organized into an Excel table, which details key information such as gene names, species information, relevant brain region locations, and pathological region divisions, ensuring that the data has a certain logical structure and readability. Specifically, it includes the following steps 301 - step 305.

[0088] Step 301, check the integrity of each structured data in the structured data set, delete incomplete structured data, and obtain the structured data set after the first cleaning.

[0089] Step 302, match the description information of each structured data in the structured data set after the first cleaning with the preset stroke keywords, and delete the structured data in the structured data set after the first cleaning whose description information does not match the preset stroke keywords, to obtain the structured data set after the second cleaning.

[0090] Step 303, uniformly standardize each structured data in the structured data set after the second cleaning to obtain the standardized structured data set.

[0091] Step 304, compare each structured data in the standardized structured data set, and delete duplicate structured data in the standardized structured data set to obtain the structured data set after the third cleaning.

[0092] Step 305, match the gene names of each structured data in the structured data set after the third cleaning with the standard gene set, and delete the structured data in the structured data set after the third cleaning whose gene names do not match the standard gene set to obtain the cleaned structured data set.

[0093] The specific cleaning process includes the following steps 401 - step 403.

[0094] Step 401, data screening.

[0095] First, conduct a preliminary screening of the original data extracted from a large number of literature. Check the integrity of the data and remove the data records that are obviously incomplete or damaged. For example, if some data items lack key information, such as unclear gene names or unlabeled brain region locations, they will be excluded.

[0096] Meanwhile, screen out the data related to stroke according to specific rules. For example, through keyword matching, only retain the data records that contain specific stroke-related keywords (such as "stroke", "cerebrovascular accident", "cerebral ischemia", etc.).

[0097] Step 402, format standardization.

[0098] Conduct unified standardization processing on the data format. The data formats in different literatures may vary, which will bring difficulties to subsequent analysis. Therefore, unify and standardize the formats of key information such as gene names, species information, brain region locations, and pathological region divisions.

[0099] For example, for gene names, uniformly adopt the internationally common nomenclature; for species information, adopt the standard taxonomic naming method; for brain region locations and pathological region divisions, use unified anatomical terms for annotation.

[0100] Step 403, remove noise data.

[0101] Use natural language processing techniques and machine learning algorithms to identify and remove noise data. Noise data may include incorrect annotations, unclear descriptions, or information unrelated to the theme.

[0102] Among them, the specific methods for identifying and removing noise data include the following steps C1 - C3.

[0103] C1, natural language processing techniques.

[0104] Semantic similarity analysis: Use a pre-trained medical domain word vector model (such as BioBERT) to calculate the semantic similarity between the text description and the preset stroke keywords (such as 'ischemic penumbra', 'apoptosis regulation'), and filter out the entries with similarity lower than the threshold.

[0105] Named entity recognition (NER): Identify key entities such as gene names and brain region names through an NER model (such as based on the BiLSTM - CRF architecture), and mark the entries lacking key entities as incomplete data and eliminate them.

[0106] C2, machine learning algorithms.

[0107] Supervised learning classification: Using a random forest model, a classifier is trained based on an annotated dataset (containing'relevant' and 'noise' labels) to automatically identify descriptions unrelated to stroke (such as 'normal physiological processes', 'other disease associations').

[0108] Anomaly detection: Use the Isolation Forest algorithm to detect abnormal patterns (such as contradictory expressions, non-standard terms) in the description information.

[0109] C3. Manual review.

[0110] The screened data is reviewed by researchers, focusing on verifying semantically ambiguous entries and algorithm classification results to ensure the accuracy and consistency of the final dataset.

[0111] For example, through text analysis algorithms, data records with unclear or semantically ambiguous descriptions are detected and manually reviewed to determine whether they need to be removed. At the same time, machine learning algorithms are used to classify the data, and information unrelated to stroke is screened out and removed.

[0112] During the data cleaning process, if there are multiple literature sources for the same gene, the following steps D1 - D3 are performed.

[0113] D1. Extract the common description as the main description, and label the different descriptions as 'additional information'.

[0114] D2. Record the titles, DOIs, and page numbers of all source literature.

[0115] D3. Prioritize the sources according to the authority of the literature to ensure that sources with high credibility are displayed first.

[0116] The above text analysis algorithms include:

[0117] Named Entity Recognition (NER): Using a model based on the BiLSTM - CRF architecture, identify gene names (such as 'TNF - α'), brain region names (such as 'hippocampus', 'cortex'), and pathological regions (such as 'infarct core region') in the text. Entries missing key entities are regarded as incomplete data and excluded;

[0118] Semantic similarity calculation: Using a pre - trained medical domain word vector model (such as BioBERT), calculate the semantic similarity between the text description and preset stroke keywords (such as 'apoptosis regulation', 'ischemic penumbra'), and filter out entries with a similarity lower than a threshold (such as 0.7);

[0119] Rule matching: Match non - standard terms (such as 'cerebral infarction','stroke') through regular expressions and standardize them into a unified term (such as 'ischemic stroke').

[0120] The above-mentioned machine learning algorithms include:

[0121] Supervised learning classification: Using a random forest model, training a classifier based on an annotated dataset (including'relevant' and 'noise' labels), with features including term frequency (TF-IDF), entity type, and semantic similarity score, to automatically identify descriptions unrelated to stroke (such as 'normal metabolism', 'cancer association');

[0122] Anomaly detection: Using the Isolation Forest algorithm to detect abnormal patterns in the description information, such as contradictory statements (such as 'gene X is highly expressed and not expressed in the infarct core area') or non-standard terms (such as 'brain region A' not matching anatomical terms).

[0123] The above-mentioned manual review specifically includes: The screened data is independently reviewed by two researchers, focusing on verifying semantically ambiguous entries and the algorithm classification results, and resolving inconsistencies through cross-comparison to ensure the final data quality.

[0124] The above-mentioned matching process specifically includes the following steps 501 - step 503.

[0125] Step 501, data preparation.

[0126] First, obtain the standard gene set of the National Center for Biotechnology Information (NCBI). This standard gene set has been strictly screened and verified, and contains detailed information on various known genes.

[0127] At the same time, organize and prepare the cleaned original data to ensure that the data format and content meet the requirements of the matching. For example, standardize the gene names to accurately match the gene names in the standard gene set.

[0128] Step 502, matching calculation.

[0129] Use advanced string matching algorithms and bioinformatics tools to match the genes in the cleaned original data with the NCBI standard gene set.

[0130] For example, an algorithm based on edit distance can be used to calculate the similarity between the gene names in the original data and the gene names in the standard gene set. If the similarity exceeds a certain threshold, it is considered a successful match.

[0131] At the same time, match by combining the functional information and biological characteristics of the genes. For example, for some genes with unknown functions, their conservation, expression patterns, etc. in different species can be compared to match with the genes in the standard gene set.

[0132] Step 503, manual review.

[0133] To ensure the accuracy of data extraction by Tongyi Qianwen 2.1 model, a manual proofreading method is adopted. That is, two researchers independently check the same batch of data, and then compare their proofreading results to identify and resolve any inconsistencies. As Figure 3 shown, for the proofreading results of a randomly presented literature, it is found that the gene information, species information, physiological brain regions, and pathological partition data extracted by Tongyi Qianwen 2.1 model are highly consistent with those recorded in the original literature. By carefully checking hundreds of records extracted by the large language model, it is found that the model shows a high degree of accuracy in gene name recognition and reference matching. These results verify the effectiveness and reliability of Tongyi Qianwen 2.1 large language model in complex literature information extraction tasks, significantly improving the quality and credibility of the final database. Figure 3 In [A], it is a screenshot of the txt format document output by the results of Tongyi Qianwen 2.1 model; Figure 3 In [B], it is the manual proofreading result of the relevant content of the specific literature.

[0134] That is, after the matching is completed, manual review is carried out to ensure the accuracy of the matching. Manual review can detect some problems that cannot be recognized by the automatic matching algorithm, such as the ambiguity of gene names, the confusion of genes with similar functions but different ones, etc.

[0135] For the problems found in the manual review, further analysis and processing are carried out to ensure the accuracy and consistency of the data.

[0136] To reveal potential protein-protein interactions in the research, the STRING database is used to construct the corresponding protein-protein interaction network. The STRING database provides a platform that can integrate information from multiple sources to identify and analyze protein-protein interactions. These interaction information includes but is not limited to text mining, experimental verification, existing database records, protein co-expression patterns, protein neighborhood relationships, gene fusion events, and gene co-occurrence situations. Among them, the parameter is set with a minimum interaction score of 0.4 to ensure that the analyzed protein-protein interactions have a certain degree of credibility.

[0137] In another exemplary embodiment, in step 105 above, based on the cleaned structured data set, the STRING database is used to construct the corresponding protein-protein interaction network, including the following steps 601-step 605.

[0138] Step 601, data preparation.

[0139] Sort out the gene list to ensure that the gene names are accurate and standardized, and consistent with the recognition format of the STRING database.

[0140] Step 602, access the database and input the gene list.

[0141] Open the STRING database website and upload or input the gene list in the specified area.

[0142] Step 603, set parameters.

[0143] Set the minimum interaction score to 0.4 to ensure that reliable interaction relationships are included.

[0144] Step 604, construct the network.

[0145] Click the button to start the construction. The STRING database comprehensively analyzes and calculates various information sources and algorithms, including text mining, experimental verification, database record integration, protein co-expression pattern analysis, neighborhood relationship inference, gene fusion event analysis, and gene co-occurrence statistics, etc. It screens out high-confidence interaction relationships according to the set parameters to generate a network.

[0146] Step 605, result acquisition and analysis.

[0147] View the graphical display of the network, obtain data such as detailed interaction information, statistical features, and functional annotations, etc., and can be downloaded for further analysis, such as revealing disease-related biological pathways and regulatory mechanisms through the analysis of key nodes and network modules.

[0148] On the basis of constructing the knowledge graph, the gene list extracted from the literature is classified according to its specific region in the brain, and further pathway enrichment analysis is performed on the retrieved genes in different regions.

[0149] In another exemplary embodiment, in step 106 above, based on the cleaned structured data set, online bioinformatics tools such as Metascape, DAVID, and KEGG are used to perform functional annotation and pathway enrichment on genes, focusing on analyzing the distribution differences of genes in different physiological and pathological regions, and interpreting the importance of spatial heterogeneity for stroke diseases according to the pathway enrichment results. In the embodiments of the present application, because spatial heterogeneity needs to be concerned, the expressions of different genes in space are different and they also play different functions. Therefore, it is necessary to see whether the genes screened from the pdf reflect their positions in the text, and then collect the gene sets in different positions to see what their functions are, including the following steps 701-step 704.

[0150] Step 701, data preparation.

[0151] Prepare the cleaned structured data set (in the form of an Excel table).

[0152] Step 702, use Metascape to analyze the distribution differences of genes in different physiological regions and pathological regions.

[0153] Upload data: Open Metascape, drag and drop the data file or upload it as required, and select the appropriate format and gene identifier type.

[0154] Set and start: After setting the parameters as needed, click the analysis button.

[0155] Obtain results: Metascape integrates multiple databases and provides functional annotations (biological processes, cellular components, molecular functions, etc.) and pathway enrichment results (displayed in charts, tables, and visualized networks) through statistical analysis, and analyzes the distribution differences of genes in different physiological and pathological regions based on this.

[0156] Step 703, use DAVID to analyze the distribution differences of genes in different physiological and pathological regions.

[0157] Upload and convert: Access DAVID, and if necessary, use the "Gene ID Conversion" tool to convert the format and then upload the gene list, and select the species.

[0158] Start analysis: Click tools such as "Functional Annotation", and start the analysis after setting the parameters.

[0159] Interpret results: DAVID integrates multiple databases and gives functional annotations, pathway enrichment tables, and visualized pathway diagrams based on statistical tests, and pays attention to the distribution differences of genes in regions related to stroke based on this.

[0160] Step 704, use KEGG to analyze the distribution differences of genes in different physiological and pathological regions.

[0161] Prepare and enter: Prepare the data and enter the KEGG official website analysis entry.

[0162] Submit analysis: Enter or upload the gene list as required and select the correct analysis options.

[0163] View results: KEGG compares genes with pathways and presents the pathway enrichment results in a table through statistical calculation, and explores the impact of the distribution differences of genes in different regions on stroke based on this.

[0164] The process of pathway enrichment analysis for different regions is as follows.

[0165] 1. Pathway enrichment analysis for different damaged regions.

[0166] During the occurrence of ischemic stroke, the sudden reduction of cerebral blood flow can lead to hypoxia and insufficient energy supply in brain tissues. Under such circumstances, cells in the infarct core and ischemic penumbra exhibit different biological responses and survival strategies. The ischemic penumbra is the potentially recoverable area of brain tissues after the occurrence of ischemic stroke. As shown in Table 1, the high enrichment and activation of apoptosis-related pathways (Regulation of neuron apoptotic process Negative regulation of apoptoticsignaling pathway) indicate that the process of cell death in the ischemic penumbra region can be controlled, reflecting the mechanism by which cells strive to maintain survival. At the same time, these pathways slow down the programmed cell death, providing a time window for clinical intervention. Meanwhile, some studies have reported that when white blood cells are recruited to the inflammatory site, they may cause damage while producing a tissue-protective effect. However, Table 1 enriches the positive regulation pathway of interleukin-1 (Positiveregulation of interleukin-1 production), reflecting the tissue-protective effect, that is, the removal of dead cells and tissues through immune regulation, which can promote the recovery and repair of cells in the ischemic penumbra. In addition, the ischemic penumbra enriches pathways related to lipid metabolism and atherosclerosis. Some studies have reported that the metabolic transformation of various lipids is closely related to the expansion of the infarct area and neurological function damage in cerebral infarction, suggesting that abnormal lipid metabolism may play a role in cell survival, inflammatory response, and subsequent repair in ischemic stroke.

[0167] Programmed cell death, as well as inflammation and autoimmune responses, play a key role in the cell death mechanism. Through the analysis of the related pathways enriched in the infarct core region genes, the mainly enriched pathways are those related to the positive regulation of apoptotic process (Positiveregulation of apoptotic process), which is exactly the opposite of the apoptotic process in the ischemic penumbra. The activation of the positive regulation of apoptotic process indicates that cell death process will occur in the infarct core region after stroke, especially programmed cell death, indicating that there are cell damage and death due to hypoxia and insufficient energy supply in the infarct core region. In addition, the enrichment and activation of necroptosis and spinal cord injury pathways indicate that in the infarct core region, not only programmed cell death occurs, but also cell death involving inflammatory and autoimmune response mechanisms.

[0168] Table 1 List of results of enrichment analysis of related gene pathways in each injury region

[0169]

[0170] Analysis of the spatial functional differences between the infarct core and the ischemic penumbra is crucial for understanding the spatial heterogeneity and pathophysiological processes of stroke and provides important information for the design and selection of treatment strategies. For example, the protective and restorative potential of the ischemic penumbra suggests that early interventions (such as restoring blood flow, anti-inflammatory therapies) may be beneficial for minimizing brain tissue damage and promoting functional recovery. The study of the infarct core, on the other hand, helps to understand the irreversible process of brain tissue damage and provides strategies for reducing inflammation and preventing the spread of damage.

[0171] 2. Pathway enrichment analysis of different brain regions.

[0172] As shown in Table 2, gene enrichment pathway analysis in the cortical region found that the pathways of negative regulation of cell population proliferation and regulation of smooth muscle cell proliferation were highly enriched, indicating that the cortical region attempts to limit excessive cell proliferation and maintain tissue structure stability after ischemia to reduce pathological changes such as scar formation, thereby affecting the recovery process. The activation and enrichment of the biological pathways of regulation of neuron apoptotic process and negative regulation of apoptotic signaling pathway suggest that the cortical region may activate protective mechanisms after stroke to reduce the death of nerve cells through the apoptotic pathway. A study showed that the apoptosis of nerve cells in the cortical region increased significantly after ischemic stroke, which also confirmed that the regulation of these pathways may be crucial for reducing damage and promoting recovery.

[0173] In addition, the positive regulation of angiogenesis pathway is highly enriched in the gene enrichment pathways in the hippocampal region, indicating that the hippocampal region promotes the formation of new blood vessels after ischemia, thereby improving local blood flow changes, which is beneficial to the recovery and functional reconstruction of the damaged area. This is somewhat different from the regulation of blood vessels in the cortical region. The enrichment of the biological pathway of the positive regulation of cytokine-mediated signaling pathway reflects that there may be a more active inflammatory response and immune regulation process in the hippocampal region, which may be a double-edged sword for the repair process after injury. It is beneficial for clearing dead cells and tissue repair, but may also exacerbate the injury. At the same time, the enrichment and regulation of the cellular homeostasis pathway indicate that the hippocampal region may play an important role in maintaining the stability of the intracellular environment, which is particularly important for resisting the damage caused by ischemia.

[0174] Table 2 List of results of enrichment analysis of related gene pathways in each physiological brain region

[0175]

[0176] In another exemplary embodiment, after step 107, the stroke spatial knowledge graph database is also displayed. To effectively display the constructed knowledge graph, an application example of the present application developed a dedicated website. The front end of the server was developed using HBuilder X 3.7.11, and the back end was implemented using AppServ 9.3.0, Apache 2.4.4, PHP 7.3.1, MySQL 8.0.1, and phpMyAdmin 4.9.1. The original data matrix, preprocessed data, and intermediate results are stored in Notepad 8.4.8. The preprocessing is performed on Python 3.7, Anaconda Navigator 2.3.2, Jupyter Notebook 6.4.12, R 4.2.2, and RStudio. The interactive visualization graph is implemented using Apache ECharts 5.4. In the website display design, users can query specific genes through the search function, and the website will display the detailed information related to the gene, including its distribution in the brain, the biological processes involved, and the pathological states. In addition, the website also provides a graphical view of the knowledge graph, enabling users to intuitively understand the spatial heterogeneity of stroke and the interactions between genes.

[0177] The home page design of the StrokeDB website focuses on providing a friendly user interface. The main function bars include options such as retrieval, browsing, network, pathway, and statistics. In addition, the home page maximally showcases the spatial heterogeneity of stroke in visual display, which includes rich brain region information, injury area information, etc. Through this design, users can quickly obtain the basic information they need; through the integrated spatial expression data, users can also view the distribution of genes in brain regions through high-resolution images and gain an intuitive understanding of the complexity of stroke. The retrieval function of the database provides users with flexible and powerful data query capabilities. Users can retrieve specific gene information, gene descriptions, and the expression locations in the stroke space through keywords, or search for specific brain regions to return detailed brain region information and related gene expression data. In addition, the website also includes a browsing function for gene lists and brain region lists, enabling users to easily compare the functions and expression patterns of different genes and deeply understand the transcriptional expression profiles and biological pathways of specific brain regions. For the protein-protein interaction information related to stroke, the database provides a detailed display of the protein-protein interaction (PPI) network data. This part of the content not only includes the basic information of PPI, such as interacting protein pairs, interaction types, and functional annotations, but also provides relevant experimental support evidence and literature citations. Through these detailed PPI information, researchers can explore the molecular mechanisms of stroke more deeply and provide a scientific basis for further experimental design and disease treatment. Finally, the data statistics page of StrokeDB summarizes and displays comprehensive information on the number of genes, detailed information on brain tissue samples, marker genes, protein interactions, and related pathway enrichments in the database, highlighting the extensive application potential and data support capabilities of StrokeDB in the field of stroke research.

[0178] Based on the above various method embodiments, the present application uses the Tongyi Qianwen 2.1 model to conduct in-depth text mining on a large number of literatures, performs pathway enrichment analysis on genes related to spatial regions, successfully constructs the stroke spatial knowledge graph database StrokeDB, and presents it in the form of a website. Through comparing the differential genes and pathway enrichment analysis between the ischemic penumbra and the infarct core, it is found that the infarct core and the ischemic penumbra exhibit different biological characteristics in aspects such as cell death, inflammatory response, cell protection, and repair. When comparing the differences in different physiological brain regions in stroke diseases, it is found that the cortical region restricts the damage range and promotes tissue stability by regulating the processes of cell death and proliferation, while the hippocampal region promotes injury repair and functional recovery by promoting angiogenesis and regulating the inflammatory response. In summary, the method of the embodiments of the present application emphasizes the importance of spatial heterogeneity in the development of stroke diseases, challenges the general conclusions obtained by traditional Bulk research methods, and also lays a foundation for the comparative study of the effects of Panax notoginseng and Panax ginseng on acute stroke.

[0179] Based on the same inventive concept, the embodiments of the present application also provide a device for constructing a stroke spatial knowledge graph database for implementing the method for constructing the stroke spatial knowledge graph database involved above. The solution provided by this device for solving problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the device for constructing a stroke spatial knowledge graph database provided below can refer to the limitations on the method for constructing a stroke spatial knowledge graph database in the above text, and will not be repeated here.

[0180] In an exemplary embodiment, a device for constructing a stroke spatial knowledge graph database is provided, including the following modules.

[0181] A paper acquisition module, configured to acquire papers related to ischemic stroke.

[0182] A gene recognition module, configured to use the Tongyi Qianwen 2.1 model to recognize the genes related to ischemic stroke and the description information of the genes recorded in each paper, and output them in the form of structured data; the structured data includes gene names, description information, and sources.

[0183] A region annotation module, configured to use the Tongyi Qianwen 2.1 model to analyze the papers recording each structured data, perform physiological region annotation and pathological region annotation on each structured data, and add the physiological region annotation and pathological region annotation to the structured data respectively to construct a structured data set.

[0184] A cleaning module for cleaning a structured data set, deleting invalid and duplicate structured data in the structured data set, obtaining a cleaned structured data set, and outputting the cleaned structured data set in the form of an Excel spreadsheet to obtain Excel spreadsheet data.

[0185] A protein-protein interaction network construction module for constructing a protein-protein interaction network using the STRING database based on the Excel spreadsheet data.

[0186] A distribution difference analysis module for analyzing the distribution differences of genes in different physiological and pathological regions using an online bioinformatics tool based on the Excel spreadsheet data, and determining the functional annotations and pathway enrichments of each gene in the Excel spreadsheet data.

[0187] A database construction module for constructing a stroke spatial knowledge graph database based on the Excel spreadsheet data, the protein-protein interaction network, and the functional annotations and pathway enrichments of each gene in the Excel spreadsheet data.

[0188] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be as Figure 4 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. The computer program, when executed by the processor, implements a method for constructing a stroke spatial knowledge graph database.

[0189] Those skilled in the art can understand, Figure 4The structure shown is only a block diagram of some of the structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements. In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0190] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0191] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0192] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0193] In each of the embodiments provided in the present application, the databases involved may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on a blockchain, etc., without limitation. In each of the embodiments provided in the present application, the processor may be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without limitation.

[0194] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0195] Specific examples are used in this article to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation of the present application.

Claims

1. A method for constructing a spatio-temporal knowledge graph database for stroke, characterized in that, Including: Obtaining papers related to ischemic stroke; Using Tongyi Qianwen 2.1 model to identify the genes related to ischemic stroke and the description information of the genes recorded in each paper, and outputting them in the form of structured data; the structured data includes gene names, description information, and sources; Using Tongyi Qianwen 2.1 model to analyze the papers recording each structured data, performing physiological region annotation and pathological region annotation on each structured data, and adding the physiological region annotation and pathological region annotation to the structured data respectively to construct a structured data set; Cleaning the structured data set, deleting the invalid and duplicate structured data in the structured data set, obtaining the cleaned structured data set, and outputting the cleaned structured data set in the form of an Excel table to obtain Excel table data; According to the Excel table data, using the STRING database to construct a protein-protein interaction network; According to the Excel table data, using online bioinformatics tools to analyze the distribution differences of genes in different physiological regions and pathological regions, and determining the functional annotations and pathway enrichments of each gene in the Excel table data; According to the Excel table data, the protein-protein interaction network, and the functional annotations and pathway enrichments of each gene in the Excel table data, constructing a stroke spatial knowledge graph database; Using Tongyi Qianwen 2.1 model to identify the genes related to ischemic stroke and the description information of the genes recorded in each paper, specifically including: Inputting the instruction "Please extract all genes related to ischemic stroke and their description information from the paper" into the Tongyi Qianwen 2.1 model; Inputting each paper into the Tongyi Qianwen 2.1 model one by one, and obtaining the genes related to ischemic stroke and the description information of the genes recorded in each paper output by the Tongyi Qianwen 2.1 model; Inputting each paper into the Tongyi Qianwen 2.1 model one by one, and obtaining the genes related to ischemic stroke and the description information of the genes recorded in each paper output by the Tongyi Qianwen 2.1 model, specifically including: Using the natural language processing function of the Tongyi Qianwen 2.1 model to segment the text in the paper; Using the deep neural network gene recognition function of the Tongyi Qianwen 2.1 model to recognize each word obtained by segmentation, and obtaining the words that can represent the gene names related to ischemic stroke and their description information; the description information is obtained by combining the descriptive words of the gene names related to ischemic stroke; the deep neural network recognition function is implemented based on a pre-trained deep neural network sub-model, and the pre-trained deep neural network sub-model is obtained by training the deep neural network model using the knowledge in the medical and biological fields.

2. The construction method of the stroke spatial knowledge graph database according to claim 1, characterized in that Obtaining papers related to ischemic stroke, specifically including: Obtaining papers related to ischemic stroke published within a preset time period from the PubMed database; Setting the file name of each paper in the form of "Stroke + serial number".

3. The method for constructing a stroke spatial knowledge graph database according to claim 1, wherein Analyze the papers recording each structured data using Tongyi Qianwen 2.1 model, and perform physiological region annotation and pathological region annotation on each structured data, specifically including: Analyze the papers containing structured data i using the natural language processing function of Tongyi Qianwen 2.1 model to obtain the words related to the region where the gene in structured data i belongs; i = 1, 2,..., I, where I is the number of structured data; Use the deep neural network region recognition function of Tongyi Qianwen 2.1 model to recognize the words related to the region where the gene in structured data i belongs, determine the physiological region and pathological region to which the gene in structured data i belongs, and perform physiological region annotation and pathological region annotation on structured data i.

4. The method for constructing a stroke spatial knowledge graph database according to claim 1, wherein Clean the structured data set, delete the invalid and duplicate structured data in the structured data set, and obtain the cleaned structured data set, specifically including: Check the integrity of each structured data in the structured data set, delete the incomplete structured data, and obtain the structured data set after the first cleaning; Match the description information of each structured data in the structured data set after the first cleaning with the preset stroke keywords, and delete the structured data in the structured data set after the first cleaning whose description information does not match the preset stroke keywords, and obtain the structured data set after the second cleaning; Perform unified standardization processing on each structured data in the structured data set after the second cleaning to obtain the standardized structured data set; Compare each structured data in the standardized structured data set, delete the duplicate structured data in the standardized structured data set, and obtain the structured data set after the third cleaning; Match the gene names of each structured data in the structured data set after the third cleaning with the standard gene set, and delete the structured data in the structured data set after the third cleaning whose gene names do not match the standard gene set, and obtain the cleaned structured data set.

5. The method for constructing a stroke spatial knowledge graph database according to claim 1, wherein The online bioinformatics tool specifically includes one or more of Metascape, DAVID, and KEGG.

6. An apparatus for constructing a spatio-temporal knowledge graph database for stroke, characterized in that The device for constructing the stroke spatial knowledge graph database applies the method for constructing the stroke spatial knowledge graph database according to any one of claims 1-5. The device for constructing the stroke spatial knowledge graph database includes: A paper acquisition module for acquiring papers related to ischemic stroke; A gene recognition module for using Tongyi Qianwen 2.1 model to recognize the genes related to ischemic stroke and the description information of the genes recorded in each paper, and output them in the form of structured data; the structured data includes gene name, description information, and source; A region annotation module for analyzing the papers recording each structured data using Tongyi Qianwen 2.1 model, performing physiological region annotation and pathological region annotation on each structured data, and adding the physiological region annotation and pathological region annotation to the structured data respectively to construct a structured data set; A cleaning module, which is used to clean the structured data set, delete the invalid and duplicate structured data in the structured data set, obtain the cleaned structured data set, and output the cleaned structured data set in the form of an Excel table to obtain Excel table data; A protein-protein interaction network construction module, which is used to construct a protein-protein interaction network according to the Excel table data by using the STRING database; A distribution difference analysis module, which is used to analyze the distribution differences of genes in different physiological regions and pathological regions according to the Excel table data by using online bioinformatics tools, and determine the functional annotations and pathway enrichments of each gene in the Excel table data; A database construction module, which is used to construct a stroke spatial knowledge graph database according to the Excel table data, the protein-protein interaction network, and the functional annotations and pathway enrichments of each gene in the Excel table data; 7. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the method for constructing the stroke spatial knowledge graph database according to any one of claims 1-5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method for constructing the stroke spatial knowledge graph database according to any one of claims 1-5.

Citation Information

Patent Citations

  • Lung cancer medical big data-based treatment pathway key node information processing method

    CN110618987A

  • Anticancer drug collaborative prediction method based on knowledge graph attention network

    CN116313147A