An automatic construction method of knowledge graph based on evidence-based medicine
By preprocessing and entity recognition of multi-source medical data based on evidence-based medicine, and using natural heuristic optimization algorithms to optimize node layout, the problem of unreasonable node layout in the knowledge graph is solved, the visualization effect and query efficiency of medical knowledge graph are improved, its usability and user experience are enhanced, and the scientificity and accuracy of evidence-based medical decision-making are improved.
Patent Information
- Application Number
- CN202510511720.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-23
AI Technical Summary
The existing technology does not consider node layout optimization during the knowledge graph construction process, resulting in unreasonable node layout, affecting the visualization effect and query efficiency of the knowledge graph, increasing computing resource consumption, and reducing overall availability and user experience.
An evidence-based medicine-based method is adopted to collect multi-source medical data, perform preprocessing and entity recognition, and use natural heuristic optimization algorithms and deep learning technology to optimize node layout to build a medical knowledge graph.
It improves the visualization effect and query efficiency of the medical knowledge graph, enhances its usability and user experience, and improves the scientificity and accuracy of evidence-based medical decisions.
Smart Images

Figure CN120032916B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical information technology, and more specifically, to a method for automatically constructing a knowledge graph based on evidence-based medicine. Background Art
[0002] With the continuous deepening of medical research and the rapid growth of clinical data, a vast amount of knowledge and information has been generated in the medical field. However, this knowledge is scattered in different literatures, electronic medical records, databases, and clinical guidelines. How to effectively integrate, manage, and apply this information has become a major challenge in modern medical research and clinical practice. Traditional knowledge management methods are no longer sufficient to handle such a large amount of data and complex knowledge systems, resulting in low efficiency in information acquisition and application, and affecting the scientificity and accuracy of medical decisions.
[0003] Evidence-based medicine (EBM), as a modern medical model, helps medical staff make scientific medical decisions by systematically evaluating existing evidence. However, in practice, there are still many difficulties in the process of obtaining, integrating, and applying this evidence. Specific problems include complex evidence acquisition channels, cumbersome integration processes, and poor real-time performance in clinical practice. To improve the efficiency of evidence-based medicine, knowledge graph technology has gradually attracted wide attention. By modeling structured and unstructured information from different data sources into a unified graph structure, a knowledge graph can describe the complex network relationships between entities and relationships, which helps to explore and apply the deep knowledge hidden in medical data.
[0004] Chinese Patent with the authorization announcement number CN111370127B discloses a cross-departmental early diagnosis decision support system for chronic kidney disease based on a knowledge graph, including a patient information model establishment module, a patient information model library storage module, a knowledge graph association module, a knowledge graph reasoning module, and a decision support feedback module. By constructing a patient information model and using the OMOP CDM standard terminology system, the present invention constructs the patient's electronic medical record data into a patient information model with unified concept coding and semantic structure, giving full play to the advantages of semantic technology in data interactivity and scalability, so that the system has good adaptability and scalability to heterogeneous data from different hospitals. At the same time, the clinical suggestions derived from the knowledge reasoning of the knowledge graph all come from clinical guidelines and physician experience that conform to evidence-based medicine. The reasoning process and the reasons for the suggestions can be traced and obtained by constructing reasoning instances, so that the reasoning process and the reasons for the suggestions can be given while giving clinical suggestions, enhancing the physician's trust in the decision support suggestions.
[0005] However, the above technologies do not take into account node layout optimization during the knowledge graph construction process. The node layout in the knowledge graph (i.e., the position of each node in the graph) will affect the visualization effect and query efficiency of the knowledge graph. If the node layout is unreasonable, it will lead to chaotic distribution of data nodes and edges, increasing the path complexity during query. In addition, more computing resources will be consumed when retrieving and processing requests, resulting in longer response time. This will lead to low query efficiency and chaotic visualization, affecting the relevance and logic between data, and thus reducing the overall usability and user experience of the knowledge graph.
[0006] In view of this, the present invention proposes an automatic construction method of a knowledge graph based on evidence-based medicine to solve the above problems. Summary of the invention
[0007] In order to overcome the above-mentioned defects of the prior art and to achieve the above-mentioned purpose, the present invention provides the following technical solution: a method for automatically constructing a knowledge graph based on evidence-based medicine, comprising:
[0008] S1: Collect medical data from multiple sources;
[0009] S2: Preprocess the multi-source medical data and mark them as multi-source processed data;
[0010] S3: Perform entity recognition on multi-source processed data to identify medical entities;
[0011] S4: Perform relationship extraction on multi-source processed data to extract the relationships between medical entities;
[0012] S5: Node layout is performed based on medical entities and the relationships between medical entities, and optimized using a nature-inspired optimization algorithm;
[0013] S6: Construct a medical knowledge graph based on the optimized node layout.
[0014] Further, different node layout algorithms are used to perform node layout based on medical entities and the relationships between medical entities, and N node sets are obtained, where N is an integer greater than 1, and each node set includes the coordinates of each node and the edges connecting the nodes. The nodes are medical entities, and the edges are the relationships between medical entities. The edges connecting the nodes are represented by tuples, and the tuples include two nodes connected to the edge.
[0015] The steps to optimize the node layout include:
[0016] Step S501: setting different digital labels for different node sets and marking them as set labels;
[0017] Step S502: Preset the population size Z and the number threshold T;
[0018] Step S503: Initialize the population. The positions of the lizards in the initialized population are defined in a one-dimensional search space. The positions of the lizards correspond one-to-one with the set labels. The iteration count t of the initialized population is 0;
[0019] Step S504: Determine the fitness function;
[0020] Step S505: Define the temperature and food intake of each lizard's position;
[0021] Step S506: Define the cave position and determine whether the lizard enters the cave or enters the foraging stage;
[0022] Step S507: Update the positions of the lizards;
[0023] Step S508: Determine whether the iteration is completed. If not, return to Step S505 and set the iteration count ; if completed, enter Step S509;
[0024] Step S509: Calculate the fitness corresponding to the position of each lizard, obtain the position of the lizard corresponding to the maximum fitness value, and obtain the node set corresponding to the corresponding set label according to the obtained lizard position.
[0025] Furthermore, in the said Step S503, initializing the population , includes Z lizards; both the set label and the range of the one-dimensional search space are ; the expression of the position of each lizard is: ; in the formula, is the initial position of the i-th lizard, is a random number between, ;
[0026] In the said Step S504, the expression of the fitness function is: ; in the formula, is the fitness, and BZ is the layout quality; the method for obtaining the layout quality index includes:
[0027] Obtain the node set corresponding to the set label corresponding to the lizard position and mark it as the analysis set, calculate the calculated edge crossing number, average edge length, and distribution uniformity corresponding to the analysis set; use the analysis set, edge crossing number, average edge length, and distribution uniformity as analysis data, and input the analysis data into the trained quality prediction model to predict the corresponding layout quality.
[0028] Furthermore, in the said Step S505, the expression of the temperature of the lizard position is: ; in the formula, is the temperature of the i-th lizard position, , are all constants;
[0029] A lizard food intake model is pre-built, and the lizard food intake model is a normal distribution model;
[0030] The expression for food intake is: ; In the formula, is the food intake of the i-th lizard, is the normalization constant, exp is the exponential function, is the standard deviation of food intake, is the optimal temperature, which is the temperature corresponding to the maximum food intake of the lizard;
[0031] In step S506, the method for determining whether the lizard has entered a cave or a foraging stage includes:
[0032] like , then the lizard enters the cave to escape the heat, and R is a constant;
[0033] like , then the lizards do not enter the cave to escape the heat, and the lizards enter the foraging stage;
[0034] The expression for the cave is: ; In the formula, For the cave location, is the position of the lizard with the highest fitness during the population iteration process, It is the position of the lizard with the highest fitness in the last population iteration;
[0035] In step S508, the method for judging whether the iteration is completed is: if the number of iterations t is less than the number threshold T, the iteration is not completed; if the number of iterations t is greater than or equal to the number threshold T, the iteration is completed.
[0036] Furthermore, in step S507, the method for updating the position of the lizard entering the cave includes:
[0037] Each lizard that enters the cave is given a random number , for A random number between
[0038] like , then the calculation method for the position of the lizard entering the cave includes:
[0039] ;
[0040] ;
[0041] In the formula, is the position of the lizard that enters the cave, is a decreasing curve;
[0042] If , the calculation method for the position of the lizard entering the cave is as follows: ; where is the position of a randomly selected lizard in the population;
[0043] The method for updating the position of the lizard entering the foraging stage includes:
[0044] The expression for the food position is: ; where is the food position;
[0045] The expression for the food size is: ; where is the food size, is the food factor, = 3, is the fitness corresponding to the position of the i-th lizard, is the fitness corresponding to the food position, is a random number between;
[0046] If , the expression for the food position after shredding the food is: ; where is the food position after shredding the food; The calculation method for the position of the lizard entering the foraging stage includes:
[0047] ;
[0048] where is the position of the i-th lizard after entering the foraging stage, is a random number between;
[0049] If , the lizard moves directly towards the food and eats, and the calculation method for the position of the lizard entering the foraging stage is: .
[0050] Furthermore, the calculation method for the edge crossing number includes:
[0051] Mark the two node coordinates corresponding to each edge in the analysis set as the first coordinate and the second coordinate respectively. Subtract the second coordinate corresponding to each edge from the first coordinate to obtain the corresponding direction vector for each edge, and mark it as the edge vector; Randomly combine all the edge vectors, take any two edge vectors as an edge set, and all the edge sets are different;
[0052] Label the two edge vectors in each edge set as the first vector and the second vector respectively; subtract the first coordinate of the corresponding first vector from the first coordinate of the second vector in each edge set, and then multiply by the first vector to obtain the first cross product of each edge set; subtract the first coordinate of the corresponding first vector from the second coordinate of the second vector in each edge set, and then multiply by the first vector to obtain the second cross product of each edge set; subtract the first coordinate of the corresponding second vector from the first coordinate of the first vector in each edge set, and then multiply by the second vector to obtain the third cross product of each edge set; subtract the first coordinate of the corresponding second vector from the second coordinate of the first vector in each edge set, and then multiply by the second vector to obtain the fourth cross product of each edge set; where the first coordinate of the edge vector is the first coordinate used when calculating the edge vector, and the second coordinate of the edge vector is the second coordinate used when calculating the edge vector;
[0053] Label the edge sets where the signs of the first cross product and the second cross product are different, and the signs of the third cross product and the fourth cross product are different as the cross sets; count the number of cross sets as the edge crossing number.
[0054] Furthermore, the calculation method of the average side length is as follows: subtract the abscissa of the second coordinate from the abscissa of the first coordinate corresponding to each edge to obtain the horizontal difference corresponding to each edge; subtract the ordinate of the second coordinate from the ordinate of the first coordinate corresponding to each edge to obtain the vertical difference corresponding to each edge; take the square root of the sum of the square of the horizontal difference and the square of the vertical difference to obtain the length of each edge; count the number of edges in the analysis set, add the lengths of each edge in turn and divide by the number of edges to obtain the average side length;
[0055] The calculation method of the distribution uniformity is as follows: count the number of node coordinates in the analysis set and label it as the number of nodes; add the abscissas of each node coordinate in turn and divide by the number of nodes to obtain the horizontal mean value; add the ordinates of each node coordinate in turn and divide by the number of nodes to obtain the vertical mean value; subtract the horizontal mean value from the abscissa of each node coordinate in turn to obtain the horizontal mean difference value; subtract the vertical mean value from the ordinate of each node coordinate in turn to obtain the vertical mean difference value; add the squares of each horizontal mean difference value in turn and divide by the number of nodes to obtain the horizontal variance; add the squares of each vertical mean difference value in turn and divide by the number of nodes to obtain the vertical variance; take the square root of the sum of the horizontal variance and the vertical variance to obtain the distribution uniformity;
[0056] The training process of the quality prediction model includes:
[0057] Pre-collect the analysis data of group b, set the corresponding layout quality for the analysis data of group b, where b is an integer greater than 1, and convert the analysis data and the corresponding layout quality into a corresponding set of feature vectors; use each set of feature vectors as the input of the quality prediction model. The quality prediction model outputs a set of predicted layout qualities corresponding to each set of analysis data, uses the actual layout quality corresponding to each set of analysis data as the prediction target, and the actual layout quality is the pre-set layout quality corresponding to the analysis data; use minimizing the sum of prediction errors of all analysis data as the training target; train the quality prediction model until the sum of prediction errors converges and then stop training; the quality prediction model is a deep neural network model.
[0058] Further, the multi-source medical data is medical-related data from different sources; the steps for preprocessing the multi-source medical data include:
[0059] Step S201: Clean the multi-source medical data;
[0060] Step S202: Normalize the multi-source medical data;
[0061] Step S203: Transform the multi-source medical data;
[0062] Step S204: Integrate the multi-source medical data;
[0063] The method for entity recognition of the multi-source processed data includes:
[0064] Input the multi-source processed data into the trained entity extraction model to identify medical entities; the entity extraction model is a BERT model, and the principle of the BERT model is:
[0065] For an input text, decompose it into words, and then convert each word into its corresponding embedding vector; the embedding vector includes an embedded word and an embedded sentence position;
[0066] The embedded word Et is ; where is the embedding matrix of the word, is the word vector of the original word;
[0067] The embedded sentence position Es is ; where is the embedding matrix of the position, is the position encoding of the original word in the sentence;
[0068] The final input representation is ;
[0069] At the core of BERT is a Transformer encoder composed of multiple attention heads; the input representation Einput is fed into the Transformer encoder, and through multiple layers of self-attention mechanisms and feed-forward networks, the context output representation is obtained;
[0070] Preset entity labels. After inputting the multi-source processed data into the entity extraction model, obtain the labels of each word, and screen out the words corresponding to the entity labels from all the labels, and mark them as medical entities.
[0071] Furthermore, the method for extracting relationships from the multi-source processed data includes:
[0072] The methods for relationship extraction include rule reasoning, graph algorithm reasoning, and model reasoning;
[0073] Rule reasoning is to use pre-defined rules to extract the relationships between medical entities through the rules and discover new relationships through logical reasoning;
[0074] Graph algorithm reasoning is to use graph algorithms to extract the relationships between medical entities;
[0075] Model reasoning is to use a graph neural network model to extract the relationships between medical entities;
[0076] The principle of the graph neural network model includes:
[0077] Graph structure representation: For a graph , where V is the set of nodes and A is the set of edges. For each node there exists a feature vector representing the feature information of the node;
[0078] Message passing: For node , it receives the messages from adjacent nodes and updates the feature vector corresponding to node ;
[0079] The process of message passing includes two parts: and , where is the feature vector of node in the th convolutional layer of the graph neural network, is the set of adjacent nodes of node , is the feature vector of the edge connecting node and , including the weight and type information of the edge; Agg is an aggregation function that aggregates the information of the adjacent nodes of node , including summation and averaging, is an activation function, For the connection node and the eigenvector of the edge in the L-th convolutional layer of the graph neural network;
[0080] The eigenvector of the node is updated in each convolutional layer of the graph neural network, and through message passing in multiple convolutional layers, an eigenvector containing information and context connections corresponding to multiple nodes is obtained;
[0081] Relationship reasoning: Relationship reasoning is performed by learning the eigenvector of the node and the eigenvector of the edge corresponding to the connected node.
[0082] Furthermore, the steps of constructing the medical knowledge graph include:
[0083] Step S601: According to the optimized node layout, each medical entity is used as a node, and an edge is created between the nodes according to the relationship between the medical entities;
[0084] Step S602: Select a graph database to store the created nodes and edges, and import the created nodes and edges into the graph database in the supported format of the graph database;
[0085] Step S603: Use a visualization tool to display the nodes and edges in the graph database to complete the construction of the medical knowledge graph.
[0086] The technical effects and advantages of a method for automatically constructing a knowledge graph based on evidence-based medicine according to the present invention:
[0087] By collecting and preprocessing multi-source medical data, using various methods for entity recognition and relationship extraction, and adopting a nature-inspired optimization algorithm and deep learning technology to optimize the node layout, a well-structured and high-query-efficiency medical knowledge graph is automatically constructed from massive medical data; it can effectively improve the integration and reasoning ability of medical knowledge, and at the same time improve the visualization effect and usage efficiency of the medical knowledge graph by optimizing the node layout, enhance the usability and user experience of the medical knowledge graph, and thus improve the scientificity and accuracy of evidence-based medical decision-making. Brief Description of the Drawings
[0088] Figure 1 It is a flowchart of a method for automatically constructing a knowledge graph based on evidence-based medicine according to Embodiment 1 of the present invention;
[0089] Figure 2 It is a flowchart of a method for optimizing the node layout according to Embodiment 1 of the present invention;
[0090] Figure 3 It is a schematic diagram of an electronic device according to Embodiment 2 of the present invention;
[0091] Figure 4Schematic diagram of the storage medium according to Embodiment 3 of the present invention. Detailed implementation manners
[0092] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0093] Embodiment 1
[0094] Please refer to Figure 1 As shown, a method for automatically constructing a knowledge graph based on evidence-based medicine in this embodiment includes:
[0095] S1: Collect multi-source medical data.
[0096] The multi-source medical data is medical-related data from different sources; the multi-source medical data is collected through multiple data sources such as medical literature, clinical trials, case reports, and electronic medical records; the multi-source medical data includes, for example, electronic medical record (EMR) data (such as patient information, medical visit records, laboratory reports, prescription records, etc.), medical imaging data (such as CT, MRI, X-ray, etc.), physiological signal monitoring data (such as electrocardiogram, electroencephalogram, etc.), social media and Internet data (such as patient discussion information from online forums and social platforms), clinical trial data (such as data recorded in various clinical trials), etc.
[0097] It should be noted that the purpose of collecting multi-source medical data is to provide a rich source of medical knowledge: the knowledge graph needs to mine and extract knowledge from various medical data sources, and multi-source data can provide broader and more comprehensive data support; support knowledge superposition and comparison: different aspects of medical knowledge can be discovered from data sources with different perspectives and types, and the knowledge graph can be enriched through complementary comparison; help with knowledge representation and extraction: different data sources express knowledge from different perspectives and in different formats, and can be used as references for each other to standardize knowledge representation and discover extraction rules; verify and update the knowledge graph: cross-verify the results of knowledge graph construction through multi-source data, discover new knowledge, and update and iterate the graph; support research on complex problems: complex medical problems such as etiology, diagnosis, and treatment need to be studied from multiple angles and levels, and multi-source data becomes an important information support.
[0098] S2: Preprocess the multi-source medical data and label it as multi-source processed data.
[0099] The steps for preprocessing the multi-source medical data include:
[0100] Step S201: Clean the multi-source medical data; remove the incorrect, redundant, and incomplete data in the multi-source medical data to improve the accuracy and consistency of the data.
[0101] Step S202: Normalize the multi-source medical data; unify the representation forms of the multi-source medical data to ensure the data consistency and compatibility of different data sources.
[0102] Step S203: Convert the multi-source medical data; ensure that different types of data (such as image data, text data) are represented in a consistent standard format, which helps to improve the data consistency and operability.
[0103] Step S204: Integrate the multi-source medical data; merge the data from different sources into a unified and comprehensive dataset for subsequent knowledge graph construction.
[0104] The above preprocessing methods are all existing technologies, and the specific processes are not elaborated here.
[0105] S3: Perform entity recognition on the multi-source processed data to identify medical entities.
[0106] The methods for performing entity recognition on the multi-source processed data include:
[0107] Input the multi-source processed data into a trained entity extraction model to identify medical entities; medical entities include, for example, diseases (such as diabetes, hypertension, etc.), symptoms (such as headache, cough, etc.), drugs (such as amoxicillin, ibuprofen, etc.), medical devices (such as CT scanners, electrocardiographs, etc.), etc.
[0108] The entity extraction model is specifically the BERT model, and the principle of the BERT model is:
[0109] For an input text, first decompose it into words, and then convert each word into its corresponding embedding vector; the embedding vector includes an embedded word and an embedded sentence position.
[0110] The embedded word Et is ; where is the embedding matrix of the word, is the word vector of the original word;
[0111] The embedded sentence position Es is ; where is the embedding matrix of the position, is the position encoding of the original word in the sentence;
[0112] The final input representation is ;
[0113] At the core of BERT is a Transformer encoder composed of multiple attention heads; the input representation Einput is fed into the Transformer encoder, and through multiple layers of self-attention mechanisms and feed-forward networks, a context-rich output representation is obtained;
[0114] To adapt to specific tasks, a task-specific layer, such as a fully connected layer, is added to the output of BERT; for the entity extraction task, a fully connected layer Wner can be used to predict whether each word is a named entity: and according to the specific requirements of the task, rules are defined to label and extract medical entities. For example, dates usually have a specific format and can be matched through regular expressions;
[0115] Preset entity labels. After inputting multi-source processed data into the entity extraction model, obtain the labels of each word, and screen out the words corresponding to the entity labels from all labels, and mark them as medical entities.
[0116] The BERT model is a prior art, so the training process of the entity extraction model will not be elaborated here; the entity labels are preset by those skilled in the art.
[0117] S4: Perform relation extraction on the multi-source processed data to extract the relationships between medical entities.
[0118] The methods for performing relation extraction on the multi-source processed data include:
[0119] The methods for relation extraction include rule reasoning, graph algorithm reasoning, and model reasoning; the relationships between medical entities are, for example, the relationship between a drug or treatment method and a disease (such as amoxicillin treating bacterial infections), the causal relationship between a virus and a disease (such as the influenza virus causing influenza), the relationship describing symptoms and diseases (such as fever being related to infection), etc.
[0120] Rule reasoning is to use pre-defined rules to extract the relationships between medical entities through the rules, and discover new relationships through logical reasoning. For example, if one medical entity is a doctor and another medical entity is a patient, and there is a diagnosis record in the multi-source medical data, then there is a diagnosis relationship between these two medical entities; for another example, if disease A and disease B use the same drug, and disease B and disease C also use the same drug, then there is also a relationship between disease A and disease C.
[0121] Graph algorithm reasoning is to use graph algorithms to extract the relationships between medical entities, such as graph traversal algorithms, shortest path algorithms, etc. For the shortest path algorithm, for example, use depth-first search or breadth-first search to obtain the shortest path between two medical entities. The shortest path can reveal the shortest communication or shortest connection method between medical entities, thereby indicating the degree of closeness of the relationship between medical entities.
[0122] Model inference uses a graph neural network model to extract the relationships between medical entities; examples of graph neural network models include graph convolutional network (GCN), graph attention network (GAT), and graph isomorphism network (GIN), etc.
[0123] The principles of the graph neural network model include:
[0124] Graph structure representation:
[0125] For a graph , where V is the set of nodes and A is the set of edges. For each node , there exists a feature vector representing the feature information of the node.
[0126] Message passing:
[0127] The core of the graph neural network model is to enable nodes to aggregate information from adjacent nodes through a message passing mechanism. For node , it receives messages from adjacent nodes and updates the feature vector corresponding to node ;
[0128] The process of message passing includes two parts: and , where is the feature vector of node in the -th convolutional layer of the graph neural network, is the set of adjacent nodes of node , is the feature vector of the edge connecting node and , including the weight and type information of the edge; Agg is an aggregation function that aggregates the information of the adjacent nodes of node , including summation and averaging. The specific choice depends on the task and network design, is the activation function, is the feature vector of the edge connecting node and in the
[0129] -th convolutional layer of the graph neural network;
[0130] The feature vector of the node is updated in each convolutional layer of the graph neural network. Through message passing in multiple convolutional layers, a feature vector containing the information and context connections corresponding to multiple nodes is obtained.
[0131] By learning the feature vectors of the nodes and the feature vectors of the edges corresponding to the connected nodes, relationship reasoning is performed, such as performing classification, regression, and other tasks.
[0132] S5: Perform node layout based on medical entities and the relationships between medical entities, and optimize it using a nature-inspired optimization algorithm.
[0133] Use different node layout algorithms to perform node layout based on medical entities and the relationships between medical entities, and obtain N sets of nodes, where N is an integer greater than 1. Each set of nodes includes the coordinates of each node and the edges connecting the nodes. The nodes are medical entities, and the edges are the relationships between medical entities. The edges connecting the nodes are represented by tuples, and the tuples include the two nodes connected by the edge. For example, (node 1, node 2); Node layout algorithms such as hierarchical layout, circular layout, principal component analysis layout, force-directed layout, etc.
[0134] As Figure 2 shown, the steps to optimize the node layout include:
[0135] Step S501: Set different digital labels for different sets of nodes and mark them as set labels;
[0136] Step S502: Preset the population size Z and the number threshold T;
[0137] Step S503: Initialize the population. The positions of the lizards in the initialized population are defined in a one-dimensional search space. The positions of the lizards correspond one-to-one with the set labels, and the iteration number t of the initialized population is 0;
[0138] Step S504: Determine the fitness function;
[0139] Step S505: Define the temperature and food intake of each lizard's position;
[0140] Step S506: Define the cave position and determine whether the lizard enters the cave or enters the foraging stage;
[0141] Step S507: Update the positions of the lizards;
[0142] Step S508: Determine whether the iteration is completed. If not, return to Step S505 and let the iteration number ; If completed, enter Step S509;
[0143] Step S509: Calculate the fitness corresponding to the position of each lizard, obtain the position of the lizard corresponding to the maximum fitness value, and obtain the set of nodes corresponding to the set label according to the obtained lizard position.
[0144] In the above step S502, the population size Z is determined by those skilled in the art. During the historical node layout optimization process, under multiple different node set conditions, multiple different population sizes are set for the same node set, and the lizard optimization algorithm is performed multiple times. After the same number of iterations, the corresponding set labels are obtained. The population size with the set label being the same as the actual set label is used as the population size corresponding to this set of test data; the actual set label is the digital label of the node set that best matches this node set, and the actual set label is obtained through experiments by those skilled in the art based on actual experience; and so on to obtain the population size corresponding to each node set, and the average value of multiple population sizes is used as the preset population size Z.
[0145] The iteration number threshold T is determined by those skilled in the art. During the historical node layout optimization process, under multiple different node set conditions, the lizard optimization algorithm is performed multiple times on the same node set to obtain multiple set labels, where the number of iterations of each lizard optimization algorithm is different, and the population size is the same and all are Z; the iteration number corresponding to the set label that is closest to the actual set label is used as the iteration number corresponding to this node set; and so on to obtain the iteration number corresponding to each node set, and the average value of multiple iteration numbers is used as the iteration number threshold T.
[0146] It should be understood that the population size Z determines the breadth of the search. A larger population size can explore more solution spaces, while the iteration number threshold T determines the termination condition of the algorithm and can control the convergence speed of the algorithm.
[0147] In the above step S503, initialize the population , including Z lizards.
[0148] The range of both the set label and the one-dimensional search space is ; The expression for the position of each lizard is: ; In the formula, is the initial position of the i-th lizard, is a random number between, .
[0149] In the above step S504, the expression of the fitness function is: ; In the formula, is the fitness, BZ is the layout quality; the method for obtaining the layout quality index includes:
[0150] Obtain the node set corresponding to the set label of the lizard position and mark it as the analysis set, calculate the computational edge crossing number, average edge length, and distribution uniformity corresponding to the analysis set; use the analysis set, edge crossing number, average edge length, and distribution uniformity as analysis data, and input the analysis data into the trained quality prediction model to predict the corresponding layout quality.
[0151] The training process of the quality prediction model includes:
[0152] Pre-collect b sets of analysis data, set corresponding layout qualities for each of the b sets of analysis data, where b is an integer greater than 1, and convert the analysis data and the corresponding layout qualities into a corresponding set of feature vectors; the layout quality corresponding to the analysis data is collected by those skilled in the art during the layout optimization process at historical nodes. For each of the b sets of analysis data, under the conditions of each analysis set, the layout quality is evaluated according to the actual situation; corresponding layout qualities are set for the b sets of analysis data in sequence;
[0153] Take each set of feature vectors as the input of the quality prediction model. The quality prediction model outputs a set of predicted layout qualities corresponding to each set of analysis data, uses the actual layout quality corresponding to each set of analysis data as the prediction target, and the actual layout quality is the pre-set layout quality corresponding to the analysis data; use minimizing the sum of the prediction errors of all analysis data as the training target; where the calculation formula for the prediction error is , where is the prediction error, k is the group number of the feature vector corresponding to the analysis data, is the predicted layout quality corresponding to the k-th set of analysis data, is the actual layout quality corresponding to the k-th set of analysis data; train the quality prediction model until the sum of the prediction errors reaches convergence and then stop training.
[0154] The above quality prediction model is specifically a deep neural network model; it includes an input layer, a hidden layer, and an output layer; each hidden layer includes multiple neurons, and there are connections between each neuron and the neurons in the next layer. The connections contain weights, which determine the importance and influence of data transmission in the neural network; an activation function is applied to each neuron between the hidden layer and the output layer. The activation function introduces non-linearity and allows the network to learn more complex patterns and features.
[0155] The calculation method of the edge crossing number includes:
[0156] Mark the two node coordinates corresponding to each edge in the analysis set as the first coordinate and the second coordinate respectively. Subtract the second coordinate corresponding to each edge from the first coordinate to obtain the direction vector corresponding to each edge, and mark it as the edge vector; randomly combine all the edge vectors, take any two edge vectors as an edge set, and all the edge sets are different;
[0157] Label the two edge vectors in each edge set as the first vector and the second vector respectively; subtract the first coordinate of the corresponding first vector from the first coordinate of the second vector in each edge set, and then multiply by the first vector to obtain the first cross product of each edge set; subtract the first coordinate of the corresponding first vector from the second coordinate of the second vector in each edge set, and then multiply by the first vector to obtain the second cross product of each edge set; subtract the first coordinate of the corresponding second vector from the first coordinate of the first vector in each edge set, and then multiply by the second vector to obtain the third cross product of each edge set; subtract the first coordinate of the corresponding second vector from the second coordinate of the first vector in each edge set, and then multiply by the second vector to obtain the fourth cross product of each edge set; where the first coordinate of the edge vector is the first coordinate used when calculating the edge vector, and the second coordinate of the edge vector is the second coordinate used when calculating the edge vector;
[0158] Label the edge sets where the signs of the first cross product and the second cross product are different, and the signs of the third cross product and the fourth cross product are different as the cross sets; count the number of cross sets as the edge crossing number; Exemplarily, for an edge set, the first cross product is 2, the second cross product is -1, the third cross product is 3, and the fourth cross product is -4, so this edge set is labeled as a cross set.
[0159] The calculation method of the average side length includes:
[0160] Subtract the abscissa of the second coordinate from the abscissa of the first coordinate corresponding to each edge to obtain the horizontal difference corresponding to each edge; subtract the ordinate of the second coordinate from the ordinate of the first coordinate corresponding to each edge to obtain the vertical difference corresponding to each edge; take the square root of the sum of the square of the horizontal difference and the square of the vertical difference to obtain the length of each edge; count the number of edges in the analysis set, add up the lengths of each edge in turn and divide by the number of edges to obtain the average side length.
[0161] The calculation method of the distribution uniformity includes:
[0162] Count the number of node coordinates in the analysis set and label it as the number of nodes; add up the abscissas of each node coordinate in turn and divide by the number of nodes to obtain the horizontal mean value; add up the ordinates of each node coordinate in turn and divide by the number of nodes to obtain the vertical mean value; subtract the horizontal mean value from the abscissa of each node coordinate in turn to obtain the horizontal mean difference value; subtract the vertical mean value from the ordinate of each node coordinate in turn to obtain the vertical mean difference value; add up the squares of each horizontal mean difference value in turn and divide by the number of nodes to obtain the horizontal variance; add up the squares of each vertical mean difference value in turn and divide by the number of nodes to obtain the vertical variance; take the square root of the sum of the horizontal variance and the vertical variance to obtain the distribution uniformity.
[0163] It should be noted that the number of edge crossings, average edge length, and distribution uniformity are all factors affecting the layout quality. The reason is that the number of edge crossings reflects whether there are crossings in the connections between nodes in the knowledge graph; a lower number of edge crossings usually means a clearer node layout, more intuitive information transmission, thus improving readability and comprehensibility, so the layout quality is better; if there are too many edge crossings, it may lead to information chaos and reduce the usability of the knowledge graph, so the layout quality is poor; the average edge length represents the average length of the edges connecting nodes; a shorter edge length means closer connections between nodes, which is conducive to rapid information transmission and enhances the overall visibility and interactivity of the knowledge graph, so the layout quality is better; relatively longer edges indicate that the layout is not compact enough, affecting the information transmission efficiency, so the layout quality is poor; the distribution uniformity measures the distribution state of nodes in the graph; a uniform distribution can ensure information balance and visual harmony, making it easier for users to capture important information and relationships, so the layout quality is better; if the node distribution is uneven, it may cause some areas to be too dense and other areas to be sparse, imposing a cognitive burden on users, so the layout quality is poor; in summary, these three indicators jointly affect the readability, query efficiency, and visualization effect of the knowledge graph, so they are important parameters for evaluating the layout quality of the knowledge graph.
[0164] In the above step S505, the expression for the temperature at the lizard's position is: ; where is the temperature at the position of the i-th lizard, and are both constants, , .
[0165] A lizard food intake model is pre-constructed. The lizard food intake model is a normal distribution model used to predict the food intake of lizards at different temperatures; the lizard food intake model is pre-constructed by those skilled in the art by pre-collecting the food intake data of lizards at different temperatures.
[0166] The expression for the food intake is: ; where is the food intake of the i-th lizard, is the normalization constant used to adjust the lizard food intake model to ensure that the sum of the probabilities of all values is 1, exp is the exponential function, is the standard deviation of the food intake, is the optimal temperature, which is the temperature corresponding to the maximum food intake of the lizard, and the standard deviation of the food intake and the optimal temperature are both obtained according to the lizard food intake model.
[0167] In the above step S506, the methods for determining whether the lizard enters the cave or the foraging stage include:
[0168] If , the lizard enters the cave to avoid the heat, indicating that the temperature at the location where the lizard is located is too high. R is a constant, ;
[0169] If , the lizard does not enter the cave to avoid the heat, indicating that the temperature at the location where the lizard is located is suitable for the lizard to feed, and the lizard enters the foraging stage.
[0170] The expression of the cave is: ; In the formula, is the cave location, is the location of the lizard with the maximum fitness during the population iteration process, is the location of the lizard with the maximum fitness in the previous population iteration; where and The difference is that is the location of the lizard with the maximum fitness during multiple population iterations, while is only the location of the lizard with the maximum fitness in the previous population iteration.
[0171] In the above step S507, the method for updating the location of the lizard entering the cave includes:
[0172] Assign a random number to each lizard entering the cave, is a random number between;
[0173] If , it means that the cave entered by the corresponding lizard has no competition from other lizards. Then the calculation method for the location of the lizard entering the cave includes:
[0174] ;
[0175] ;
[0176] In the formula, is the location of the i-th lizard entering the cave, is a decreasing curve.
[0177] If , it means that the cave entered by the corresponding lizard has competition from other lizards. Then the calculation method for the location of the lizard entering the cave includes:
[0178] ;
[0179] In the formula, is the location of a random lizard in the population.
[0180] The method for updating the location of the lizard entering the foraging stage includes:
[0181] When feeding, the lizard will choose whether to tear the food according to the size of the food; if the size of the food is appropriate, the lizard will directly ingest the food; if the food is too large, the lizard will use its sharp teeth and powerful jaw muscles to tear the food before ingesting it.
[0182] The expression for the food position is: ; where is the food position.
[0183] The expression for the food size is: ; where is the food size, is the food factor, = 3, is the fitness corresponding to the position of the i-th lizard, is the fitness corresponding to the food position, is a random number between.
[0184] If , it means the food is too large and the lizard needs to tear the food; the expression for the food position after tearing the food is: ; where is the food position after tearing the food.
[0185] The calculation method for the position of the lizard entering the foraging stage includes:
[0186] ;
[0187] where is the position of the i-th lizard after entering the foraging stage, is a random number between.
[0188] If , the lizard will move directly towards the food and feed, and the calculation method for the position of the lizard entering the foraging stage includes:
[0189] .
[0190] In the above step S508, the method for judging whether the iteration is completed is: if the iteration number t is less than the number threshold T, the iteration is not completed, and let the iteration number be the value of the iteration number after adding one to it and then assigning it to the iteration number; if the iteration number t is greater than or equal to the number threshold T, the iteration is completed.
[0191] It should be noted that the reason for using the lizard optimization algorithm to optimize the node layout is as follows: The lizard algorithm simulates the behaviors of lizards such as foraging, avoiding heat, and competing in different environments, and can perform global search in a complex multi-dimensional search space; by simulating the foraging and cave behaviors of lizards, the algorithm can effectively avoid falling into local optima, thereby increasing the probability of finding the global optimal solution; and the lizard algorithm can adapt to different node layout optimization problems by adjusting the population size Z and the number threshold T; the optimization process can dynamically adjust the population size and the number of iterations according to historical data to ensure the applicability and efficiency of the algorithm; at the same time, by reasonably setting the fitness function and the cave update mechanism, the lizard optimization algorithm can accelerate convergence and shorten the optimization time, that is, it can quickly find the optimal node layout method and improve the efficiency of knowledge graph construction; in addition, the lizard optimization algorithm can dynamically adjust the behaviors of lizards (such as directly eating or tearing food) during the foraging stage to achieve refined processing of local search; this search strategy that takes into account both local and global aspects improves the performance of the algorithm in complex multi-modal optimization problems.
[0192] S6: Construct a medical knowledge graph according to the optimized node layout;
[0193] The steps of constructing a medical knowledge graph include:
[0194] Step S601: According to the optimized node layout, take each medical entity as a node, and create edges between the nodes according to the relationships between medical entities;
[0195] Step S602: Select a graph database (such as Neo4j, ArangoDB, etc.) to store the created nodes and edges, and import the created nodes and edges into the graph database in the supported format of the graph database (such as CSV, JSON, etc.);
[0196] Step S603: Use a visualization tool (such as Gephi, Cytoscape, etc.) to display the nodes and edges in the graph database to complete the construction of the medical knowledge graph.
[0197] In this embodiment, by collecting and preprocessing multi-source medical data, using various methods for entity recognition and relationship extraction, and adopting a nature-inspired optimization algorithm and deep learning technology to optimize the node layout, a well-structured and highly query-efficient medical knowledge graph is automatically constructed from a large amount of medical data; it can effectively improve the integration and reasoning ability of medical knowledge, and at the same time improve the visualization effect and usage efficiency of the medical knowledge graph by optimizing the node layout, enhance the usability and user experience of the medical knowledge graph, thereby improving the scientificity and accuracy of evidence-based medical decision-making.
[0198] Embodiment 2
[0199] Please refer to Figure 3As shown, the present application also provides an electronic device 500. The electronic device 500 may include one or more processors and one or more memories. Among them, computer-readable code is stored in the memory, and when the computer-readable code is run by one or more processors, it can execute a method for automatically constructing a knowledge graph based on evidence-based medicine as described above.
[0200] The method or system according to an embodiment of the present application can also be implemented by means of Figure 3 the architecture of the electronic device shown. As Figure 3 shown, the electronic device 500 may include a bus 501, one or more CPUs 502, a ROM 503, a RAM 504, a communication port 505 connected to a network, an input / output 506, a hard disk 507, etc. The storage device in the electronic device 500, such as the ROM 503 or the hard disk 507, can store a method for automatically constructing a knowledge graph based on evidence-based medicine provided by the present application. Further, the electronic device 500 may further include a user interface 508. Of course, Figure 3 the architecture shown is only exemplary. When implementing different devices, one or more components in the Figure 3 shown electronic device can be omitted according to actual needs.
[0201] Embodiment 3
[0202] Please refer to Figure 4 shown. An embodiment of the present application discloses a computer-readable storage medium 600. Computer-readable instructions are stored on the computer-readable storage medium 600. When the computer-readable instructions are run by a processor, a method for automatically constructing a knowledge graph based on evidence-based medicine according to an embodiment of the present application described with reference to the above drawings can be executed. The storage medium 600 includes but is not limited to, for example, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and cache memory, etc. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.
[0203] In addition, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the present application provides a non-transitory machine-readable storage medium storing machine-readable instructions that can be run by a processor to execute instructions corresponding to the method steps provided by the present application, such as: a method for automatically constructing a knowledge graph based on evidence-based medicine. When the computer program is executed by a central processing unit (CPU), the above functions defined in the method of the present application are executed.
[0204] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims described above.
[0205] Finally: The above description is only the preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An automatic construction method for a knowledge graph based on evidence-based medicine, characterized in that, Including: S1: Collect multi-source medical data; S2: Preprocess the multi-source medical data and label it as multi-source processed data; S3: Perform entity recognition on the multi-source processed data to identify medical entities; S4: Extract relationships between medical entities from the multi-source processed data; S5: Perform node layout based on medical entities and the relationships between medical entities, and optimize it using a nature-inspired optimization algorithm; Using different node layout algorithms, perform node layout based on medical entities and the relationships between medical entities to obtain N node sets, where N is an integer greater than 1; The steps for optimizing the node layout include: Step S101: Set different digital labels for different node sets and label them as set labels; Step S502: Preset the population size Z and the number threshold T; Step S503: Initialize the population. The positions of the lizards in the initialized population are defined in a one-dimensional search space. The positions of the lizards correspond one-to-one with the set labels, and the iteration number t of the initialized population is 0; Step S504: Determine the fitness function; Step S505: Define the temperature and food intake of each lizard's position; Step S506: Define the cave position and determine whether the lizard enters the cave or the foraging stage; Step S507: Update the positions of the lizards; Step S508: Determine whether the iteration is completed. If not, return to Step S505 and set the iteration count ; if completed, proceed to Step S509; Step S509: Calculate the fitness corresponding to each lizard's position, obtain the position of the lizard corresponding to the maximum fitness value, and obtain the node set corresponding to the set label according to the obtained lizard position; In the said step S504, the expression of the fitness function is as follows: ; in the formula, is the fitness, and BZ is the layout quality; the method for obtaining the layout quality index includes: Obtain the node set corresponding to the set label of the lizard position and label it as the analysis set. Calculate the number of edge crossings, average edge length, and distribution uniformity corresponding to the analysis set; Use the analysis set, number of edge crossings, average edge length, and distribution uniformity as analysis data, and input the analysis data into the trained quality prediction model to predict the corresponding layout quality; S6: Construct a medical knowledge graph based on the optimized node layout.
2. The automatic construction method of the knowledge graph based on evidence-based medicine according to claim 1, wherein Each node set includes the coordinates of each node and the edges connecting the nodes. The nodes are medical entities, and the edges are the relationships between medical entities. The edges connecting the nodes are represented by tuples, and the tuples include the two nodes connected by the edge.
3. The automatic construction method of a knowledge graph based on evidence-based medicine according to claim 2, characterized in that In the step S503, initialize the population , including Z lizards; the range of both the set label and the one-dimensional search space is ; the expression of the position of each lizard is: ; in the formula, is the initial position of the i-th lizard, is a random number between, .
4. The automatic construction method of a knowledge graph based on evidence-based medicine according to claim 3, characterized in that, In the step S505, the expression of the temperature at the position of the lizard is: ; in the formula, is the temperature at the position of the i-th lizard, , are both constants; Pre-construct a lizard food intake model, and the lizard food intake model is a normal distribution model; The expression for the food intake is as follows: ; where is the food intake of the i-th lizard, is the normalization constant, exp is the exponential function, is the standard deviation of the food intake, is the optimal temperature, and the optimal temperature is the temperature corresponding to the maximum food intake of the lizard; In the step S506, the method for determining whether the lizard enters the cave or the foraging stage includes: If , the lizard enters the cave to avoid the heat, where R is a constant; If , the lizard does not enter the cave to avoid the heat, and the lizard enters the foraging stage; The expression of the cave is as follows: ; where is the cave location, is the location of the lizard with the maximum fitness during the population iteration process, is the location of the lizard with the maximum fitness in the previous population iteration; In the step S508, the method for determining whether the iteration is completed is: If the iteration number t is less than the number threshold T, the iteration is not completed; If the iteration number t is greater than or equal to the number threshold T, the iteration is completed.
5. The automatic construction method of a knowledge graph based on evidence-based medicine according to claim 4, characterized in that, In the step S507, the method for updating the positions of the lizards entering the cave includes: A random number is assigned to each lizard entering the cave , which is a random number between; If , the calculation method for the position of the lizard entering the cave includes: ; ; wherein, is the position of the i-th lizard entering the cave, is a decreasing curve; If , the calculation method for the position of the lizard entering the cave is as follows: ; where is the position of a randomly selected lizard in the population; The method for updating the positions of the lizards entering the foraging stage includes: The expression for the food position is: ; where is the food position; The expression for the food size is: ; where is the food size, is the food factor, = 3, is the fitness corresponding to the position of the i-th lizard, is the fitness corresponding to the food position, is a random number between; If , then the expression for the food position after shredding the food is: ; where is the food position after shredding the food; the calculation method for the position of the lizard entering the foraging stage includes: ; wherein, is the position of the i-th lizard after entering the foraging stage, is a random number between; If , the lizard moves directly towards the food and eats. The calculation method for the position of the lizard entering the foraging stage is as follows: .
6. The automatic construction method of a knowledge graph based on evidence-based medicine according to claim 5, characterized in that, The calculation method for the number of edge crossings includes: Mark the coordinates of the two nodes corresponding to each edge in the analysis set as the first coordinate and the second coordinate respectively. Subtract the second coordinate corresponding to each edge from the first coordinate to obtain the corresponding direction vector for each edge, and label it as the edge vector; Randomly combine all the edge vectors, and use any two edge vectors as an edge set. All the edge sets are different; Label the two edge vectors in each edge set as the first vector and the second vector respectively; subtract the first coordinate of the corresponding first vector from the first coordinate of the second vector in each edge set, and then multiply by the first vector to obtain the first cross product of each edge set; subtract the first coordinate of the corresponding first vector from the second coordinate of the second vector in each edge set, and then multiply by the first vector to obtain the second cross product of each edge set; subtract the first coordinate of the corresponding second vector from the first coordinate of the first vector in each edge set, and then multiply by the second vector to obtain the third cross product of each edge set; subtract the first coordinate of the corresponding second vector from the second coordinate of the first vector in each edge set, and then multiply by the second vector to obtain the fourth cross product of each edge set; where the first coordinate of the edge vector is the first coordinate used when calculating the edge vector, and the second coordinate of the edge vector is the second coordinate used when calculating the edge vector; Label the edge sets where the signs of the first cross product and the second cross product are different, and the signs of the third cross product and the fourth cross product are different as the cross sets; count the number of cross sets as the edge crossing number.
7. A method for automatically constructing a knowledge graph based on evidence-based medicine according to claim 6, characterized in that, The calculation method of the average side length is as follows: subtract the abscissa of the second coordinate from the abscissa of the first coordinate corresponding to each edge to obtain the horizontal difference corresponding to each edge; subtract the ordinate of the second coordinate from the ordinate of the first coordinate corresponding to each edge to obtain the vertical difference corresponding to each edge; take the square root of the sum of the square of the horizontal difference and the square of the vertical difference to obtain the length of each edge; Count the number of edges in the statistical analysis set, add up the lengths of each edge in turn and divide by the number of edges to obtain the average side length; The calculation method of the distribution uniformity is as follows: count the number of node coordinates in the statistical analysis set and label it as the number of nodes; add up the abscissas of each node coordinate in turn and divide by the number of nodes to obtain the horizontal mean value; add up the ordinates of each node coordinate in turn and divide by the number of nodes to obtain the vertical mean value; subtract the horizontal mean value from the abscissa of each node coordinate in turn to obtain the horizontal mean difference value; subtract the vertical mean value from the ordinate of each node coordinate in turn to obtain the vertical mean difference value; add up the squares of each horizontal mean difference value in turn and divide by the number of nodes to obtain the horizontal variance; add up the squares of each vertical mean difference value in turn and divide by the number of nodes to obtain the vertical variance; Take the square root of the sum of the horizontal variance and the vertical variance to obtain the distribution uniformity; The training process of the quality prediction model includes: Pre-collect b sets of analysis data, set corresponding layout qualities for the b sets of analysis data, where b is an integer greater than 1, and convert the analysis data and the corresponding layout qualities into a corresponding set of feature vectors; use each set of feature vectors as the input of the quality prediction model, the quality prediction model outputs a set of predicted layout qualities corresponding to each set of analysis data, uses the actual layout quality corresponding to each set of analysis data as the prediction target, and the actual layout quality is the pre-set layout quality corresponding to the analysis data; use minimizing the sum of the prediction errors of all analysis data as the training target; train the quality prediction model until the sum of the prediction errors reaches convergence and then stop training; the quality prediction model is a deep neural network model.
8. A method for automatically constructing a knowledge graph based on evidence-based medicine according to claim 7, characterized in that, The multi-source medical data is medical-related data from different sources; The steps for preprocessing multi-source medical data include: Step S201: Clean the multi-source medical data; Step S202: Normalize the multi-source medical data; Step S203: Transform the multi-source medical data; Step S204: Integrate the multi-source medical data; The methods for entity recognition of multi-source processed data include: Input the multi-source processed data into a trained entity extraction model to identify medical entities; the entity extraction model is a BERT model, and the principle of the BERT model is: For an input text, decompose it into words, and then convert each word into its corresponding embedding vector; the embedding vector includes an embedded word and an embedded sentence position; The embedded word Et is ; where is the embedding matrix of the word, is the word vector of the original word; The embedding sentence position Es is ; where is the embedding matrix of the position, is the position encoding of the original word in the sentence; The final input representation is ; The core of BERT is a Transformer encoder composed of multiple attention heads; the input representation Einput is input into the Transformer encoder, and through multiple layers of self-attention mechanisms and feed-forward networks, the output representation of the context is obtained; Preset entity labels. After inputting the multi-source processed data into the entity extraction model, obtain the labels of each word, and screen out the words corresponding to the entity labels from all the labels and mark them as medical entities.
9. The automatic construction method of a knowledge graph based on evidence-based medicine according to claim 8, wherein The methods for relationship extraction of multi-source processed data include: The methods for relationship extraction include rule reasoning, graph algorithm reasoning, and model reasoning; Rule reasoning is to use pre-defined rules to extract the relationships between medical entities through the rules and discover new relationships through logical reasoning; Graph algorithm reasoning is to use graph algorithms to extract the relationships between medical entities; Model reasoning is to use a graph neural network model to extract the relationships between medical entities; The principle of the graph neural network model includes: Graph structure representation: For a graph , where V is the set of nodes and A is the set of edges. For each node , there exists a feature vector representing the feature information of the node; Message passing: For node , upon receiving a message from an adjacent node , update the corresponding eigenvector of node . The process of message passing includes two parts: and , where is the feature vector of node in the -th convolutional layer of the graph neural network, is the set of adjacent nodes of node , is the feature vector of the edge connecting node and , including the weight and type information of the edge; Agg is an aggregation function that aggregates the information of the adjacent nodes of node , including summation and averaging, is the activation function, is the feature vector of the edge connecting node and in the L-th convolutional layer of the graph neural network; The feature vectors of the nodes are updated in each convolutional layer of the graph neural network, and through message passing in multiple convolutional layers, a feature vector containing the information and context connections corresponding to multiple nodes is obtained; Relationship reasoning: Through the learned feature vectors of the nodes and the feature vectors of the edges corresponding to the connected nodes, relationship reasoning is carried out.
10. A method for automatically constructing a knowledge graph based on evidence-based medicine according to claim 9, characterized in that, The steps for constructing a medical knowledge graph include: Step S601: According to the optimized node layout, take each medical entity as a node, and create edges between the nodes according to the relationships between the medical entities; Step S602: Select a graph database to store the created nodes and edges, and import the created nodes and edges into the graph database in the supported format of the graph database; Step S603: Use a visualization tool to display the nodes and edges in the graph database to complete the construction of the medical knowledge graph.
Citation Information
Patent Citations
A knowledge graph-based cross-departmental decision support system for early diagnosis of chronic kidney disease
CN111370127B
Knowledge graph visualization construction method suitable for use and display
CN113157942A
Multi-source knowledge graph construction method and system for chronic disease diagnosis and treatment
CN118820486A
Signal management system based on dual-frequency signal
CN119603746A
Cited By
Atrial fibrillation knowledge graph self-evolution system driven by evidence-based medicine and construction method
CN122474345A