Knowledge graph automatic construction method based on evidence-based medicine
By introducing a natural heuristic optimization algorithm into the knowledge graph construction method to optimize node layout, the problem of low query efficiency and visual confusion caused by unreasonable node layout in the existing technology is solved, and the construction of medical knowledge graphs with efficient and good visual effects is achieved.
Patent Information
- Application Number
- CN202510511720.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-04-23
AI Technical Summary
The existing technology does not consider node layout optimization during the knowledge graph construction process, resulting in confusion in the distribution of data nodes and edges, increasing the complexity of query paths, consuming more computing resources, and reducing query efficiency and visualization effects.
An automatic construction method of knowledge graph based on evidence-based medicine is adopted. By collecting multi-source medical data, preprocessing and entity recognition, the relationship between medical entities is extracted, and node layout optimization is optimized using natural heuristic optimization algorithm.
Build a medical knowledge graph with a good structure and high query efficiency, improve the integration and reasoning ability of medical knowledge, improve visualization and use efficiency, and enhance the usability and user experience of the knowledge graph.
Smart Images

Figure CN120032916A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical information technology, and more specifically, to a method for automatically constructing a knowledge graph based on evidence-based medicine. Background Art
[0002] With the continuous deepening of medical research and the rapid growth of clinical data, a huge amount of knowledge and information has been generated in the medical field. However, this knowledge is scattered in different documents, electronic medical records, databases and clinical guidelines. How to effectively integrate, manage and apply this information has become a major challenge in modern medical research and clinical practice. Traditional knowledge management methods have been unable to cope with such a huge amount of data and complex knowledge systems, resulting in low efficiency in information acquisition and application, affecting the scientificity and accuracy of medical decision-making. Evidence-Based Medicine (EBM), as a modern medical model, helps medical staff make scientific medical decisions by systematically evaluating existing evidence. However, in practice, the process of obtaining, integrating and applying this evidence still faces many difficulties. Specific problems include the complexity of evidence acquisition, the cumbersome integration process and the poor real-time performance in clinical practice. In order to improve the efficiency of evidence-based medicine, knowledge graph technology has gradually attracted widespread attention. Knowledge graphs can describe the complex network relationships between entities and relationships by modeling structured and unstructured information in different data sources into a unified graphical structure, which helps to mine and apply deep knowledge hidden in medical data. The Chinese patent with the authorization announcement number CN111370127B discloses a cross-departmental chronic kidney disease early diagnosis decision support system based on knowledge graph; it includes a patient information model building module, a patient information model library storage module, a knowledge graph association module, a knowledge graph reasoning module and a decision support feedback module; the present invention constructs a patient information model and uses the OMOP CDM standard terminology system to construct the patient's electronic medical record data into a patient information model with unified concept coding and unified semantic structure; it takes advantage of semantic technology in data interactivity and scalability, so that the system has better adaptability and scalability to heterogeneous data from different hospitals. At the same time, the clinical recommendations derived from knowledge reasoning based on the knowledge graph are all derived from clinical guidelines and physician experience that conform to evidence-based medicine. The reasoning process and reasons for the recommendations can be traced back by constructing reasoning instances, so that the reasoning process and reasons for the recommendations can be given while giving clinical recommendations, thereby enhancing the physician's trust in the decision support recommendations; However, the above technologies do not take into account node layout optimization during the knowledge graph construction process. The node layout in the knowledge graph (i.e., the position of each node in the graph) will affect the visualization effect and query efficiency of the knowledge graph. If the node layout is unreasonable, it will lead to chaotic distribution of data nodes and edges, increasing the path complexity during query. In addition, more computing resources will be consumed when retrieving and processing requests, resulting in longer response time. This will lead to low query efficiency and chaotic visualization, affecting the relevance and logic between data, and thus reducing the overall usability and user experience of the knowledge graph. In view of this, the present invention proposes an automatic construction method of a knowledge graph based on evidence-based medicine to solve the above problems. Summary of the invention
[0003] In order to overcome the above-mentioned defects of the prior art and to achieve the above-mentioned purpose, the present invention provides the following technical solution: a method for automatically constructing a knowledge graph based on evidence-based medicine, comprising: S1: Collect medical data from multiple sources; S2: Preprocess the multi-source medical data and mark them as multi-source processed data; S3: Perform entity recognition on multi-source processed data to identify medical entities; S4: Perform relationship extraction on multi-source processed data to extract the relationships between medical entities; S5: Node layout is performed based on medical entities and the relationships between medical entities, and optimized using a nature-inspired optimization algorithm; S6: Construct a medical knowledge graph based on the optimized node layout.
[0004] Further, different node layout algorithms are used to perform node layout based on medical entities and the relationships between medical entities, and N node sets are obtained, where N is an integer greater than 1, and each node set includes the coordinates of each node and the edges connecting the nodes. The nodes are medical entities, and the edges are the relationships between medical entities. The edges connecting the nodes are represented by tuples, and the tuples include two nodes connected to the edge. The steps to optimize the node layout include: Step S501: setting different digital labels for different node sets and marking them as set labels; Step S502: Preset the population size Z and the number threshold T; Step S503: Initialize the population. The positions of the lizards in the initialized population are defined in a one-dimensional search space. The positions of the lizards correspond to the set labels one by one. The number of iterations t of the initialized population is 0. Step S504: determining a fitness function; Step S505: defining the temperature and food intake of each lizard's location; Step S506: defining the location of the cave and determining whether the lizard has entered the cave or entered the foraging stage; Step S507: updating the position of the lizard; Step S508: Determine whether the iteration is complete. If not, return to step S505 and set the number of iterations to ; If completed, proceed to step S509; Step S509: Calculate the fitness corresponding to the position of each lizard, obtain the position of the lizard corresponding to the fitness with the largest value, and obtain the node set corresponding to the corresponding set label according to the obtained lizard position.
[0005] Furthermore, in step S503, the population is initialized , including Z lizards; the range of the set label and the one-dimensional search space are ; The expression for the position of each lizard is: ; In the formula, is the initial position of the i-th lizard, for A random number between ; In step S504, the fitness function is expressed as: ; In the formula, is the fitness, BZ is the layout quality; the method for obtaining the layout quality index includes: Get the node set corresponding to the set label corresponding to the lizard position and mark it as the analysis set, calculate the edge crossing number, average edge length and distribution uniformity corresponding to the analysis set; use the analysis set, edge crossing number, average edge length and distribution uniformity as analysis data, input the analysis data into the trained quality prediction model to predict the corresponding layout quality.
[0006] Furthermore, in step S505, the expression of the temperature at the lizard position is: ; In the formula, is the temperature of the i-th lizard’s position, , are all constants; A lizard food intake model is pre-built, and the lizard food intake model is a normal distribution model; The expression for food intake is: ; In the formula, is the food intake of the i-th lizard, is the normalization constant, exp is the exponential function, is the standard deviation of food intake, is the optimal temperature, which is the temperature corresponding to the maximum food intake of the lizard; In step S506, the method for determining whether the lizard has entered a cave or a foraging stage includes: like , then the lizard enters the cave to escape the heat, and R is a constant; like , then the lizards do not enter the cave to escape the heat, and the lizards enter the foraging stage; The expression for the cave is: ; In the formula, For the cave location, is the position of the lizard with the highest fitness during the population iteration process, It is the position of the lizard with the highest fitness in the last population iteration; In step S508, the method for judging whether the iteration is completed is: if the number of iterations t is less than the number threshold T, the iteration is not completed; if the number of iterations t is greater than or equal to the number threshold T, the iteration is completed.
[0007] Furthermore, in step S507, the method for updating the position of the lizard entering the cave includes: Each lizard that enters the cave is given a random number , for A random number between like , then the calculation method for the position of the lizard entering the cave includes: ; ; In the formula, is the position of the lizard that enters the cave, is a decreasing curve; like , then the position of the lizard entering the cave is calculated as: ; In the formula, is the position of a random lizard in the population; Methods for updating the position of a lizard entering the foraging phase include: The expression for food position is: ; In the formula, for food location; The expression for food size is: ; In the formula, For food size, For food factors, =3, is the fitness corresponding to the position of the i-th lizard, is the fitness corresponding to the food position, for A random number between like , then the expression for the position of the food after it is shredded is: ; In the formula, is the position of the food after it is torn into pieces; the calculation method of the position of the lizard entering the foraging stage includes: ; In the formula, is the position of the i-th lizard after it enters the foraging phase, for A random number between like , then the lizard moves directly to the food and eats it. The calculation method for the position of the lizard entering the foraging stage is: .
[0008] Furthermore, the edge crossing number calculation method includes: Mark the two node coordinates corresponding to each edge in the analysis set as the first coordinate and the second coordinate respectively, subtract the corresponding second coordinate from the first coordinate of each edge, obtain each corresponding direction vector, and mark it as the edge vector; randomly combine all edge vectors, regard any two edge vectors as an edge set, and all edge sets are different; Mark the two edge vectors in each edge set as the first vector and the second vector respectively; subtract the first coordinate of the corresponding first vector from the first coordinate of the second vector of each edge set, and then multiply by the first vector to obtain the first cross product of each edge set; subtract the first coordinate of the corresponding first vector from the second coordinate of the second vector of each edge set, and then multiply by the first vector to obtain the second cross product of each edge set; subtract the first coordinate of the corresponding second vector from the first coordinate of the first vector of each edge set, and then multiply by the second vector to obtain the third cross product of each edge set; subtract the first coordinate of the corresponding second vector from the second coordinate of the first vector of each edge set, and then multiply by the second vector to obtain the fourth cross product of each edge set; wherein the first coordinate of the edge vector is the first coordinate used when calculating the edge vector, and the second coordinate of the edge vector is the second coordinate used when calculating the edge vector; The edge sets whose first cross product and second cross product have different signs, and whose third cross product and fourth cross product have different signs are marked as intersection sets; the number of intersection sets is counted as the edge intersection number.
[0009] Furthermore, the average side length is calculated by: subtracting the horizontal coordinate of the second coordinate from the horizontal coordinate of each side corresponding to the first coordinate to obtain the horizontal difference corresponding to each side; subtracting the vertical coordinate of the second coordinate from the vertical coordinate of each side corresponding to the first coordinate to obtain the vertical difference corresponding to each side; adding the square of the horizontal difference to the square of the vertical difference and then taking the square root to obtain the length of each side; statistically analyzing the number of sides in the set, adding the length of each side in turn and dividing by the number of sides to obtain the average side length; The calculation method of distribution uniformity is as follows: statistically analyze the number of node coordinates in the set and mark them as the number of nodes; add the horizontal coordinates of each node coordinate in turn, and then divide them by the number of nodes to obtain the horizontal mean; add the vertical coordinates of each node coordinate in turn, and then divide them by the number of nodes to obtain the vertical mean; subtract the horizontal mean from the horizontal coordinate of each node coordinate in turn to obtain the horizontal mean difference; subtract the vertical mean from the vertical coordinate of each node coordinate in turn to obtain the vertical mean difference; add the squares of each horizontal mean difference in turn and divide them by the number of nodes to obtain the horizontal variance; add the squares of each vertical mean difference in turn and divide them by the number of nodes to obtain the vertical variance; take the square root of the horizontal variance plus the vertical variance to obtain the distribution uniformity; The training process of the quality prediction model includes: Collect b groups of analysis data in advance, set corresponding layout qualities for all b groups of analysis data, b is an integer greater than 1, and convert the analysis data and the corresponding layout qualities into a corresponding set of feature vectors; use each set of feature vectors as input to a quality prediction model, the quality prediction model uses a set of predicted layout qualities corresponding to each set of analysis data as output, and uses the actual layout quality corresponding to each set of analysis data as a prediction target, the actual layout quality is the pre-set layout quality corresponding to the analysis data; minimize the sum of prediction errors of all analysis data as a training target; train the quality prediction model until the sum of prediction errors converges and the training is stopped; the quality prediction model is a deep neural network model.
[0010] Furthermore, the multi-source medical data is medical-related data from different sources; the step of preprocessing the multi-source medical data includes: Step S201: performing data cleaning on multi-source medical data; Step S202: performing data normalization processing on multi-source medical data; Step S203: performing data conversion on multi-source medical data; Step S204: performing data integration on multi-source medical data; Methods for entity recognition on multi-source processing data include: The multi-source processed data is input into the trained entity extraction model to identify medical entities; the entity extraction model is the BERT model, and the principle of the BERT model is: For an input text, decompose it into words, and then convert each word into its corresponding embedding vector; the embedding vector includes the embedded word and the embedded sentence position; The word embedding Et is ;in is the word embedding matrix, is the word vector of the original word; The embedding sentence position Es is ;in is the embedding matrix of the position, Encode the position of the original word in the sentence; The final input is expressed as ; The core of BERT is the Transformer encoder composed of multiple attention heads. The input representation Einput is input into the Transformer encoder, and after multiple layers of self-attention mechanism and feed-forward network, the output representation of the context is obtained. After presetting entity labels, the multi-source processed data is input into the entity extraction model, the label of each word is obtained, and the words corresponding to the entity labels are filtered out from all labels and marked as medical entities.
[0011] Furthermore, the method for extracting relationships from multi-source processed data includes: Relation extraction methods include rule reasoning, graph algorithm reasoning, and model reasoning; Rule reasoning is to use pre-defined rules to extract the relationship between medical entities and discover new relationships through logical reasoning; Graph algorithm reasoning is to use graph algorithms to extract the relationship between medical entities; Model reasoning uses a graph neural network model to extract the relationship between medical entities; The principles of the graph neural network model include: Graph structure representation: For a graph , where V is the node set and A is the edge set. For each node There is a feature vector that represents the feature information of the node; Message passing: For nodes , received from the neighboring node Message, update node The corresponding eigenvector; The message passing process consists of two parts: as well as ,in For Node In graph neural networks The feature vector of the convolutional layer, For Node The set of adjacent nodes of To connect nodes and The feature vector of the edge, including the weight and type information of the edge; Agg is the aggregation function, which aggregates the nodes Aggregate the information of neighboring nodes, including summation and averaging, is the activation function, To connect nodes and The feature vector of the edge in the L-th convolutional layer of the graph neural network; The feature vector of the node is updated at each convolutional layer in the graph neural network, and the feature vector containing the information and contextual connections corresponding to multiple nodes is obtained through message passing through multiple convolutional layers; Relational reasoning: Relational reasoning is performed by learning the feature vectors of nodes and the feature vectors of the edges connecting the nodes.
[0012] Furthermore, the step of constructing a medical knowledge graph includes: Step S601: according to the optimized node layout, each medical entity is regarded as a node, and edges are created between the nodes according to the relationship between the medical entities; Step S602: Select a graph database to store the created nodes and edges, and import the created nodes and edges into the graph database using a format supported by the graph database; Step S603: Use visualization tools to display the nodes and edges in the graph database to complete the construction of the medical knowledge graph.
[0013] Technical effects and advantages of the method for automatically constructing a knowledge graph based on evidence-based medicine of the present invention: By collecting and preprocessing multi-source medical data, using a variety of methods for entity recognition and relationship extraction, and adopting nature-inspired optimization algorithms and deep learning techniques to optimize node layout, a well-structured and highly query-efficient medical knowledge graph is automatically constructed from massive medical data. It can effectively improve the integration and reasoning capabilities of medical knowledge, and at the same time improve the visualization and use efficiency of the medical knowledge graph by optimizing the node layout, enhance the usability and user experience of the medical knowledge graph, and thus improve the scientificity and accuracy of evidence-based medical decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 This is a flow chart of a method for automatically constructing a knowledge graph based on evidence-based medicine according to Embodiment 1 of the present invention; Figure 2 This is a flow chart of a node layout optimization method according to Embodiment 1 of the present invention; Figure 3 This is a schematic diagram of an electronic device according to Embodiment 2 of the present invention; Figure 4 This is a schematic diagram of a storage medium according to Embodiment 3 of the present invention. DETAILED DESCRIPTION
[0015] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0016] Example 1 See also Figure 1 As shown, the method for automatically constructing a knowledge graph based on evidence-based medicine described in this embodiment includes: S1: Collect multi-source medical data.
[0017] Multi-source medical data refers to medical-related data from different sources; multi-source medical data is collected through multiple data sources such as medical literature, clinical trials, case reports, electronic medical records, etc.; multi-source medical data includes: electronic medical record (EMR) data (such as patient information, medical records, laboratory reports, prescription records, etc.), medical imaging data (such as CT, MRI, X-ray, etc.), physiological signal monitoring data (such as electrocardiogram, electroencephalogram, etc.), social media and Internet data (such as patient discussion information from online forums and social platforms), clinical trial data (such as data recorded in various clinical trials), etc.
[0018] It should be noted that the purpose of collecting multi-source medical data is to provide a rich source of medical knowledge: the knowledge graph needs to mine and extract knowledge from various medical data sources, and multi-source data can provide more extensive and comprehensive data support; support knowledge superposition and comparison: different aspects of medical knowledge can be discovered from data sources of different perspectives and types, and the knowledge graph can be enriched through comparison and complementarity; help knowledge representation and extraction: different data sources have different perspectives and formats for expressing knowledge. They can refer to each other for standardization of knowledge representation and discovery of extraction rules; verify and update the knowledge graph: cross-validate the results of knowledge graph construction through multi-source data. Discover new knowledge to update and iterate the graph; support the study of complex problems: complex medical problems such as etiology, diagnosis, and treatment. They need to be studied from multiple perspectives and levels. Multi-source data has become an important information support.
[0019] S2: Preprocess the multi-source medical data and mark them as multi-source processed data.
[0020] The steps for preprocessing multi-source medical data include: Step S201: performing data cleaning on multi-source medical data; removing erroneous, redundant and incomplete data in the multi-source medical data to improve the accuracy and consistency of the data; Step S202: performing data normalization processing on multi-source medical data; unifying the representation of multi-source medical data to ensure data consistency and compatibility of different data sources; Step S203: performing data conversion on multi-source medical data; ensuring that different types of data (such as image data, text data) are represented in a consistent standard format, which helps to improve the consistency and operability of the data; Step S204: Perform data integration on multi-source medical data; merge data from different sources into a unified, comprehensive data set to facilitate subsequent knowledge graph construction.
[0021] The above pretreatment methods are all prior art, and the specific processes will not be described in detail here.
[0022] S3: Perform entity recognition on multi-source processed data to identify medical entities.
[0023] Methods for entity recognition on multi-source processing data include: Input the multi-source processed data into the trained entity extraction model to identify medical entities; medical entities include: diseases (such as diabetes, hypertension, etc.), symptoms (such as headache, cough, etc.), drugs (such as amoxicillin, ibuprofen, etc.), medical equipment (such as CT scanners, electrocardiographs, etc.), etc.
[0024] The entity extraction model is specifically a BERT model. The principle of the BERT model is: For an input text, first decompose it into words, and then convert each word into its corresponding embedding vector; the embedding vector includes the embedded word and the embedded sentence position; The word embedding Et is ;in is the word embedding matrix, is the word vector of the original word; The embedding sentence position Es is ;in is the embedding matrix of the position, Encode the position of the original word in the sentence; The final input is expressed as ; The core of BERT is the Transformer encoder composed of multiple attention heads. The input representation Einput is input into the Transformer encoder, and after multiple layers of self-attention mechanism and feed-forward network, the context-rich output representation is obtained. To adapt to specific tasks, add a task-specific layer, such as a fully connected layer, to the output of BERT; for entity extraction tasks, a fully connected layer Wner can be used to predict whether each word is a named entity: and according to the specific requirements of the task, define rules to mark and extract medical entities, such as dates usually have a specific format that can be matched by regular expressions; After presetting entity labels, the multi-source processed data is input into the entity extraction model, the label of each word is obtained, and the words corresponding to the entity labels are filtered out from all labels and marked as medical entities.
[0025] The BERT model is an existing technology, so the training process of the entity extraction model will not be described in detail here; the entity labels are pre-set by technicians in this field.
[0026] S4: Perform relationship extraction on multi-source processed data to extract the relationships between medical entities.
[0027] Methods for extracting relationships from multi-source processing data include: Methods of relationship extraction include rule reasoning, graph algorithm reasoning, and model reasoning; relationships between medical entities, for example: the relationship between drugs or treatments and diseases (such as amoxicillin for bacterial infections), the causal relationship between viruses and diseases (such as influenza viruses causing influenza), and descriptions of the relationship between symptoms and diseases (such as fever is related to infection).
[0028] Rule reasoning uses predefined rules to extract relationships between medical entities and discover new relationships through logical reasoning. For example, if one medical entity is a doctor and another medical entity is a patient, and there are diagnosis records in multi-source medical data, then there is a diagnostic relationship between the two medical entities; for example, if disease A and disease B use the same medicine, and disease B and disease C also use the same medicine, then there is also a relationship between disease A and disease C.
[0029] Graph algorithm reasoning uses graph algorithms to extract the relationship between medical entities, such as graph traversal algorithms, shortest path algorithms, etc. The shortest path algorithm, for example, uses depth-first search or breadth-first search to obtain the shortest path between two medical entities. The shortest path can reveal the shortest communication or shortest connection method between medical entities, thereby indicating the closeness of the relationship between medical entities.
[0030] Model reasoning uses graph neural network models to extract the relationship between medical entities; graph neural network models include graph convolutional networks (GCN), graph attention networks (GAT) and graph isomorphism networks (GIN).
[0031] The principles of the graph neural network model include: Graph structure representation: For a graph , where V is the node set and A is the edge set. For each node There is a feature vector that represents the feature information of the node.
[0032] Messaging: The core of the graph neural network model is to enable nodes to aggregate information from adjacent nodes through a message passing mechanism. , received from the neighboring node Message, update node The corresponding eigenvector; The message passing process consists of two parts: as well as ,in For Node In graph neural networks The feature vector of the convolutional layer, For Node The set of adjacent nodes of To connect nodes and The feature vector of the edge, including the weight and type information of the edge; Agg is the aggregation function, which aggregates the nodes Aggregate the information of neighboring nodes, including summation and averaging. The specific choice depends on the task and network design. is the activation function, To connect nodes and The feature vector of the edge in the L-th convolutional layer of the graph neural network; The feature vector of the node is updated in each convolutional layer in the graph neural network, and a feature vector containing information corresponding to multiple nodes and context connections is obtained through message passing through multiple convolutional layers.
[0033] Relational Reasoning: By learning the feature vectors of nodes and the feature vectors of the edges connecting the nodes, relational reasoning can be performed, such as classification and regression tasks.
[0034] S5: Node layout is performed based on medical entities and the relationships between medical entities, and optimized using a nature-inspired optimization algorithm.
[0035] Different node layout algorithms are used to perform node layout based on medical entities and the relationships between medical entities, and N node sets are obtained, where N is an integer greater than 1. Each node set includes the coordinates of each node and the edges connecting the nodes. The nodes are medical entities, and the edges are the relationships between medical entities. The edges connecting the nodes are represented by tuples, which include two nodes connected to the edges, such as (node 1, node 2). Node layout algorithms include hierarchical layout, circular layout, principal component analysis layout, force-guided layout, etc.
[0036] like Figure 2 As shown, the steps for optimizing the node layout include: Step S501: setting different digital labels for different node sets and marking them as set labels; Step S502: Preset the population size Z and the number threshold T; Step S503: Initialize the population. The positions of the lizards in the initialized population are defined in a one-dimensional search space. The positions of the lizards correspond to the set labels one by one. The number of iterations t of the initialized population is 0. Step S504: determining a fitness function; Step S505: defining the temperature and food intake of each lizard's location; Step S506: defining the location of the cave and determining whether the lizard has entered the cave or entered the foraging stage; Step S507: updating the position of the lizard; Step S508: Determine whether the iteration is complete. If not, return to step S505 and set the number of iterations to ; If completed, proceed to step S509; Step S509: Calculate the fitness corresponding to the position of each lizard, obtain the position of the lizard corresponding to the fitness with the largest value, and obtain the node set corresponding to the corresponding set label according to the obtained lizard position.
[0037] In the above step S502, the population size Z is determined by a technician in this field. During the historical node layout optimization process, under multiple different node set conditions, multiple different population sizes are set for the same node set, and the lizard optimization algorithm is performed multiple times. After the same number of iterations, the corresponding set label is obtained, and the population size with the same set label as the actual set label is used as the population size corresponding to the group of test data; the actual set label is the digital label of the node set that best matches the node set, and the actual set label is obtained by a technician in this field through experiments based on actual experience; and the population size corresponding to each node set is obtained in this way, and the average of the multiple population sizes is used as the preset population size Z.
[0038] The number threshold T is determined by a technician in this field. During the historical node layout optimization process, under multiple different node set conditions, the same node set is subjected to multiple lizard optimization algorithms to obtain multiple set labels, wherein the number of iterations of the lizard optimization algorithm is different each time, and the population size is the same and is Z; the number of iterations that is closest to the set label and the actual set label is used as the number of iterations corresponding to the node set; and the number of iterations corresponding to each node set is obtained by analogy, and the average of multiple iterations is used as the number threshold T.
[0039] It should be understood that the population size Z determines the breadth of the search. A larger population size can explore more solution spaces, while the number threshold T determines the termination condition of the algorithm and can control the speed at which the algorithm converges.
[0040] In the above step S503, the population is initialized , including Z lizards.
[0041] The range of the set label and the one-dimensional search space are ; The expression for the position of each lizard is: ; In the formula, is the initial position of the i-th lizard, for A random number between .
[0042] In the above step S504, the fitness function is expressed as: ; In the formula, is the fitness, BZ is the layout quality; the method for obtaining the layout quality index includes: Get the node set corresponding to the set label corresponding to the lizard position and mark it as the analysis set, calculate the edge crossing number, average edge length and distribution uniformity corresponding to the analysis set; use the analysis set, edge crossing number, average edge length and distribution uniformity as analysis data, input the analysis data into the trained quality prediction model to predict the corresponding layout quality.
[0043] The training process of the quality prediction model includes: Collect b groups of analysis data in advance, set corresponding layout qualities for the b groups of analysis data, b is an integer greater than 1, and convert the analysis data and the corresponding layout qualities into a corresponding set of feature vectors; the layout qualities corresponding to the analysis data are collected by technicians in this field during the historical node layout optimization process, and the layout qualities are evaluated according to the actual situation under the conditions of each analysis set; and the corresponding layout qualities are set for the b groups of analysis data in sequence; Each set of feature vectors is used as the input of the quality prediction model. The quality prediction model takes a set of predicted layout qualities corresponding to each set of analysis data as output, and takes the actual layout quality corresponding to each set of analysis data as the prediction target. The actual layout quality is the pre-set layout quality corresponding to the analysis data. Minimizing the sum of the prediction errors of all analysis data is used as the training target. The calculation formula of the prediction error is: ,in is the prediction error, k is the group number of the eigenvector corresponding to the analysis data, is the predicted layout quality corresponding to the kth group of analysis data, is the actual layout quality corresponding to the kth group of analysis data; the quality prediction model is trained until the sum of the prediction errors reaches convergence and the training is stopped.
[0044] The above-mentioned quality prediction model is specifically a deep neural network model; it includes an input layer, a hidden layer and an output layer; each hidden layer includes multiple neurons, each neuron is connected to the neurons in the next layer, and the connection contains weights, which determine the importance and influence of data transmission in the neural network; an activation function is applied to each neuron between the hidden layer and the output layer, and the activation function introduces nonlinearity, allowing the network to learn more complex patterns and features.
[0045] The calculation method of edge crossing number includes: Mark the two node coordinates corresponding to each edge in the analysis set as the first coordinate and the second coordinate respectively, subtract the corresponding second coordinate from the first coordinate of each edge, obtain each corresponding direction vector, and mark it as the edge vector; randomly combine all edge vectors, regard any two edge vectors as an edge set, and all edge sets are different; Mark the two edge vectors in each edge set as the first vector and the second vector respectively; subtract the first coordinate of the corresponding first vector from the first coordinate of the second vector of each edge set, and then multiply by the first vector to obtain the first cross product of each edge set; subtract the first coordinate of the corresponding first vector from the second coordinate of the second vector of each edge set, and then multiply by the first vector to obtain the second cross product of each edge set; subtract the first coordinate of the corresponding second vector from the first coordinate of the first vector of each edge set, and then multiply by the second vector to obtain the third cross product of each edge set; subtract the first coordinate of the corresponding second vector from the second coordinate of the first vector of each edge set, and then multiply by the second vector to obtain the fourth cross product of each edge set; wherein the first coordinate of the edge vector is the first coordinate used when calculating the edge vector, and the second coordinate of the edge vector is the second coordinate used when calculating the edge vector; Mark the edge set whose first cross product and second cross product have different signs, and whose third cross product and fourth cross product have different signs as a cross set; count the number of cross sets as the edge crossing number; illustratively, the edge set corresponds to the first cross product of 2, the second cross product of -1, the third cross product of 3, and the fourth cross product of -4, so the edge set is marked as a cross set.
[0046] The calculation method of average side length includes: Subtract the horizontal coordinate of the second coordinate from the horizontal coordinate of each side corresponding to the first coordinate to obtain the horizontal difference corresponding to each side; subtract the vertical coordinate of the second coordinate from the vertical coordinate of each side corresponding to the first coordinate to obtain the vertical difference corresponding to each side; add the square of the horizontal difference to the square of the vertical difference and then take the square root to obtain the length of each side; statistically analyze the number of sides in the set, add the length of each side in turn and divide it by the number of sides to obtain the average side length.
[0047] The calculation methods of distribution uniformity include: The number of node coordinates in the statistical analysis set is marked as the number of nodes; the horizontal coordinates of each node coordinate are added in turn, and then divided by the number of nodes to obtain the horizontal mean; the vertical coordinates of each node coordinate are added in turn, and then divided by the number of nodes to obtain the vertical mean; the horizontal mean is subtracted from the horizontal coordinate of each node coordinate in turn to obtain the horizontal mean difference; the vertical mean is subtracted from the vertical coordinate of each node coordinate in turn to obtain the vertical mean difference; the squares of each horizontal mean difference are added in turn and divided by the number of nodes to obtain the horizontal variance; the squares of each vertical mean difference are added in turn and divided by the number of nodes to obtain the vertical variance; the horizontal variance plus the vertical variance are squared to obtain the distribution uniformity.
[0048] It should be noted that the number of edge crossings, average edge length and distribution uniformity are all factors affecting layout quality. The reason is that the number of edge crossings reflects whether there are crossings between the lines between nodes in the knowledge graph; a lower number of edge crossings usually means a clearer node layout and more intuitive information transmission, which improves readability and comprehension, so the layout quality is better; if there are too many edge crossings, it may cause information confusion and reduce the usability of the knowledge graph, so the layout quality is poor; the average edge length indicates the average length of the edges connecting the nodes; shorter edge lengths mean that the nodes are more closely connected, which is conducive to the rapid transmission of information and enhances the overall visibility and interactivity of the knowledge graph. The longer the edge, the lower the quality of the layout; the longer the edge, the less compact the layout, which affects the efficiency of information transmission, so the layout quality is poor; the uniformity of distribution measures the distribution status of nodes in the graph; uniform distribution can ensure the balance of information and visual harmony, making it easier for users to capture important information and relationships, so the layout quality is good; if the nodes are unevenly distributed, some areas may be too dense and other areas sparse, causing cognitive burden on users, so the layout quality is poor; in summary, these three indicators jointly affect the readability, query efficiency and visualization effect of the knowledge graph, and are therefore important parameters for evaluating the layout quality of the knowledge graph.
[0049] In the above step S505, the expression of the temperature at the lizard position is: ; In the formula, is the temperature of the i-th lizard’s position, , are constants, , .
[0050] Pre-constructing a lizard food intake model, which is a normal distribution model and is used to predict the food intake of lizards at different temperatures; the lizard food intake model is pre-constructed by a technician in the field who collects food intake data of lizards at different temperatures in advance; The expression for food intake is: ; In the formula, is the food intake of the i-th lizard, is a normalization constant used to adjust the lizard food intake model to ensure that the sum of the probabilities of all values is 1, exp is an exponential function, is the standard deviation of food intake, is the optimal temperature, which is the temperature corresponding to the maximum food intake of the lizard. and optimal temperature All are obtained based on the lizard food intake model.
[0051] In the above step S506, the method for determining whether the lizard has entered a cave or entered a foraging stage includes: like , then the lizard enters the cave to escape the heat, indicating that the temperature where the lizard is located is too high. R is a constant. ; like , the lizard does not enter the cave to escape the heat, which means that the temperature where the lizard is located is suitable for the lizard to eat, and the lizard enters the foraging stage.
[0052] The expression for the cave is: ; In the formula, For the cave location, is the position of the lizard with the highest fitness during the population iteration process, is the position of the lizard with the highest fitness in the last population iteration; and The difference is that is the position of the lizard with the highest fitness during multiple iterations of the population, and Only the position of the lizard with the highest fitness in the last population iteration.
[0053] In the above step S507, the method for updating the position of the lizard entering the cave includes: Each lizard that enters the cave is given a random number , for A random number between like , indicating that the corresponding cave entered by the lizard has no competition from other lizards. The calculation method of the position of the lizard entering the cave includes: ; ; In the formula, is the position of the lizard that enters the cave, It is a decreasing curve.
[0054] like , indicating that the cave entered by the corresponding lizard has other lizards competing. The calculation method of the position of the lizard entering the cave includes: ; In the formula, is the position of a random lizard in the population.
[0055] Methods for updating the position of a lizard entering the foraging phase include: When eating, lizards will choose whether to tear the food into pieces based on the size of the food; if the food is of appropriate size, the lizard will directly ingest the food; if the food is too large, the lizard will use its sharp teeth and powerful jaw muscles to tear the food into pieces before ingesting it.
[0056] The expression for food position is: ; In the formula, For food location.
[0057] The expression for food size is: ; In the formula, For food size, For food factors, =3, is the fitness corresponding to the position of the i-th lizard, is the fitness corresponding to the food position, for A random number between .
[0058] like , indicating that the food is too big and the lizard needs to tear it into pieces; the expression of the position of the food after tearing it into pieces is: ; In the formula, The position of food after it has been shredded.
[0059] The calculation method for the position of the lizard entering the foraging phase includes: ; In the formula, is the position of the i-th lizard after it enters the foraging phase, for A random number between .
[0060] like , then the lizard moves directly to the food and eats it. The calculation method of the position of the lizard entering the foraging stage includes: .
[0061] In the above step S508, the method for determining whether the iteration is completed is: if the number of iterations t is less than the number threshold T, the iteration is not completed, and the number of iterations is set to That is, the value of the number of iterations is increased by one and then assigned the number of iterations; if the number of iterations t is greater than or equal to the number threshold T, the iteration is completed.
[0062] It should be noted that the reason for using the lizard optimization algorithm to optimize the node layout is that the lizard algorithm simulates the behaviors of lizards such as foraging, summer heat avoidance and competition in different environments, and can perform global search in complex multi-dimensional search space; by simulating the foraging and burrowing behaviors of lizards, the algorithm can effectively avoid falling into local optimality, thereby increasing the probability of finding the global optimal solution; and the lizard algorithm can adapt to different node layout optimization problems by adjusting the population size Z and the number threshold T; the optimization process can dynamically adjust the population size and the number of iterations according to historical data to ensure the applicability and efficiency of the algorithm; at the same time, by reasonably setting the fitness function and the burrowing update mechanism, the lizard optimization algorithm can accelerate convergence and shorten the optimization time, which means that it can quickly find the optimal node layout method and improve the efficiency of knowledge graph construction; in addition, the lizard optimization algorithm can dynamically adjust the behavior of lizards (such as directly eating or tearing food into pieces) during the foraging stage to achieve refined processing of local search; this search strategy that takes both local and global considerations improves the performance of the algorithm in complex multi-peak optimization problems.
[0063] S6: Construct a medical knowledge graph based on the optimized node layout; The steps to build a medical knowledge graph include: Step S601: according to the optimized node layout, each medical entity is regarded as a node, and edges are created between the nodes according to the relationship between the medical entities; Step S602: Select a graph database (such as Neo4j, ArangoDB, etc.) to store the created nodes and edges, and import the created nodes and edges into the graph database using a format supported by the graph database (such as CSV, JSON, etc.); Step S603: Use visualization tools (such as Gephi, Cytoscape, etc.) to display the nodes and edges in the graph database to complete the construction of the medical knowledge graph.
[0064] This embodiment collects and preprocesses multi-source medical data, uses a variety of methods for entity recognition and relationship extraction, and adopts nature-inspired optimization algorithms and deep learning technologies to optimize node layout, so as to automatically construct a medical knowledge graph with good structure and high query efficiency from massive medical data; it can effectively improve the integration and reasoning capabilities of medical knowledge, and at the same time improve the visualization effect and use efficiency of the medical knowledge graph by optimizing the node layout, enhance the usability and user experience of the medical knowledge graph, and thus improve the scientificity and accuracy of evidence-based medical decision-making.
[0065] Example 2 See also Figure 3As shown, the present application also provides an electronic device 500. The electronic device 500 may include one or more processors and one or more memories. The memories store computer readable codes, and when the computer readable codes are executed by one or more processors, the method for automatically constructing a knowledge graph based on evidence-based medicine as described above may be executed.
[0066] The method or system according to the embodiment of the present application can also be used by Figure 3 The electronic device architecture shown in FIG. Figure 3 As shown, the electronic device 500 may include a bus 501, one or more CPUs 502, ROM 503, RAM 504, a communication port 505 connected to a network, an input / output 506, a hard disk 507, etc. The storage device in the electronic device 500, such as ROM 503 or hard disk 507, may store a method for automatically constructing a knowledge graph based on evidence-based medicine provided in the present application. Furthermore, the electronic device 500 may also include a user interface 508. Of course, Figure 3 The architecture shown is only exemplary and can be omitted according to actual needs when implementing different devices. Figure 3 One or more components of an electronic device are shown.
[0067] Example 3 See also Figure 4 As shown, one embodiment of the present application discloses a computer-readable storage medium 600. Computer-readable instructions are stored on the computer-readable storage medium 600. When the computer-readable instructions are executed by the processor, a method for automatically constructing a knowledge graph based on evidence-based medicine according to an embodiment of the present application described with reference to the above figures can be executed. The storage medium 600 includes, but is not limited to, for example, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and cache memory (cache), etc. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.
[0068] In addition, according to the embodiments of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the present application provides a non-transitory machine-readable storage medium, which stores machine-readable instructions, and the machine-readable instructions can be executed by a processor to execute instructions corresponding to the method steps provided in the present application, for example: a method for automatically constructing a knowledge graph based on evidence-based medicine. When the computer program is executed by a central processing unit (CPU), the above functions defined in the method of the present application are executed.
[0069] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
[0070] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for automatically constructing a knowledge graph based on evidence-based medicine, characterized in that: include: S1: Collect medical data from multiple sources; S2: Preprocess the multi-source medical data and mark them as multi-source processed data; S3: Perform entity recognition on multi-source processed data to identify medical entities; S4: Perform relationship extraction on multi-source processed data to extract the relationships between medical entities; S5: Node layout is performed based on medical entities and the relationships between medical entities, and optimized using a nature-inspired optimization algorithm; S6: Construct a medical knowledge graph based on the optimized node layout.
2. The method for automatically constructing a knowledge graph based on evidence-based medicine according to claim 1, characterized in that: Different node layout algorithms are used to perform node layout based on medical entities and the relationships between medical entities, and N node sets are obtained, where N is an integer greater than 1. Each node set includes the coordinates of each node and the edges connecting the nodes. The nodes are medical entities, and the edges are the relationships between medical entities. The edges connecting the nodes are represented by tuples, and the tuples include the two nodes connected to the edge. The steps to optimize the node layout include: Step S501: setting different digital labels for different node sets and marking them as set labels; Step S502: Preset the population size Z and the number threshold T; Step S503: Initialize the population. The positions of the lizards in the initialized population are defined in a one-dimensional search space. The positions of the lizards correspond to the set labels one by one. The number of iterations t of the initialized population is 0. Step S504: determining a fitness function; Step S505: defining the temperature and food intake of each lizard's location; Step S506: defining the location of the cave and determining whether the lizard has entered the cave or entered the foraging stage; Step S507: updating the position of the lizard; Step S508: Determine whether the iteration is complete. If not, return to step S505 and set the number of iterations to ; If completed, proceed to step S509; Step S509: Calculate the fitness corresponding to the position of each lizard, obtain the position of the lizard corresponding to the fitness with the largest value, and obtain the node set corresponding to the corresponding set label according to the obtained lizard position.
3. The method for automatically constructing a knowledge graph based on evidence-based medicine according to claim 2, characterized in that: In step S503, the population is initialized , including Z lizards; the range of the set label and the one-dimensional search space are ; The expression for the position of each lizard is: ; In the formula, is the initial position of the i-th lizard, for A random number between ; In step S504, the fitness function is expressed as: ; In the formula, is the fitness, BZ is the layout quality; Methods for obtaining layout quality indicators include: Get the node set corresponding to the set label corresponding to the lizard position and mark it as the analysis set, calculate the edge crossing number, average edge length and distribution uniformity corresponding to the analysis set; use the analysis set, edge crossing number, average edge length and distribution uniformity as analysis data, input the analysis data into the trained quality prediction model to predict the corresponding layout quality.
4. The method for automatically constructing a knowledge graph based on evidence-based medicine according to claim 3, characterized in that: In step S505, the expression of the temperature at the lizard position is: ; In the formula, is the temperature of the i-th lizard’s position, , are all constants; A lizard food intake model is pre-built, and the lizard food intake model is a normal distribution model; The expression for food intake is: ; In the formula, is the food intake of the i-th lizard, is the normalization constant, exp is the exponential function, is the standard deviation of food intake, is the optimal temperature, which is the temperature corresponding to the maximum food intake of the lizard; In step S506, the method for determining whether the lizard has entered a cave or a foraging stage includes: like , then the lizard enters the cave to escape the heat, and R is a constant; like , then the lizards do not enter the cave to escape the heat, and the lizards enter the foraging stage; The expression for the cave is: ; In the formula, For the cave location, is the position of the lizard with the highest fitness during the population iteration process, It is the position of the lizard with the highest fitness in the last population iteration; In step S508, the method for judging whether the iteration is completed is: if the number of iterations t is less than the number threshold T, the iteration is not completed; if the number of iterations t is greater than or equal to the number threshold T, the iteration is completed.
5. The method for automatically constructing a knowledge graph based on evidence-based medicine according to claim 4, characterized in that: In step S507, the method for updating the position of the lizard entering the cave includes: Each lizard that enters the cave is given a random number , for A random number between like , then the calculation method for the position of the lizard entering the cave includes: ; ; In the formula, is the position of the lizard that enters the cave, is a decreasing curve; like , then the position of the lizard entering the cave is calculated as: ; In the formula, is the position of a random lizard in the population; Methods for updating the position of a lizard entering the foraging phase include: The expression for food position is: ; In the formula, for food location; The expression for food size is: ; In the formula, For food size, For food factors, =3, is the fitness corresponding to the position of the i-th lizard, is the fitness corresponding to the food position, for A random number between like , then the expression for the position of the food after it is shredded is: ; In the formula, is the position of the food after it is torn into pieces; the calculation method of the position of the lizard entering the foraging stage includes: ; In the formula, is the position of the i-th lizard after it enters the foraging phase, for A random number between like , then the lizard moves directly to the food and eats it. The calculation method for the position of the lizard entering the foraging stage is: .
6. The method for automatically constructing a knowledge graph based on evidence-based medicine according to claim 5, characterized in that: The calculation method of edge crossing number includes: Mark the two node coordinates corresponding to each edge in the analysis set as the first coordinate and the second coordinate respectively, subtract the corresponding second coordinate from the first coordinate of each edge, obtain each corresponding direction vector, and mark it as the edge vector; randomly combine all edge vectors, regard any two edge vectors as an edge set, and all edge sets are different; Mark the two edge vectors in each edge set as the first vector and the second vector respectively; subtract the first coordinate of the corresponding first vector from the first coordinate of the second vector of each edge set, and then multiply by the first vector to obtain the first cross product of each edge set; subtract the first coordinate of the corresponding first vector from the second coordinate of the second vector of each edge set, and then multiply by the first vector to obtain the second cross product of each edge set; subtract the first coordinate of the corresponding second vector from the first coordinate of the first vector of each edge set, and then multiply by the second vector to obtain the third cross product of each edge set; subtract the first coordinate of the corresponding second vector from the second coordinate of the first vector of each edge set, and then multiply by the second vector to obtain the fourth cross product of each edge set; wherein the first coordinate of the edge vector is the first coordinate used when calculating the edge vector, and the second coordinate of the edge vector is the second coordinate used when calculating the edge vector; The edge sets whose first cross product and second cross product have different signs, and whose third cross product and fourth cross product have different signs are marked as intersection sets; the number of intersection sets is counted as the edge intersection number.
7. The method for automatically constructing a knowledge graph based on evidence-based medicine according to claim 6, characterized in that: The calculation method of the average side length is as follows: subtract the horizontal coordinate of the second coordinate from the horizontal coordinate of each side corresponding to the first coordinate to obtain the horizontal difference corresponding to each side; subtract the vertical coordinate of the second coordinate from the vertical coordinate of each side corresponding to the first coordinate to obtain the vertical difference corresponding to each side; add the square of the horizontal difference to the square of the vertical difference and then take the square root to obtain the length of each side; Statistically analyze the number of edges in the set, add the length of each edge in turn and divide it by the number of edges to get the average edge length; The calculation method of distribution uniformity is as follows: statistically analyze the number of node coordinates in the set and mark it as the number of nodes; add the horizontal coordinates of each node coordinate in turn, and then divide it by the number of nodes to obtain the horizontal mean; add the vertical coordinates of each node coordinate in turn, and then divide it by the number of nodes to obtain the vertical mean; subtract the horizontal mean from the horizontal coordinate of each node coordinate in turn to obtain the horizontal mean difference; subtract the vertical mean from the vertical coordinate of each node coordinate in turn to obtain the vertical mean difference; add the squares of each horizontal mean difference in turn and divide it by the number of nodes to obtain the horizontal variance; add the squares of each vertical mean difference in turn and divide it by the number of nodes to obtain the vertical variance; Add the horizontal variance to the vertical variance and then take the square root to obtain the uniformity of distribution; The training process of the quality prediction model includes: Collect b groups of analysis data in advance, set corresponding layout qualities for all b groups of analysis data, b is an integer greater than 1, and convert the analysis data and the corresponding layout qualities into a corresponding set of feature vectors; use each set of feature vectors as input to a quality prediction model, the quality prediction model uses a set of predicted layout qualities corresponding to each set of analysis data as output, and uses the actual layout quality corresponding to each set of analysis data as a prediction target, the actual layout quality is the pre-set layout quality corresponding to the analysis data; minimize the sum of prediction errors of all analysis data as a training target; train the quality prediction model until the sum of prediction errors converges and the training is stopped; the quality prediction model is a deep neural network model.
8. The method for automatically constructing a knowledge graph based on evidence-based medicine according to claim 7, characterized in that: The multi-source medical data are medical-related data from different sources; The step of preprocessing the multi-source medical data comprises: Step S201: performing data cleaning on multi-source medical data; Step S202: performing data normalization processing on multi-source medical data; Step S203: performing data conversion on multi-source medical data; Step S204: performing data integration on multi-source medical data; Methods for entity recognition on multi-source processing data include: The multi-source processed data is input into the trained entity extraction model to identify medical entities; the entity extraction model is the BERT model, and the principle of the BERT model is: For an input text, decompose it into words, and then convert each word into its corresponding embedding vector; the embedding vector includes the embedded word and the embedded sentence position; The word embedding Et is ;in is the word embedding matrix, is the word vector of the original word; The embedding sentence position Es is ;in is the embedding matrix of the position, Encode the position of the original word in the sentence; The final input is expressed as ; The core of BERT is the Transformer encoder composed of multiple attention heads. The input representation Einput is input into the Transformer encoder, and after multiple layers of self-attention mechanism and feed-forward network, the output representation of the context is obtained. After presetting entity labels, the multi-source processed data is input into the entity extraction model, the label of each word is obtained, and the words corresponding to the entity labels are filtered out from all labels and marked as medical entities.
9. The method for automatically constructing a knowledge graph based on evidence-based medicine according to claim 8, characterized in that: The method for extracting relations from multi-source processed data comprises: Relation extraction methods include rule reasoning, graph algorithm reasoning, and model reasoning; Rule reasoning is to use pre-defined rules to extract the relationship between medical entities and discover new relationships through logical reasoning; Graph algorithm reasoning is to use graph algorithms to extract the relationship between medical entities; Model reasoning uses a graph neural network model to extract the relationship between medical entities; The principles of the graph neural network model include: Graph structure representation: For a graph , where V is the node set and A is the edge set. For each node There is a feature vector that represents the feature information of the node; Message passing: For nodes , received from the neighboring node Message, update node The corresponding eigenvector; The message passing process consists of two parts: as well as ,in For Node In graph neural networks The feature vector of the convolutional layer, For Node The set of adjacent nodes of To connect nodes and The feature vector of the edge, including the weight and type information of the edge; Agg is the aggregation function, which aggregates the nodes Aggregate the information of neighboring nodes, including summation and averaging, is the activation function, To connect nodes and The feature vector of the edge in the Lth convolutional layer of the graph neural network; The feature vector of the node is updated at each convolutional layer in the graph neural network, and the feature vector containing the information and contextual connections corresponding to multiple nodes is obtained through message passing through multiple convolutional layers; Relational reasoning: Relational reasoning is performed by learning the feature vectors of nodes and the feature vectors of the edges connecting the nodes.
10. The method for automatically constructing a knowledge graph based on evidence-based medicine according to claim 9, characterized in that: The steps of constructing the medical knowledge graph include: Step S601: according to the optimized node layout, each medical entity is regarded as a node, and edges are created between the nodes according to the relationship between the medical entities; Step S602: Select a graph database to store the created nodes and edges, and import the created nodes and edges into the graph database using a format supported by the graph database; Step S603: Use visualization tools to display the nodes and edges in the graph database to complete the construction of the medical knowledge graph.
Citation Information
Patent Citations
A knowledge graph-based cross-departmental decision support system for early diagnosis of chronic kidney disease
CN111370127B
Knowledge map building system and method
CN109804364A
Knowledge graph visualization construction method suitable for use and display
CN113157942A
Multi-source knowledge graph construction method and system for chronic disease diagnosis and treatment
CN118820486A
Novel feature selection method suitable for disease diagnosis and based on horny lizard algorithm
CN119400400A