Medical text-oriented intelligent disease diagnosis causal atlas construction method and system
Through the intelligent disease diagnosis causal map construction method for medical text, and the feature enhancement model and entity unified technology are used to solve the problem of the ‘black box’ of the intelligent diagnosis system and the inefficient construction of causal knowledge graphs, and an intelligent diagnosis system with high accuracy and interpretability is achieved.
Patent Information
- Application Number
- CN202510267229.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-24
AI Technical Summary
The existing intelligent diagnostic system has the characteristics of ‘black box’ in model construction and reasoning, which is difficult to inspire doctors to analyze and make decisions on complex diseases, and there are problems of inefficiency and incomplete knowledge and data when building a causal knowledge graph for large-scale disease diagnosis.
Using an intelligent disease diagnosis causal map construction method for medical text, through medical named entity recognition and relationship extraction, the MedKE model enhanced by knowledge features and the GCNSE model enhanced by structural feature, combined with entity embedding and path reasoning, a causal relationship map is constructed, and the graph redundancy is reduced through entity uniformity.
The automated construction of causal knowledge graphs is realized, the accuracy and interpretability of intelligent diagnosis are improved, the redundancy of graphs is reduced, and the domain adaptability of the model is enhanced.
Smart Images

Figure CN120197678A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technology of disease causal map processing, and in particular to an intelligent disease diagnosis causal map construction method and system for medical texts. Background Art
[0002] At present, the clinical medical diagnosis and treatment process is a logical causal thinking process facing complex situations with uncertainty. An effective intelligent diagnosis system must provide interpretable and persuasive diagnostic decisions, and intuitively present the causal logic matching the etiology and pathogenesis. Doctors and patients not only need to know what the conclusion is, but also how the conclusion is derived. However, many current machine learning methods have the "black box" feature in model construction and reasoning methods. When the model expression cannot effectively correspond to medical knowledge, the reasoning and calculation process is simplified and mechanized, making it difficult to inspire doctors' analysis and decision-making for complex conditions, which greatly limits the clinical application value of intelligent medical systems.
[0003] Causal knowledge graphs, with their intuitive and clear logical forms and uncertainty quantification expression mechanisms, can clearly and accurately express the causal relationships between various diseases and various symptoms, signs, examination and test factors. They can assist physicians in deeply understanding the pathogenesis and pathophysiological processes of diseases, mining the causal correlation between diseases and corresponding symptoms, and improving the accuracy and interpretability of intelligent diagnosis.
[0004] In the process of constructing a large-scale disease diagnosis causal knowledge graph, if it completely relies on the professional knowledge of clinical doctors and medical experts, there will be problems such as low construction efficiency, incomplete and unreliable knowledge and data. There is a vast amount of information on medical research literature, clinical reports, electronic medical records, online disease consultations, diagnosis and treatment suggestions, etc. on the Internet. These medical texts contain rich and detailed descriptions of disease information, diagnostic paths, and causal relationship explanations. In-depth mining and automated analysis of these network medical text materials are of great significance for assisting the construction of disease diagnosis causal knowledge graphs and disease case libraries, and can effectively assist the research and development of intelligent clinical diagnosis systems. Summary of the Invention
[0005] The main purpose of the present invention is to provide an intelligent disease diagnosis causal map construction method and system for medical texts, which can extract medical causal knowledge and data from a vast amount of medical texts and automatically construct a causal relationship map for medical named entity recognition and relationship extraction tasks.
[0006] According to one aspect of the present invention, there is provided an intelligent disease diagnosis causal map construction method for medical texts, including:
[0007] Obtain data related to medical texts, including medical research literature, clinical reports, electronic medical records, online disease consultations, diagnostic and treatment suggestions, etc.; among them, unstructured data is stored for causal entity recognition, and structured data is stored for training;
[0008] For causal entity recognition, a named entity recognition method based on feature enhancement is adopted, that is, the medical text representation ability of the model is enhanced from two aspects of knowledge and structure. Specifically, it is the MedKE model with knowledge feature enhancement, which integrates external medical knowledge into the embedding layer of the model to supplement medical knowledge information for feature representation; the entity recognition model GCNSE with structure feature enhancement uses the structural features generated by GCN to enhance the medical entity structure information based on the BERT model;
[0009] For causal relationship extraction, a medical causal relationship extraction method based on entity embedding and path reasoning is used. That is, the EF-CRE model based on entity feature embedding introduces local features of entities by designing entity markers, and then combines the global features of sentences as the basis for text classification; according to the characteristics of causal relationships themselves, a causal strength quantification method based on trigger word similarity is designed, and a Multi-PCR extraction method based on multi-hop path causal reasoning is established to complement the long-distance document-level causal relationships;
[0010] For entity unification, an entity unification method combining edit distance and word vector similarity is adopted to solve the problem of entity inconsistency. By calculating the edit distance and word vector similarity, entities with the same meaning but different expressions are unified to reduce the redundancy of the knowledge graph.
[0011] Furthermore, the MedKE model with knowledge feature enhancement includes:
[0012] Adopt an entity recognition method with knowledge feature enhancement to supplement the professional knowledge information lacking in word vectors by introducing medical domain knowledge;
[0013] The MedKE model integrates medical domain knowledge into the embedding representation layer of the model. By retrieving external knowledge and encoding and classifying the input medical text in two ways, an embedded layer statement containing knowledge is generated to enhance the knowledge information of the text. Then, the model is fine-tuned on medical corpora to generate a feature vector representation with prior semantic knowledge and medical domain knowledge;
[0014] Denote the external medical knowledge source as MedK, and use MedK i to represent any one of the data in it, and its triple is represented as MedK i ={K i ,R i ,V i}, K i represents the head entity, V i represents the tail entity, Ri Represents the relationship between the head entity and the tail entity;
[0015] The input sentence Sentence is segmented into a word sequence through Chinese word segmentation, i.e., Sentence = {A0, A1, …, A n}, where n represents the length of the input sentence, and A i represents each word after Chinese word segmentation; when each Sentence is input into the knowledge retrieval encoding layer, the model will match and retrieve A i with K i in the external knowledge base. When the corresponding medical knowledge is retrieved, i.e., A i = K i , through the knowledge triple (K i , R i , V i ), the entity V i corresponding to the K i entity is retrieved, and the corresponding C i + V i word vector is generated; if the corresponding medical knowledge is not retrieved, the corresponding C i + V0 word vector is generated; E i represents the word vector generated by C i , as shown in formula (1):
[0016]
[0017] Through the knowledge query calculation of the knowledge retrieval encoding layer, the knowledge of the external medical database is injected into this Sentence to obtain the corresponding word vector: Embedding = {E0, E1, …, E n}. In the embedding layer of the model, the sequence vector is added as Token Embedding to Segment Embedding and Position Embedding, so as to fuse the medical knowledge fragments into the corpus text, and thus obtain the Embedding that contains both rich context information and rich medical knowledge;
[0018] The vector representation is input into the Transformer encoder layer, and the text feature representation corresponding to each token is generated through the calculation of the multi-head attention mechanism, so as to obtain the output vector of the last layer of the encoder, denoted as: H = {H1, H2, …, H n}. The obtained text feature representation H is input into formula (2) to obtain the corresponding probability distribution. p represents the probability of each token in the input sequence corresponding to each label. W is the weight matrix, which stores the weights between the connected neurons; b is the bias, which provides an adjustable threshold for each neuron to adjust its activation state;
[0019] p = softmax(W * H + b) (2)
[0020] Then, the predicted category of any character, i.e., the entity category y, is calculated through formula (3); then, the loss is calculated using the cross-entropy loss function through formula (4) and the gradient is backpropagated to update the parameters, where p is the true label value of a certain token and q is the predicted label value of that token;
[0021] Repeat multiple rounds of training to obtain the predicted scores of the labels corresponding to each character in the input sentence, and select the label with the highest probability as the final prediction result of the model. After the 'BIO' label conversion, the entity category to which the character belongs is obtained;
[0022] y = argmax(p) (3)
[0023]
[0024] Furthermore, the entity recognition model GCNSE with enhanced structural features includes:
[0025] To address the problem that the model lacks structural information between entities, a fusion model of BERT and GCN is adopted to supplement the lacking structural information in the word vectors using GCN;
[0026] The GCNSE model simultaneously utilizes the context semantic representation ability of the Bert model and the structural representation ability of the graph convolutional neural network GCN to improve the effect of medical entity recognition;
[0027] The model adopts the extraction method proposed in this patent to extract the causal knowledge in medical texts into several quadruples, and then forms a causal knowledge graph with these quadruples as the training dataset of the GCN model;
[0028] The constructed causal knowledge graph is represented by an adjacency matrix A: A(n, n) represents a graph structure with n nodes of medical entities; where the nodes of the adjacency matrix are extracted from the dataset through the entity recognition method, and each sentence in the dataset is still represented by Sentence = {C0, C1, …, C n}; the edges of the adjacency matrix are extracted by the causal relationship extraction model, and the weights of the adjacency matrix are determined by the causal intensity quantization method. Specifically, for the i-th node and the j-th node (i, j = 1, 2,..., n) in the graph, the element A[i][j] in the adjacency matrix A is represented as follows:
[0029] If there is an edge connecting vertex i and vertex j, then A[i][j] = p (causal intensity);
[0030] If there is no edge connecting vertex i and vertex j, then A[i][j] = 0;
[0031] The feature vector matrix of the medical entity is represented as X, which can be default-assigned as a zero matrix or an identity matrix, and its dimension is (n, F 0 ), where F 0 is the feature dimension of the input vector. To align with the word vectors of the BERT model, set the value of F 0 to 768;
[0032] The model first processes the adjacency matrix A to obtain a normalized adjacency matrix. Secondly, for each node, it aggregates the information of its neighbor nodes, combines the aggregated neighbor node features with the features of the target node itself, and then extracts the new feature representation of the node through a linear transformation and a non-linear activation function, as shown in formula (5): W represents the parameter matrix, H represents the features of each layer. For the input layer of the network, H is the input X vector;
[0033]
[0034] The model sets the GCN network to two layers. After multiple trainings, after the iterative calculation in the last round, the feature representation of each node is obtained: And the obtained feature representation is used as the structural feature vector of the medical entity; Inject Sentence into the BERT module, and denote the medical entity in it as C T Use the semantic vector in BERT as Then add the graph structure vector H G generated by GCN to the semantic vector H C generated by the BERT pre-trained language model to get a new vector H E , where m is an adjustable ratio parameter in formula (6). The larger the value of m, the higher the proportion of GCN. When the value of m is 0, it means that the current word vector only relies on the word vector representation of the BERT model;
[0035]
[0036] Input the sequence vector H E into Transformer to perform feature extraction through the calculation of the self-attention mechanism, so as to fuse the structural fragments into the original text data and obtain the structural information of the medical text, making up for the problem of insufficient structural information in the pre-trained language model; For each input vector perform a linear transformation, multiply by the corresponding weight matrix to obtain the corresponding query vector, key vector and value vector;
[0037]
[0038]
[0039]
[0040] Then, use the attention mechanism to calculate the features of words, pass the results calculated by the model to the loss function, measure the difference between the predicted value and the true value of the model by calculating the loss, and then use the backpropagation of the loss value to update the gradient of the model and optimize the performance of the model;
[0041] Finally, obtain the predicted scores of the labels corresponding to each character in the sentence, and finally select the maximum predicted score as the label of the character. After 'BIO' conversion, obtain the corresponding named entity;
[0042]
[0043]
[0044] multiHead(Q, K, V) = Concat(head i )W(12).
[0045] Furthermore, the EF-CRE model based on entity feature embedding includes:
[0046] Adopt a relation extraction model based on entity feature embedding, and classify by introducing entity feature basis and combining the features of the whole sentence;
[0047] In the relation extraction task, the input sequence contains not only a large amount of text information but also medical entities. However, the input sentence is not clearly labeled or distinguished, resulting in no highlighting of its impact on the causal relationship during calculation; the EF-CRE model marks these medical entities at the input layer, introduces the features corresponding to the entity marks during the model training stage, and sets a proportional parameter to adjust the weight of the medical entity in the calculation of the features of the whole sentence, calculates the joint loss function of the features of the whole sentence and the entity features, and uses this joint feature as the basis for relation classification to determine the relation type;
[0048] In the input stage, for a given input sequence Sentence = {W0, W1, T1,..., T2,..., W n}, and the medical entities T1 and T2 existing in this input sequence, assuming that T1 is before T2; denote P1 and P2 as the subscripts of the first characters of T1 and T2 respectively, that is, the starting subscripts, then obtain their ending subscripts:
[0049] Q i = P i + Len(T i ) (13)
[0050] Save the two medical entity subscripts {P1, Q1, P2, Q2} contained in each sentence as a quadruple and store it in the list Index for subsequent finding of the corresponding medical entities in the model training module; then input Sentence into BERTTokenizer to generate the corresponding Token Embedding, denoted as Emb = {[CLS], E1, E2, …, E n , [SEP]};
[0051] According to the start and end subscripts of each medical entity, embed the special marker $ into the Embedding generated from the original sentence to explicitly mark the entity. $ serves as a marker like [CLS], denoted by U. It has no semantics itself, but through the calculation of the multi-head attention mechanism, it will contain information of neighboring entities. The input sequence Emb after inserting the special token is {[CLS], E1, E2, [U P1 ,..., [U Q1 , …, [U P2 , …, [U Q2 , …, E n , [SEP]};
[0052] In the model fine-tuning training stage, during the fine-tuning stage, add medical entity information to the TokenEmbedding of the pre-trained model. The model calculates through multiple Transformer modules to obtain the preliminary feature representations of [CLS], E i and U i ; The [CLS] marker represents the global feature, which captures the semantic information of the entire sentence, while the U i marker focuses on the local features of the entities in the sentence and can indicate the specific positions of the entities in the sentence;
[0053] Fuse the global feature and the local feature in a certain proportion and concatenate them into a new feature vector; The vector corresponding to the main entity T1 of the medical causal relationship is calculated as H[U1], and the vector corresponding to the guest entity T2 is calculated as H[U2] as shown in formula (14):
[0054]
[0055] The vector H[CLS] corresponding to [CLS], and the dimensions of the three vectors are all the BERT standard vector dimension: 768 dimensions; m is an adjustable proportion parameter. The new vector representation of the <T1, T2> entity pair finally obtained is as shown in formula (15), briefly denoted as H;
[0056] H(T1, T2) = Concat(H[CLS] + m * (H[U1] + H[U2])) (15)
[0057] The obtained new feature vector H is passed into the joint loss function to compare the difference between the label probabilities given by the model and the probability distribution of the true labels, quantify the inconsistency between the predicted value and the actual value, and then use the backpropagation algorithm to update the parameters of the model to minimize the joint loss function, continuously adjusting the weights and biases of the model so that the relationship categories output by the model are closer and closer to the true labels.
[0058] Further, the method for quantifying causal strength based on trigger word similarity includes:
[0059] Create a vocabulary containing explicit causal relationship trigger words based on the Internet corpus word bank. This vocabulary contains a series of causal trigger words, and then through the MedKE pre-trained model, encode the trigger word list into corresponding word vectors;
[0060] Select an explicit causal trigger word W from the word list and denote its corresponding word vector as E. For any set of causal relationships, denote the head entity and the tail entity as c1 and c2 respectively, and find the sentence containing this pair of entities, denoted as S = {S1, S2, …, S n}, and then perform Chinese word segmentation on each sentence S i , denoted as S i = {W i1 , W i2 , …, W in}, and then encode the segmented words with the model and denote their word vectors as E i = {E i1 , E i2 , …, E in};
[0061] F ij represents the word vector similarity value between the j-th word in the i-th sentence and the causal trigger word. Denote a similarity threshold as m, where F ij is calculated as shown in formula (16):
[0062]
[0063] Finally, the causal strength score of c1 and c2 is F as shown in formula (17):
[0064]
[0065] After obtaining a preliminary causal strength F, amplify or reduce the causal strength according to other trigger words in the trigger word list. Appropriately amplify the weight for the connection that conforms to the strong causal relationship, and appropriately reduce the weight for the connection that conforms to the strong non-causal relationship. k is the scaling coefficient, and the final causal strength is obtained from formula (18);
[0066] F = k * F(c1, c2) (18).
[0067] Furthermore, the Multi-PCR extraction method based on multi-hop path causal reasoning includes:
[0068] After obtaining a number of causal knowledge quadruples through the causal knowledge extraction task, perform multi-hop path causal reasoning;
[0069] Denote Xp and Xq as two medical entities. There is no direct edge connection between them in the causal knowledge graph, but there is a multi-hop causal path connection between them. Denote Rij as the causal strength of the j-th edge of the i-th causal path. The value range of i is [1, n], where n is the number of reachable causal paths; the value range of j is [1, k], and k is the number of hops, indicating that Xq can be reached from Xp after m hops. Denote P(Xp→Xq) as the probability that there is a causal relationship between Xp and Xq, as shown in formula (19):
[0070]
[0071] If the probability P is greater than a certain threshold, it indicates that there is a certain causal relationship between Xp and Xq; otherwise, there is no causal relationship between Xp and Xq.
[0072] Furthermore, the entity unification method combining edit distance and word vector similarity includes:
[0073] Preprocess by calculating the edit distance of entity pairs and select the candidate entity pairs with a smaller number. Initialize the two-dimensional matrix dp, whose size is (m + 1)*(n + 1), where m and n are the lengths of the entities C1 and C2 to be compared respectively. The first row and the first column of the matrix are initialized to consecutive integers from 0 to m and n;
[0074] dp[0,j] =j (20)
[0075] dp[i, 0] =i (21)
[0076] Among them, dp[0][j]=j means that converting an empty string into the first j characters of entity C2 requires j deletion operations, and dp[i][0]=i means that converting the first i characters of entity C1 into an empty string requires i insertion operations. Starting from dp[1][1], fill the two-dimensional array row by row and column by column. The filling value is as shown in formula (22). Finally, the value of dp[m][n] is the minimum number of edit operations required to convert entity C1 into entity C2. The smaller the threshold of the edit distance, the more similar the two entities are;
[0077] dp[i][j]=min(dp[i-1][j]+1,dp[i][j-1]+1,dp[i-1][j-1]+1) (22)
[0078] Select entity pairs with an edit distance equal to 1, denoted as {U1, U2}, put them into the MedKE model, obtain their corresponding word vectors {E1, E2}, use the word vector similarity as a consideration for the similarity of the two entities, and calculate the similarity score through formula (23);
[0079] Score = cos(E1 * E2) (23)
[0080] For entity pairs with scores exceeding the threshold m, it is considered that they are semantically similar and need to be unified entities. Select the TF-IDF value as the basis for unifying entities. Determine whether the entity pair {U1, U2} is unified into entity U1 or entity U2 to obtain the final entity U;
[0081]
[0082] According to another aspect of the present invention, there is provided an intelligent disease diagnosis causal graph construction system for medical texts, including:
[0083] A module for obtaining medical text-related data, which is used to obtain data such as medical research literature, clinical reports, electronic medical records, online disease consultations, diagnosis and treatment suggestions, etc.; among them, unstructured data is stored for causal entity recognition, and structured data is stored for training;
[0084] A causal entity recognition module, which is used to adopt a named entity recognition method based on feature enhancement, that is, to enhance the medical text representation ability of the model from both knowledge and structure aspects. Specifically, it is a MedKE model with knowledge feature enhancement, which integrates external medical knowledge into the embedding layer of the model to supplement medical knowledge information for feature representation; a structure feature-enhanced entity recognition model GCNSE, which uses the structure features generated by GCN to enhance the medical entity structure information based on the BERT model;
[0085] A causal relationship extraction module, which is used for a medical causal relationship extraction method based on entity embedding and path reasoning, that is, an EF-CRE model based on entity feature embedding. By designing entity markers to introduce local features of entities, and then combining the global features of sentences as the basis for text classification; according to the characteristics of causal relationships themselves, design a causal intensity quantification method based on trigger word similarity, and establish a Multi-PCR extraction method based on multi-hop path causal reasoning to supplement long-distance document-level causal relationships;
[0086] An entity unification module, which is used to adopt an entity unification method that combines edit distance and word vector similarity to solve the problem of entity inconsistency. By calculating the edit distance and word vector similarity, entities with the same meaning but different expressions are unified to reduce the redundancy of the graph.
[0087] Advantages of the present invention:
[0088] In the entity recognition stage, the present invention enhances the features of the BERT model from two aspects of knowledge and structure, proposes a knowledge-enhanced MedKE model, introduces medical domain knowledge to supplement the knowledge information of word vectors; establishes a fusion model GCNSE, combines the BERT and GCN methods, enabling the model to capture both the context relationship and semantic information between words, and to learn the structural relationship and syntactic information between words, further improving the domain adaptation ability of the model. In addition, entity unification is performed by combining the edit distance and word vector similarity to reduce the redundancy of the knowledge graph. In the relation extraction stage, the present invention proposes an EF-CRE model based on entity feature embedding, classifies by introducing entity feature bases and combining full-sentence features, and establishes a Multi-PCR method to supplement the causal relationship between long-distance entities.
[0089] The method and system for constructing an intelligent disease diagnosis causal knowledge graph for medical texts of the present invention realize the models and methods for automatically constructing the proposed causal knowledge graph, and provide functions such as dynamic uncertain causal knowledge graph editing, clinical intelligent auxiliary diagnosis, and automatic extraction of electronic medical records. Through the multi-source fusion of medical text data and expert experience knowledge, a disease causal knowledge graph with certain clinical application value can be established.
[0090] In addition to the purposes, features, and advantages described above, the present invention has other purposes, features, and advantages. The following will refer to the drawings for a further detailed description of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0091] The drawings constituting a part of this application are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention.
[0092] Figure 1 is a flowchart of the method for constructing a disease diagnosis causal knowledge graph for medical texts of the present invention;
[0093] Figure 2 is a schematic diagram of the automatic construction process of the causal knowledge graph of the present invention;
[0094] Figure 3 is the interface of the causal knowledge base editing system of the present invention;
[0095] Figure 4 is the disease diagnosis causal sub-graph of the present invention;
[0096] Figure 5 is an example of the merged overall causal graph of the present invention;
[0097] Figure 6 is the automatically generated case library graph (partial) of the present invention;
[0098] Figure 7 This is an example diagram for automatically extracting disease cases of the present invention. Detailed implementation manners
[0099] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0100] Refer to Figures 1 to 7 , the present invention provides a method and system for constructing an intelligent disease diagnosis causal map for medical texts.
[0101] Embodiment 1
[0102] 1. A method for constructing a disease diagnosis causal map for medical texts, which extracts disease information descriptions, diagnosis paths and causal relationship explanations from a large amount of medical texts on the Internet to complete the construction task of a medical causal relationship map;
[0103] The method specifically includes the following steps:
[0104] (1) Obtain relevant data of medical texts
[0105] Obtain data such as medical research literature, clinical reports, electronic medical records, online disease consultations, diagnosis and treatment suggestions, etc.; among them, unstructured data is stored for causal entity recognition, and structured data is stored for training;
[0106] (2) Causal entity recognition
[0107] Adopt a named entity recognition method based on feature enhancement, that is, enhance the medical text representation ability of the model from two aspects of knowledge and structure. Specifically, it is the MedKE model (Medical Knowledge Enhanced model, MedKE) enhanced by knowledge features, which integrates external medical knowledge into the embedding layer of the model to supplement medical knowledge information for feature representation; the entity recognition model GCNSE (GCN Structure Enhanced model, GCNSE) enhanced by structural features, which uses the structural features generated by GCN to enhance the medical entity structure information based on the BERT model;
[0108] (3) Causal relationship extraction
[0109] Medical Causal Relationship Extraction Method Based on Entity Embedding and Path Reasoning, namely the EF-CRE model (Entity Feature embed Causal Relational Extraction, EF-CRE) based on entity feature embedding, introduces local features of entities by designing entity tags, and then combines the global features of sentences as the basis for text classification; according to the characteristics of causal relationships themselves, designs a causal intensity quantification method based on trigger word similarity, and establishes a Multi-PCR extraction method (Multi-hop Path Causal Reasoning, Multi-PCR) based on multi-hop path causal reasoning to complement long-distance document-level causal relationships;
[0110] (4) Entity Unification
[0111] Finally, an entity unification method combining edit distance and word vector similarity is adopted to solve the entity inconsistency problem. By calculating the edit distance and word vector similarity, entities with the same meaning but different expressions are unified, reducing the redundancy of the knowledge graph.
[0112] The MedKE model enhanced with knowledge features includes:
[0113] To address the problem that the model lacks medical knowledge unique to the medical field, an entity recognition method enhanced with knowledge features is adopted. By introducing medical domain knowledge, professional knowledge information lacking in word vectors is supplemented;
[0114] The MedKE model integrates medical domain knowledge into the embedding representation layer of the model. By retrieving external knowledge and encoding and classifying the input medical text in two ways, an embedded layer statement containing knowledge is generated to enhance the knowledge information of the text. Then, the model is fine-tuned on medical corpora to generate a feature vector representation with prior semantic knowledge and medical domain knowledge;
[0115] Denote the external medical knowledge source as MedK, and use MedK i to represent any piece of data in it, and its triple is represented as MedK i ={K i , R i , V i}, K i represents the head entity, V i represents the tail entity, and R i represents the relationship between the head entity and the tail entity;
[0116] The input sentence Sentence is segmented into a word sequence through Chinese word segmentation, that is, Sentence={A0, A1, …, A n}, where n represents the length of the input sentence, and A iRepresents each word after Chinese word segmentation; when each Sentence is input into the knowledge retrieval encoding layer, the model will compare A i with K in the external knowledge base i for matching retrieval. When the corresponding medical knowledge is retrieved, that is, A i =K i , through the knowledge triple (K i , R i , V i ), the entity V corresponding to the K i entity is retrieved, and the corresponding C i +V i word vector is generated; if the corresponding medical knowledge is not retrieved, the corresponding C i +V0 word vector is generated; E i represents the word vector generated by C i , as shown in formula (1): i
[0117]
[0118] Through the knowledge query calculation of the knowledge retrieval encoding layer, the knowledge of the external medical database can be injected into the Sentence to obtain the corresponding word vector: Embedding = {E0, E1,..., E n}. In the embedding layer of the model, the sequence vector is added to the Token Embedding, Segment Embedding, and Position Embedding, so as to fuse the medical knowledge fragments into the corpus text, and thus obtain the Embedding that contains both rich context information and rich medical knowledge;
[0119] The vector representation is input into the Transformer encoder layer, and the text feature representation corresponding to each token is generated through the calculation of the multi-head attention mechanism, so as to obtain the output vector of the last layer of the encoder, denoted as: H = {H1, H2,..., H n}. The obtained text feature representation H is input into formula (2) to obtain the corresponding probability distribution. p represents the probability of each label corresponding to each token in the input sequence. W is the weight matrix, which stores the weights between the connected neurons; b is the bias, which provides an adjustable threshold for each neuron to adjust its activation state;
[0120] p = softmax(W * H + b) (2)
[0121] Then, the predicted category of any character, that is, the entity category y, is calculated through formula (3); then, the loss is calculated using the cross-entropy loss function through formula (4) and the gradient is backpropagated to update the parameters, where p is the true label value of a certain token and q is the predicted label value of that token;
[0122] Repeat multiple rounds of training to obtain the predicted scores of the labels corresponding to each character in the input sentence, and select the label with the highest probability as the final prediction result of the model. After the 'BIO' label conversion, the entity category to which the character belongs is obtained.
[0123] y = argmax(p) (3)
[0124]
[0125] The entity recognition model GCNSE with enhanced structural features includes:
[0126] To address the problem that the model lacks structural information between entities, a fusion model of BERT and GCN is adopted to use GCN to supplement the lacking structural information in the word vectors;
[0127] The GCNSE model simultaneously utilizes the context semantic representation ability of the Bert model and the structural representation ability of the graph convolutional neural network GCN to improve the effect of medical entity recognition;
[0128] The model uses the extraction method proposed in this patent to extract the causal knowledge in the medical text into several quadruples, and then forms a causal knowledge graph with these quadruples as the training dataset of the GCN model;
[0129] The constructed causal knowledge graph is represented by an adjacency matrix A: A(n,n) represents a graph structure with n medical entities as nodes; among them, the nodes of the adjacency matrix are extracted from the dataset through the entity recognition method, and each sentence in the dataset is still represented by Sentence = {C0, C1, …, C n}; the edges of the adjacency matrix are extracted by the causal relationship extraction model, and the weights of the adjacency matrix are determined by the causal intensity quantization method. Specifically, for the i-th node and the j-th node (i, j = 1, 2,..., n) in the graph, the element A[i][j] in the adjacency matrix A is represented as follows:
[0130] (1) If there is an edge connecting vertex i and vertex j, then A[i][j] = p (causal intensity);
[0131] (2) If there is no edge connecting vertex i and vertex j, then A[i][j] = 0;
[0132] Represent the feature vector matrix of medical entities as X, which can be defaultly assigned as a zero matrix or an identity matrix, and its dimension is (n, F 0 ), where F 0 is the feature dimension of the input vector. To align with the word vectors of the BERT model, set the value of F 0 to 768;
[0133] The model first processes the adjacency matrix A to obtain the normalized adjacency matrix. Secondly, for each node, it aggregates the information of its neighbor nodes, combines the aggregated neighbor node features with the features of the target node itself, and then extracts the new feature representation of the node through a linear transformation and a non-linear activation function, as shown in formula (5): W represents the parameter matrix, H represents the features of each layer. For the input layer of the network, H is the input X vector;
[0134]
[0135] The model sets the network of GCN to two layers. After multiple trainings, after the iterative calculation of the last round, it obtains the feature representation of each node: And regard the obtained feature representation as the structural feature vector of medical entities; Inject Sentence into the BERT module, and denote the medical entity in it as C T Use the semantic vector in BERT as Then add the graph structure vector H G generated by GCN to the semantic vector H C generated by the BERT pre-trained language model to get a new vector H E , where m is an adjustable ratio parameter in formula (6). The larger the value of m, the higher the proportion of GCN. When the value of m is 0, it means that the current word vector only relies on the word vector representation of the BERT model;
[0136]
[0137] Input the sequence vector H E into Transformer to perform feature extraction through the calculation of the self-attention mechanism, so as to fuse the structural fragments into the original text data and obtain the structural information of medical texts, making up for the problem of insufficient structural information in the pre-trained language model; For each input vector perform a linear transformation, multiply by the corresponding weight matrix to obtain the corresponding query vector, key vector and value vector;
[0138]
[0139]
[0140]
[0141] Then, use the attention mechanism to calculate the features of words, pass the results calculated by the model to the loss function, measure the difference between the model prediction value and the true value by calculating the loss, and then use the backpropagation of the loss value to update the gradient of the model and optimize the performance of the model;
[0142] Finally, obtain the prediction scores of the labels corresponding to each character in the sentence, and finally select the largest prediction score as the label of the character. After 'BIO' conversion, obtain the corresponding named entity.
[0143]
[0144]
[0145] multiHead(Q,K,V) = Concat(head i )W (12)
[0146] The EF-CRE model based on entity feature embedding includes:
[0147] Aiming at the limitations brought by the pipeline approach strategy in text classification, a relation extraction model based on entity feature embedding is adopted, and classification is carried out by introducing entity feature basis and combining the features of the whole sentence;
[0148] In the relation extraction task, the input sequence contains not only a large amount of text information but also medical entities, but the input sentences are not clearly marked or distinguished, resulting in no highlighting of their impact on the causal relationship during calculation; the EF-CRE model marks these medical entities at the input layer, introduces the features corresponding to the entity marks during the model training stage, and sets a proportional parameter to adjust the weight of the medical entities in the calculation of the features of the whole sentence, calculates the joint loss function of the features of the whole sentence and the entity features, and uses this joint feature as the basis for relation classification to determine the relation type;
[0149] (1) Input stage
[0150] For a given input sequence Sentence = {W0, W1, T1, …, T2, …, W n}, and the medical entities T1 and T2 existing in this input sequence, assuming that T1 is before T2; denote P1 and P2 as the subscripts of the first characters of T1 and T2 respectively, that is, the starting subscripts, then obtain their ending subscripts:
[0151] Q i = P i + Len(T i ) (13)
[0152] Save the two medical entity subscripts {P1, Q1, P2, Q2} contained in each sentence as a quadruple and store it in the list Index for subsequent finding of the corresponding medical entities in the model training module; then input Sentence into BERTTokenizer to generate the corresponding Token Embedding, denoted as Emb = {[CLS], E1, E2, …, E n , [SEP]};
[0153] According to the start and end subscripts of each medical entity, embed the special marker "$" into the Embedding generated from the original sentence to explicitly mark the entity. "$" serves as a marker like [CLS], denoted as U. It has no semantics itself, but through the calculation of the multi-head attention mechanism, it will contain information of neighboring entities. The input sequence Emb after inserting the special token is {[CLS], E1, E2, [U P1 ,..., [U Q1 , …, [U P2 , …, [U Q2 , …, E n , [SEP]};
[0154] (2) Model fine-tuning training stage
[0155] In the fine-tuning stage, add medical entity information to the Token Embedding of the pre-trained model. The model calculates through multiple Transformer modules to obtain the preliminary feature representations of [CLS], E i and U i ; The [CLS] marker represents the global feature, which captures the semantic information of the entire sentence, while the U i marker focuses on the local features of the entities in the sentence, which can indicate the specific positions of the entities in the sentence;
[0156] Fuse the global feature and the local feature in a certain proportion and concatenate them into a new feature vector; The vector corresponding to the main entity T1 of the medical causal relationship is H[U1], and the vector corresponding to the guest entity T2 is calculated as in formula (14):
[0157]
[0158] The vector H[CLS] corresponding to [CLS], and the dimensions of the three vectors are all the BERT standard vector dimension: 768 dimensions; m is an adjustable proportion parameter. The new vector representation of the <T1, T2> entity pair finally obtained is as in formula (15), simply denoted as H;
[0159] H(T1,T2) = Concat(H[CLS] + m * (H[U1] + H[U2])) (15)
[0160] The obtained new feature vector H is input into the joint loss function to compare the difference between the label probability given by the model and the probability distribution of the true label, quantify the inconsistency between the predicted value and the actual value, and then use the backpropagation algorithm to update the parameters of the model to minimize the joint loss function, continuously adjusting the weights and biases of the model so that the relationship category output by the model is closer and closer to the true label.
[0161] The method for quantifying the causal strength based on the trigger word similarity includes:
[0162] Create a vocabulary containing explicit causal relationship trigger words based on the Internet corpus word bank. This vocabulary will contain a series of causal trigger words, and then through the MedKE pre-trained model, encode the trigger word list into corresponding word vectors;
[0163] Select an explicit causal trigger word W from the word list and denote its corresponding word vector as E. For any set of causal relationships, denote the head entity and the tail entity as c1 and c2 respectively, and find the sentences containing this pair of entities, denoted as S = {S1, S2, …, S n}, and then perform Chinese word segmentation on each sentence S i and denote it as S i = {W i1 , W i2 , …, W in}, and then encode the segmented words with the model and denote their word vectors as E i = {E i1 , E i2 , …, E in};
[0164] F ij represents the word vector similarity value between the j-th word in the i-th sentence and the causal trigger word. Denote a similarity threshold as m, where F ij is calculated as shown in formula (16):
[0165]
[0166] Finally, the causal strength score of c1 and c2 is F as shown in formula (17):
[0167]
[0168] After obtaining a preliminary causal strength F, the causal strength is amplified or reduced according to other trigger words in the trigger word list. The connections that conform to strong causal relationships are appropriately weighted up, and the connections that conform to strong non-causal relationships are appropriately weighted down. k is the scaling coefficient, and the final causal strength is obtained from formula (18).
[0169] F = k * F(c1, c2) (18)
[0170] The Multi-PCR extraction method based on multi-hop path causal reasoning includes:
[0171] After obtaining a number of causal knowledge quadruples <head entity, tail entity, causal relationship, causal strength> through the causal knowledge extraction task, multi-hop path causal reasoning is carried out;
[0172] Let Xp and Xq be two medical entities. There is no direct edge connection between them in the causal knowledge graph, but there is a multi-hop causal path connection between them. Let Rij be the causal strength of the j-th edge of the i-th causal path. The value range of i is [1, n], where n is the number of reachable causal paths; the value range of j is [1, k], where k is the number of hops, indicating that Xq can be reached from Xp after m hops. Let P(Xp→Xq) be the probability that there is a causal relationship between Xp and Xq, as shown in formula (19):
[0173]
[0174] If the probability P is greater than a certain threshold, it indicates that there is a certain causal relationship between Xp and Xq; otherwise, there is no causal relationship between Xp and Xq.
[0175] The entity unification method combining edit distance and word vector similarity includes:
[0176] Calculate the edit distance of the entity pair for preprocessing, and select the candidate entity pairs with fewer numbers; initialize the two-dimensional matrix dp, whose size is (m + 1) * (n + 1), where m and n are the lengths of the entity C1 and entity C2 to be compared respectively. The first row and the first column of the matrix are initialized to consecutive integers from 0 to m and n;
[0177] dp[0, j] = j (20)
[0178] dp[i, 0] = i (21)
[0179] Among them, dp[0][j] = j indicates that converting an empty string to the first j characters of entity C2 requires j deletion operations, and dp[i][0] = i indicates that converting the first i characters of entity C1 to an empty string requires i insertion operations; starting from dp[1][1], the two-dimensional array is filled row by row and column by column, and the filled value is as shown in formula (22). Finally, the value of dp[m][n] is the minimum number of edit operations required to convert entity C1 to entity C2. The smaller the threshold of the edit distance, the more similar the two entities are;
[0180] dp[i][j] = min(dp[i - 1][j] + 1, dp[i][j - 1] + 1, dp[i - 1][j - 1] + 1) (22)
[0181] Select entity pairs with an edit distance equal to 1, denoted as {U1, U2}, put them into the MedKE model, obtain their corresponding word vectors {E1, E2}, select the word vector similarity as the consideration of the similarity between two entities, and calculate the similarity score through formula (23);
[0182] Score = cos(E1 * E2) (23)
[0183] For entity pairs with a score exceeding the threshold m, it is considered that they are semantically similar and need to be unified entities. Select the TF-IDF value as the basis for unifying entities. By judging whether the entity pair {U1, U2} is unified into entity U1 or entity U2, the final entity U can be obtained.
[0184]
[0185] Embodiment 2
[0186] An intelligent disease diagnosis causal graph construction system for medical texts, comprising:
[0187] A module for obtaining medical text-related data, which is used to obtain data such as medical research literature, clinical reports, electronic medical records, online disease consultations, diagnosis and treatment suggestions, etc.; among them, unstructured data is stored for causal entity recognition, and structured data is stored for training;
[0188] A causal entity recognition module, which is used to adopt a named entity recognition method based on feature enhancement, that is, to enhance the medical text representation ability of the model from both knowledge and structure aspects. Specifically, it is a MedKE model enhanced by knowledge features, which integrates external medical knowledge into the embedding layer of the model to supplement medical knowledge information for feature representation; an entity recognition model GCNSE enhanced by structural features, which uses the structural features generated by GCN to enhance the medical entity structure information based on the BERT model;
[0189] The causal relationship extraction module is used for the medical causal relationship extraction method based on entity embedding and path reasoning, that is, the EF-CRE model based on entity feature embedding, which introduces the local features of entities by designing entity tags, and then combines the global features of sentences as the basis for text classification; according to the characteristics of the causal relationship itself, a causal strength quantification method based on the similarity of trigger words is designed, and a Multi-PCR extraction method based on multi-hop path causal reasoning is established to complete the long-distance document-level causal relationship;
[0190] The entity unification module is used to solve the problem of entity inconsistency by using an entity unification method that combines edit distance and word vector similarity. By calculating edit distance and word vector similarity, entities with the same meaning but different expressions are unified to reduce graph redundancy.
[0191] The medical text-oriented intelligent disease diagnosis causal graph construction system of the present invention can construct and edit a causal knowledge base;
[0192] Causal knowledge base construction realizes the automated construction technology of disease diagnosis causal knowledge graph;
[0193] The causal knowledge base editor establishes a disease diagnosis causal graph editing system for medical experts to edit and modify disease diagnosis causal knowledge graphs, as well as automatically extract and manage cases.
[0194] The construction of the causal knowledge base includes: constructing a disease diagnosis causal knowledge graph composed of the four-tuple <head entity, tail entity, causal relationship, causal strength> through the proposed medical named entity recognition, medical causal relationship extraction, medical causal strength quantification and medical entity unification method, and then normalizing it into a data structure supported by the system to automatically establish a disease diagnosis causal knowledge graph.
[0195] Causal knowledge base editing includes: domain experts can add, modify and delete variables, and edit the causal connections and logical relationships between variables; disease-related variables are divided into subgraphs through the International Classification of Diseases codes, each subgraph represents the causal mechanism of a specific disease, and supports the management, replication and consistency checking of subgraphs to ensure the consistency of causal connections; multiple subgraphs can be merged to generate a diagnostic causal graph for a large category of diseases; in addition, entity variables and corresponding status values in the consultation text information can be extracted and stored in the database in batches, automatically matching the established disease diagnosis causal knowledge graph to assist in the implementation of disease testing for diagnostic reasoning.
[0196] Based on the visual development platform Unity, the present invention realizes the automated construction technology of the proposed disease diagnosis causal knowledge graph. Based on medical texts related to diseases, through the proposed medical named entity recognition, medical causal relationship extraction, medical causal intensity quantification, and medical entity unification methods, a disease diagnosis causal knowledge graph composed of quadruples <head entity, tail entity, causal relationship, causal intensity> is constructed, and then it is normalized into the data structure supported by the system to automatically establish a disease diagnosis causal knowledge graph.
[0197] A disease diagnosis causal graph editing system is established for medical experts to edit and modify the disease causal knowledge graph, as well as automatically extract and manage cases, providing assistance for clinical applications. The specific functions are as follows:
[0198] Graph editing and management:
[0199] Domain experts can add, modify, and delete variables, and edit the causal connections and logical relationships between variables; through the International Classification of Diseases coding, disease-related variables are divided into subgraphs, each subgraph representing the causal mechanism of a specific disease, and at the same time supporting the management, copying, and consistency checking of subgraphs to ensure the consistency of causal connections; multiple subgraphs can be merged to generate a general diagnosis causal graph of a major category of diseases;
[0200] Structured extraction of cases:
[0201] During the automatic construction of the disease causal knowledge graph, based on the corresponding relationship between the identified entity variables and the text, the editing system can further automatically extract disease cases from the text, extract the entity variables and corresponding status values in the interrogation text information, batch store them in the database, and automatically match the established disease diagnosis causal knowledge graph to assist in realizing the disease test of diagnostic reasoning; the data in the database represents the values of the extracted cases in the corresponding status of each variable, while the diagnostic causal graph in the editing system shows the clinical observation evidence of the case and the corresponding causal relationship.
[0202] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for constructing an intelligent disease diagnosis causal graph for medical text, characterized in that: include: Obtain medical text-related data, including medical research literature, clinical reports, electronic medical records, online disease consultations, diagnosis and treatment recommendations, etc. The unstructured data is stored for causal entity recognition, and the structured data is stored for training. Causal entity recognition uses a named entity recognition method based on feature enhancement, which enhances the model's medical text representation ability through knowledge and structure. Specifically, the MedKE model with knowledge feature enhancement integrates external medical knowledge into the model's embedding layer to supplement feature representation with medical knowledge information. The structural feature enhanced entity recognition model GCNSE uses the structural features generated by GCN to enhance the medical entity structure information based on the BERT model; Causal relationship extraction, a medical causal relationship extraction method based on entity embedding and path reasoning, namely the EF-CRE model based on entity feature embedding, which introduces the local features of entities by designing entity tags, and then combines the global features of sentences as the basis for text classification; based on the characteristics of causal relationships themselves, a causal strength quantification method based on trigger word similarity is designed, and a Multi-PCR extraction method based on multi-hop path causal reasoning is established to complete long-distance document-level causal relationships; Entity unification, adopts the entity unification method that combines edit distance and word vector similarity to solve the problem of entity inconsistency. Through the calculation of edit distance and word vector similarity, entities with the same meaning but different expressions are unified to reduce graph redundancy.
2. The method for constructing an intelligent disease diagnosis causal graph for medical text according to claim 1, characterized in that: The knowledge feature enhanced MedKE model includes: Adopting the entity recognition method enhanced by knowledge features, by introducing medical domain knowledge, supplementing the lack of professional knowledge information in word vectors; The MedKE model integrates medical knowledge into the model's embedding representation layer. By retrieving external knowledge and encoding and classifying the input medical text, it generates embedding layer sentences containing knowledge to enhance the knowledge information of the text. The model is then fine-tuned on medical corpus to generate feature vector representations with prior semantic knowledge and medical domain knowledge. The external medical knowledge source is denoted as MedK, and MedK i Represents any piece of data, and its triplet is represented by MedK i ={K i ,R i ,V i },K i Represents the head entity, V i Represents the tail entity, R i Represents the relationship between the head entity and the tail entity; The input sentence Sentence is segmented into a word sequence, that is, Sentence = {A0, A1, ..., A n }, where n represents the length of the input sentence, A i Represents each word after Chinese word segmentation; when each Sentence is input into the knowledge retrieval encoding layer, the model will convert A i K with external knowledge base i Perform matching search, and when the corresponding medical knowledge is retrieved, that is, A i =K i When the knowledge triple (K i ,R i ,V i ), retrieved the i V corresponding to the entity i Entity, generate the corresponding C i +V i Word vector; if the corresponding medical knowledge is not retrieved, generate the corresponding C i +V0 word vector; E i Represents C i The generated word vector is shown in formula (1): After the knowledge query calculation in the knowledge retrieval encoding layer, the knowledge of the external medical database is injected into the Sentence to obtain the corresponding word vector: Embedding = {E0, E1, ..., E n }, in the embedding layer of the model, the sequence vector is added as a token embedding to the segment embedding and position embedding, so as to integrate the medical knowledge fragment into the corpus text, thereby obtaining an embedding that contains both rich contextual information and rich medical knowledge; The vector representation is input into the Transformer encoder layer, and the text feature representation corresponding to each token is generated through the calculation of the multi-head attention mechanism, so as to obtain the output vector of the last layer of the encoder, which is denoted as: H = {H1, H2, ..., H n }, pass the obtained text feature representation H into formula (2) to obtain the corresponding probability distribution, p represents the probability of each label corresponding to each token in the input sequence, W is the weight matrix, which stores the weights between connected neurons; b is the bias, which provides an adjustable threshold for each neuron to adjust its activation state; p=softmax(W*H+b) (2) Then, the predicted category of any character, i.e., the entity category y, is calculated by formula (3). Then, the loss is calculated by using the cross entropy loss function and the gradient is back-propagated to update the parameters, where p is the real label value of a token and q is the predicted label value of the token. Repeat multiple rounds of training to obtain the predicted score of the label corresponding to each character in the input sentence, and select the label with the highest probability as the final prediction result of the model. After the 'BIO' label conversion, the entity category to which the character belongs is obtained; y=argmax(p) (3) 3. The method for constructing an intelligent disease diagnosis causal graph for medical text according to claim 1, characterized in that: The structural feature enhanced entity recognition model GCNSE includes: To address the problem that the model lacks structural information between entities, a fusion model of BERT and GCN is adopted, using GCN to supplement the lack of structural information in word vectors; The GCNSE model uses both the contextual semantic representation capability of the Bert model and the structural representation capability of the graph convolutional neural network (GCN) to improve the effect of medical entity recognition. The model uses the extraction method proposed in this patent to extract the causal knowledge in the medical text into several quadruples, and then the quadruples form a causal knowledge graph as the training data set of the GCN model; The constructed causal knowledge graph is represented by an adjacency matrix A: A(n,n) represents a graph structure with n medical entities as nodes; the nodes of the adjacency matrix are extracted from the dataset through entity recognition method, and each sentence in the dataset is still represented by Sentence = {C0, C1, ..., C n }; the edges of the adjacency matrix are extracted by the causal relationship extraction model, and the weights of the adjacency matrix are determined by the causal strength quantification method. Specifically, for the i-th node and the j-th node in the graph (i, j = 1, 2, ..., n), the element A[i][j] in the adjacency matrix A is expressed as follows: If there is an edge connecting vertex i and vertex j, then A[i][j] = p (causal strength); If there is no edge connecting vertex i and vertex j, then A[i][j] = 0; The eigenvector matrix of the medical entity is represented as X, which can be assigned a zero matrix or an identity matrix by default, and its dimension is (n,F 0 ), F 0 is the feature dimension of the input vector. In order to align with the word vector of the BERT model, F 0 The value of is set to 768; The model first processes the adjacency matrix A to obtain a normalized adjacency matrix. Then, for each node, it aggregates the information of its neighboring nodes, combines the aggregated neighboring node features with the target node’s own features, and then extracts the new feature representation of the node through linear transformation and nonlinear activation function, as shown in formula (5): W represents the parameter matrix, H represents the features of each layer, and for the input layer of the network, H is the input X vector; The model sets the GCN network to two layers. After multiple trainings, after the last round of iterative calculations, the feature representation of each node is obtained: H G ={H1 G ,H2 G ,...,H n G }, and use the obtained feature representation as the structural feature vector of the medical entity; inject Sentence into the BERT module, and record the medical entity in it as C T The semantic vector used in BERT is H C ={H1 C ,H2 C ,...,H n C }, and then the graph structure vector H generated by GCN G The semantic vector H generated by the BERT pre-trained language model C Add them together to get the new vector H E In formula (6), m is an adjustable ratio parameter. The larger the value of m, the higher the proportion of GCN. When the value of m is 0, it means that the word vector currently used only relies on the word vector representation of the BERT model. The sequence vector H E The input is sent to Transformer for feature extraction through the calculation of the self-attention mechanism, so as to integrate the structural fragments into the original text data, obtain the structural information of the medical text, and make up for the problem of insufficient structural information of the pre-trained language model; for each input vector Perform a linear transformation and multiply by the corresponding weight matrix to obtain the corresponding query vector, key vector and value vector; Then use the attention mechanism to calculate the features of the words, pass the model calculation results to the loss function, calculate the loss to measure the difference between the model prediction value and the true value, and then use the back propagation of the loss value to update the model gradient and optimize the model performance; Finally, the predicted score of the label corresponding to each character in the sentence is obtained, and the largest predicted score is selected as the label of the word. After the 'BIO' conversion, the corresponding named entity is obtained; head i =Attention(QW i Q ,KW i K ,VW i V ) (11) multiHead(Q,K,V)=Concat(head i )W (12)。 4. The method for constructing an intelligent disease diagnosis causal graph for medical text according to claim 1, characterized in that: The EF-CRE model based on entity feature embedding includes: A relation extraction model based on entity feature embedding is adopted to classify by introducing entity features and combining them with the features of the entire sentence; In the relationship extraction task, the input sequence contains not only a large amount of text information but also medical entities. However, the input sentences are not clearly marked or distinguished, resulting in the lack of highlighting of their impact on causal relationships during calculation. The EF-CRE model marks these medical entities at the input layer, introduces the features corresponding to the entity tags during the model training phase, and sets the ratio parameter to adjust the weight of the medical entity on the calculation of the entire sentence features, calculates the joint loss function of the entire sentence features and the entity features, and uses the joint features as the basis for relationship classification to determine the relationship type. In the input stage, for a given input sequence Sentence = {W0, W1, T1, ..., T2, ..., W n }, and the medical entities T1 and T2 existing in the input sequence, assuming that T1 is located before T2; let P1 and P2 be the subscripts of the first word of T1 and T2 respectively, that is, the starting subscript, then get its ending subscript: Q i =P i +Len(T i ) (13) The two medical entity subscripts {P1, Q1, P2, Q2} contained in each sentence are saved as a four-tuple and stored in the list Index, so that the corresponding medical entity can be found in the model training module later; then the Sentence is input into the BERT Tokenizer to generate the corresponding Token Embedding, which is recorded as Emb = {[CLS], E1, E2, …, E n ,[SEP]}; According to the start and end subscripts of each medical entity, the special token $ is embedded into the Embedding generated corresponding to the original sentence to explicitly mark the entity. $ is used as a marker like [CLS] and is represented by U. It has no semantics itself, but through the calculation of the multi-head attention mechanism, it will contain the information of neighboring entities. The input sequence after inserting the special token is Emb = {[CLS], E1, E2, [U P1 ],...,[U Q1 ],…,[U P2 ],…,[U Q2 ],…,E n ,[SEP]}; In the fine-tuning training phase, medical entity information is added to the TokenEmbedding of the pre-trained model. The model is calculated through a multi-layer Transformer module to obtain [CLS], E i and U i Preliminary feature representation; [CLS] marks the global features, which captures the semantic information of the entire sentence, while U i Labeling focuses on the local features of entities in a sentence, which can point out the specific location of the entity in the sentence; The global features and local features are combined in a certain proportion to form a new feature vector; the vector corresponding to the main entity T1 of the medical causal relationship is H[U1], and the vector corresponding to the object entity T2 is H[U2], which is calculated as shown in formula (14): [CLS] corresponds to the vector H[CLS], the dimensions of the three vectors are all BERT standard vector dimensions: 768 dimensions; m is an adjustable scale parameter, and the final<T1,T2> The new vector representation of the entity pair is as shown in formula (15), abbreviated as H; H(T1,T2)=Concat(H[CLS]+m*(H[U1]+H[U2])) (15) The new feature vector H is passed into the joint loss function, and the difference between the label probability given by the model and the probability distribution of the true label is compared to quantify the inconsistency between the predicted value and the actual value. The back propagation algorithm is then used to update the model parameters to minimize the joint loss function, and the model weights and biases are continuously adjusted to make the relationship categories output by the model closer and closer to the true labels.
5. The method for constructing an intelligent disease diagnosis causal graph for medical text according to claim 1, characterized in that: The causal strength quantification method based on trigger word similarity includes: Create a vocabulary containing explicit causal trigger words based on the Internet corpus. This vocabulary contains a series of causal trigger words. Then, use the MedKE pre-trained model to encode the trigger word vocabulary into corresponding word vectors. Select an explicit causal trigger word W from the vocabulary, and record its corresponding word vector as E. For any set of causal relationships, record the head entity and the tail entity as c1 and c2 respectively, and find the sentence containing this pair of entities, recorded as S = {S1, S2, ..., S n }, and then for each sentence S i As Chinese word segmentation, denoted as S i = {W i1 ,W i2 ,…,W in }, then encode the classified words using the model, and record their word vectors as E i ={E i1 ,E i2 ,…,E in }; F ij represents the similarity value between the jth word in the ith sentence and the word vector of the causal trigger, and a similarity threshold is m, where F ij The calculation method is as follows: Finally, the causal strength score of c1 and c2 is F as shown in formula (17): After obtaining a preliminary causal strength F, the causal strength is enlarged or reduced according to other trigger words in the trigger word list. The weight of the connection that meets the strong causal relationship is appropriately enlarged, and the weight of the connection that meets the strong non-causal relationship is appropriately reduced. k is the enlargement and reduction coefficient. The final causal strength is obtained by formula (18); F=k*F(c1,c2)(18).
6. The method for constructing an intelligent disease diagnosis causal graph for medical text according to claim 1, characterized in that: The Multi-PCR extraction method based on multi-hop path causal reasoning includes: After obtaining several causal knowledge quadruplets through the causal knowledge extraction task, causal reasoning of multi-hop paths is performed; Xp and Xq are two medical entities. There is no direct edge connecting them in the causal knowledge graph, but there is a multi-hop causal path connecting them. Rij is the causal strength of the jth edge of the i-th causal path. The value range of i is [1, n], and n is the number of reachable causal paths. The value range of j is [1, k], and k is the number of hops, indicating that Xp can reach Xq after m hops. P(Xp→Xq) is the probability that there is a causal relationship between Xp and Xq, as shown in formula (19): If the probability P is greater than a certain threshold, it means that there is a certain causal relationship between Xp and Xq, otherwise, there is no causal relationship between Xp and Xq.
7. The method for constructing an intelligent disease diagnosis causal graph for medical text according to claim 1, characterized in that: The entity unification method combining edit distance and word vector similarity includes: Calculate the entity pair edit distance for preprocessing and select a smaller number of candidate entity pairs; initialize the two-dimensional matrix dp, whose size is (m+1)*(n+1), where m and n are the lengths of the entities C1 and C2 to be compared, respectively, and the first row and first column of the matrix are initialized to consecutive integers from 0 to m and n; dp[0,j] = j (20) dp[i, 0] = i (21) Where dp[0][j]=j means that it takes j deletion operations to convert an empty string into the first j characters of entity C2, and dp[i][0]=i means that it takes i insertion operations to convert the first i characters of entity C1 into an empty string. Starting from dp[1][1], fill the two-dimensional array row by row and column by column, and the filling value is as shown in formula (22). Finally, the value of dp[m][n] is the minimum number of editing operations required to convert entity C1 into entity C2. The smaller the threshold of the edit distance, the more similar the two entities are. dp[i][j]=min(dp[i-1][j]+1,dp[i][j-1]+1,dp[i-1][j-1]+1) (22) Select an entity pair with an edit distance equal to 1, denoted as {U1, U2}, put it into the MedKE model, obtain its corresponding word vector {E1, E2}, use the word vector similarity as the consideration of the similarity between the two entities, and calculate the similarity score through formula (23); Score=cos(E1*E2) (23) For entity pairs whose scores exceed the threshold m, they are considered to be semantically similar entities that need to be unified. The TF-IDF value is selected as the basis for unifying the entities. It is determined whether the entity pair {U1, U2} is unified into entity U1 or entity U2, and the final entity U can be obtained.
8. An intelligent disease diagnosis causal graph construction system for medical texts, characterized in that: include: The module for obtaining medical text-related data is used to obtain data such as medical research literature, clinical reports, electronic medical records, online disease consultations, diagnosis and treatment recommendations; The unstructured data is stored for causal entity recognition, and the structured data is stored for training. The causal entity recognition module is used to adopt a named entity recognition method based on feature enhancement, that is, to enhance the model's medical text representation ability through knowledge and structure. Specifically, the knowledge feature enhanced MedKE model integrates external medical knowledge into the embedding layer of the model to supplement the feature representation with medical knowledge information; The structural feature enhanced entity recognition model GCNSE uses the structural features generated by GCN to enhance the medical entity structure information based on the BERT model; The causal relationship extraction module is used for the medical causal relationship extraction method based on entity embedding and path reasoning, that is, the EF-CRE model based on entity feature embedding, which introduces the local features of entities by designing entity tags, and then combines the global features of sentences as the basis for text classification; according to the characteristics of the causal relationship itself, a causal strength quantification method based on the similarity of trigger words is designed, and a Multi-PCR extraction method based on multi-hop path causal reasoning is established to complete the long-distance document-level causal relationship; The entity unification module is used to solve the problem of entity inconsistency by using an entity unification method that combines edit distance and word vector similarity. By calculating edit distance and word vector similarity, entities with the same meaning but different expressions are unified to reduce graph redundancy.
Citation Information
Cited By
Clinical research data analysis method based on machine learning
CN120409633A
Medical intelligent decision-making method based on Deepseek and time sequence causal knowledge graph
CN120636780A
Knowledge graph optimization method and system
CN120745782A
Disease and pest causal relationship identification method and system and disease and pest causal map construction method and system
CN121052349A
Geological metallogenic causal knowledge extraction method, storage medium, equipment and product
CN121301895A