Disease identification prediction method based on social media knowledge graph embedding
By constructing a user self-report knowledge graph in social media and using a combination of hyperbolic spatial embedding optimization and BERT+LSTM model, the problem of insufficient performance in processing hierarchical data in social media is solved, and a more efficient disease recognition prediction effect is achieved.
Patent Information
- Application Number
- CN202510067754.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-16
AI Technical Summary
Existing disease identification models are difficult to effectively identify patients with specific disease in social media, and traditional knowledge graph embedding methods are underperforming in processing hierarchical data.
A disease recognition prediction method based on social media knowledge graph embedding is proposed. By constructing a self-reported knowledge graph of social media users, and using the ComplEx model in hyperbolic space for knowledge embedding optimization, combining the BERT+LSTM model and the channel-space attention mechanism for knowledge fusion to achieve disease recognition prediction.
It significantly improves the ability to identify and predict patients with specific diseases in social media, enhances the ability to process hierarchical data, and improves the accuracy and recall of disease identification prediction.
Smart Images

Figure CN119993533A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing technology, and in particular to a disease recognition and prediction method based on social media knowledge graph embedding Background Art
[0002] With the widespread popularity of social media platforms, many patients record their disease symptoms on social media to seek help, providing researchers and public health departments with an effective channel to obtain data. These social media data contain a large amount of medical knowledge, disease information, patient feedback and related descriptions. However, since social media data is usually unstructured and contains a lot of noise and uncertainty, how to effectively extract valuable health information from these massive data and perform disease identification and prediction has become a challenge that needs to be solved.
[0003] In recent years, knowledge graphs, as a graphical method for expressing complex information, have been widely used in intelligent systems in various fields. Knowledge graphs represent things and their relationships through the relationship between nodes and edges. In the medical field, knowledge graphs can help integrate and associate relevant information such as diseases, symptoms, and treatment plans. Knowledge graph embedding technology can transform entities and relationships in the graph into dense representations in vector space, so that different types of information can be effectively compared and calculated in vector space. Therefore, by embedding relevant information in social media and combining it with deep learning technology, better performance can be achieved in disease identification and prediction tasks.
[0004] At present, although there have been studies on embedding knowledge graphs into disease prediction models, these methods usually rely on knowledge graphs or other structured clinical data in the field of traditional medicine. When facing epidemic infectious diseases such as SARS, Ebola virus, and new coronavirus, they are often limited by the sensitivity of patient data and the difficulty of timely acquisition and sharing. In addition, although the knowledge graph embedding model currently proposed can learn knowledge representation of entities and relationships in the knowledge graph, its representation ability is weak when facing complex data with hierarchical structures. The highly personalized and social characteristics of social media data require more flexible processing methods for disease prediction models. Summary of the invention
[0005] In view of the shortcomings of the existing technology, the present invention proposes a disease identification and prediction method based on social media knowledge graph embedding to solve the problems that the existing disease identification model is difficult to effectively identify patients with specific diseases in social media and the traditional knowledge graph embedding method has insufficient performance when processing hierarchical data.
[0006] The technical solution adopted by the present invention to solve its technical problem is:
[0007] A disease recognition and prediction method based on social media knowledge graph embedding includes the following steps:
[0008] Step 1: Obtain self-reported data from social media users and construct a social media disease recognition dataset, which includes basic information of social media users and their published disease self-report posts, and divide the social media disease recognition dataset into a training set, a test set, and a validation set.
[0009] Step 2: Construct a knowledge graph of social media user self-reports, extract user age, disease, symptoms, time, medicine and other related information from the social media user self-reports as entities, and the corresponding relationships are formed by the discretization of each entity. The social media user self-report knowledge graph will include a triple set {(h1, r1, t1), (h2, r2, t2), ...}. Among them, h represents the head entity, r represents the relationship, and t represents the tail entity.
[0010] Step 3: Construct a knowledge graph embedding optimization model KG2ER, which extends the ComplEx model embedding to the hyperbolic space, performs representation learning on the triples in the self-reported knowledge graph of social media users, and uses the Riemannian gradient descent method to optimize the Poincaré sphere model to ensure that the embedding points are always in the valid area, thereby realizing the embedding of the knowledge graph.
[0011] Step 4: Construct a disease recognition prediction model KGBertCS based on social media knowledge graph embedding. The disease recognition prediction model based on social media knowledge graph embedding includes: a social media user self-report text encoding module, a knowledge embedding module, a span representation module, a joint attention knowledge fusion module and an identification prediction module. The disease recognition prediction model based on social media knowledge graph embedding is trained using the data in the training set to obtain a trained disease recognition prediction model based on social media knowledge graph embedding.
[0012] Step 5: Use the disease recognition prediction model based on social media knowledge graph embedding trained in step 4 to identify the disease and evaluate the prediction indicators.
[0013] Further, the process of step 3 is as follows:
[0014] Step 3.1: Extend the ComplEx model embedding to the hyperbolic space, i.e., the hyperbolic space embedding, and obtain the score of the triple (h, r, t) in the hyperbolic space by defining the scoring function.
[0015] Further, the process of step 3.1 is as follows:
[0016] The Poincaré ball model is used to represent entities and relationships in hyperbolic space. The embedding vectors h, r, and t of entities and relationships are all located in the unit ball of hyperbolic space. In the hyperbolic space, the scoring function is redefined to replace the original scoring method based on the complex inner product in the ComplEx model. The hyperbolic cosine distance is used to calculate the score of the triple (h, r, t) to measure the matching degree of the triple in the embedding space. The specific formula is: in, represents the inner product of the embedded vectors h, r, t in the hyperbolic space, -cosh -1 () is the inverse hyperbolic cosine function. In the Poincaré ball model, the inner product in the hyperbolic space is defined as: <h,r·t> Represents the similarity between the head entity h and the relation r and the tail entity t. Based on the above score function, the logistic sigmoid function is used to calculate the probability that each triple is true. Specifically, for each triple (h, r, t), the probability that it is true is For each positive and negative sample pair, the cross entropy loss function of the Poincaré sphere model is calculated to quantify the difference between the model's prediction and the actual label: Among them, y i is the label of the positive sample, y i =1 indicates a positive sample, y i =0 indicates a negative sample.
[0017] Step 3.2: Use Riemannian gradient descent to optimize, ensuring that the updated embedding is still in the hyperbolic space. In each optimization step, first multiply the Euclidean gradient with the inverse of the Poincare metric tensor to get the Riemannian gradient of the loss function with respect to the embedding: Where g(θ) -1 is the inverse of the Poincare metric tensor, is the Euclidean gradient. For the D-dimensional Poincaré ball model, its metric tensor is Among them, ∥θ∥ is the norm of the embedding vector θ in the Euclidean space, is a d×d identity matrix. Then update along the gradient direction, and calculate the gradient to update the embedding vector: exponential mapping from tangent space to manifold, η is the learning rate.
[0018] Further, the process of step 4 is as follows:
[0019] Step 4.1: The social media user self-report text encoding module described in step 4. First, the social media user self-report is represented by text serialization, and the text sequence X = [x1, x2, ..., x n ]. Use the pre-trained model BERT-Base-Chinese as the text encoder to encode the text sequence X: Get text embedding vector
[0020] Step 4.2: The knowledge embedding module described in step 4. Use the knowledge graph embedding model described in step 3 to embed each entity e in the triple (h, r, t) of the social media user's self-reported knowledge graph. i and the relationship j Encode and get the embedding vector θ = [e h ,r j ,e t ], use linear transformation to transform the knowledge graph embedding vector θ into the text embedding vector H x Same dimensions: θ'=W1θ+b1, where W1 is a linear transformation weight matrix and b1 is a bias vector.
[0021] Step 4.3: The span representation module in step 4. The text embedding vector H obtained in steps 4.1 and 4.2 is x The embedding vector θ′ of the knowledge graph is input into the LSTM model. The LSTM model dynamically adjusts the representation of each word based on the contextual information in the user's self-report text, and dynamically captures the temporal characteristics of the relationship between entities in the graph to understand the semantics and structure of the text and graph. The specific process includes: calculating the cell state of each time step through the input gate, forget gate, and output gate of the LSTM model and the hidden state h t =o t ⊙tanh(c t ), where i t 、f t , o t They represent input gate, forget gate, and output gate respectively. Indicates the candidate status of the current cell.
[0022] Step 4.4: The joint attention knowledge fusion module described in step 4. Use the channel attention mechanism and the spatial attention mechanism to perform knowledge fusion, and input the hidden layer h of the text embedding vector obtained in step 4.3 as the input feature. text and the hidden layer h of the knowledge graph embedding vector kg , B is the batch size, C1 is the number of channels of text embedding, C2 is the number of channels of knowledge graph embedding, and H and W are the height and width of the feature map respectively.
[0023] The channel attention mechanism uses two convolutional layers and a batch normalization layer to generate channel attention weights. First, the number of channels of the input feature is reduced by a convolution operation: h′=Conv1(h), and the vectors h′ are obtained respectively. text and h′ kg , the number of channels is reduced from C1 to C′1, and C2 to C′2. Then the ReLU activation function is used to introduce nonlinearity and make adjustments: h″=ReLU(h′), and the vector h″ is obtained respectively. text and h″ kg The activated feature map is then restored to its original number of channels through a convolutional layer, h final =Conv2(h″), and we get vectors and The average activation value of each channel is calculated by global average pooling, and then the attention weight of each channel is obtained by Sigmoid activation function: α c =σ(GAP(h final )),in, is the Sigmoid activation function, is the global average pooling. The obtained channel attention weight and Multiply the corresponding elements of the original feature map of the text and knowledge graph: h att1 =α c h, get the weighted channel attention feature map and
[0024] The spatial attention mechanism further processes the feature map after applying channel attention, generates attention weights for emphasizing or suppressing spatial positions, and enhances the spatial feature expression ability of the model. and Perform global maximum pooling and global average pooling in the channel dimension to obtain two H×W×1 feature maps respectively. and The obtained feature map is concatenated according to the channel to obtain a feature map of size H×W×2. A 7×7 convolution operation is performed on the concatenated result to obtain a feature map of size H×W×1. Then, the Sigmoid activation function is used to obtain the spatial attention weight: α s =σ(f 7×7 [h catt——ax ;h catt_max ]), the spatial attention weight matrix is then multiplied by the original feature map hsatt =α s ·h catt , we can get the weighted spatial attention feature map and Finally, additive fusion is used to fuse the user self-report text feature map and the knowledge graph feature map processed by the channel-spatial attention mechanism:
[0025] Step 4.3: Identify the prediction module described in step 4. The feature h after the channel-spatial attention knowledge fusion fused The classifier is composed of a fully connected layer, which maps the input feature vector to a single value, that is, a linear transformation: z = W2·h fused +b2, where W2 is the weight matrix and b2 is the bias term. Since the present invention is aimed at a binary classification task, that is, predicting whether a user has a specific disease, the sigmoid activation function is used for prediction to obtain the probability that the user has the disease. If P(y|z)>=0.5, it means that the model predicts that the user suffers from the disease; if P(y|z)<0.5, it means that the model predicts that the user does not suffer from the disease.
[0026] The beneficial effects of adopting the above technical solution are:
[0027] The present invention first proposes a knowledge embedding optimization algorithm based on hyperbolic space, which extends the complex number embedding in ComplEx to the hyperbolic space, enhances the model's ability to represent complex data with hierarchical structures, and uses the Riemannian gradient descent method for optimization to ensure that the embedding points are always in the valid area, thereby realizing the embedding of social media knowledge graphs. Based on the BERT+LSTM model, the channel attention mechanism and the spatial attention mechanism are combined to fuse the self-report text of social media users with the knowledge graph, and the information in the knowledge graph is used to enhance the model's spatial feature expression ability, so as to realize the recognition and prediction of users with specific diseases in social media. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 is a flow chart of the present invention;
[0029] Figure 2 This is a network architecture diagram of the present invention. DETAILED DESCRIPTION
[0030] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but the implementation of the present invention is not limited thereto.
[0031] Embodiment 1:
[0032] Combined with Figure 1 As shown, a disease recognition and prediction method based on social media knowledge graph embedding includes:
[0033] Step 1: Obtain self-reported data from social media users and construct a social media disease recognition dataset, which includes basic information of social media users and their published disease self-report posts, and divide the social media disease recognition dataset into a training set and a test set.
[0034] Step 2: Construct a knowledge graph of social media user self-reports, extract user age, disease, symptoms, time, medicine and other related information from the social media user self-reports as entities, and the corresponding relationships are formed by the discretization of each entity. The social media user self-report knowledge graph will include a set of triples {(h1, r1, t1), (h2, r2, t2), ...}. Among them, h represents the head entity, r represents the relationship, and t represents the tail entity.
[0035] Step 3: Construct a knowledge graph embedding optimization model KG2ER, which extends the ComplEx model embedding to the hyperbolic space, performs representation learning on the triples in the self-reported knowledge graph of social media users, and uses the Riemannian gradient descent method to optimize the Poincaré sphere model to ensure that the embedding points are always in the valid area, thereby realizing the embedding of the knowledge graph.
[0036] Step 4: Construct a disease recognition prediction model KGBertCS based on social media knowledge graph embedding. The disease recognition prediction model based on social media knowledge graph embedding includes: a social media user self-report text encoding module, a knowledge embedding module, a span representation module, a joint attention knowledge fusion module and an identification prediction module. The disease recognition prediction model based on social media knowledge graph embedding is trained using the data in the training set to obtain a trained disease recognition prediction model based on social media knowledge graph embedding.
[0037] Step 5: Use the disease recognition prediction model based on social media knowledge graph embedding trained in step 4 to identify the disease and evaluate the prediction indicators.
[0038] Embodiment 2:
[0039] Based on Example 1, Figure 1 As shown, the step 3 specifically includes:
[0040] Step 3.1: Embed the ComplEx model into the hyperbolic space, i.e., hyperbolic space embedding, and obtain the score of the triple (h, r, t) in the hyperbolic space by defining a scoring function. The Poincaré ball model is used to represent entities and relationships in the hyperbolic space. The embedding vectors h, r, t of entities and relationships are all located within the unit ball of the hyperbolic space inside, and the is that the norm of each embedding vector is less than 1. Among them, represents the d-dimensional Euclidean space, that is, the set of all real vectors with dimension d; ||x|| is the norm of vector x, that is, the length or magnitude of the vector.
[0041] In the hyperbolic space, redefine the scoring function to replace the scoring method based on complex inner product in the original ComplEx model. Use the hyperbolic cosine distance to calculate the score of the triple (h, r, t) and measure the matching degree of the triple in the embedding space. The specific formula is: Among them, represents the inner product of the embedding vectors h, r, t in the hyperbolic space, and -cosh -1 () is the inverse hyperbolic cosine function. In the Poincaré ball model, the inner product in the hyperbolic space is defined as: <h, r·t> represents the similarity between the head entity h, the relationship r and the tail entity t. Since the hyperbolic space has geometric properties different from the Euclidean space, additional normalization factors (1 - ||h|| 2 ) and (1 - ||r·t|| 2 ) are used to ensure the consistency and stability of the score calculation. According to the above scoring function, use the logistic sigmoid function to calculate the probability that each triple is true. Specifically, for each triple (h, r, t), the probability that it is true For each positive and negative sample pair, calculate the cross-entropy loss function of the Poincaré ball model to quantify the difference between the model's prediction and the actual label: Among them, yi is the label of the positive sample, yi = 1 indicates a positive sample, and yi = 0 indicates a negative sample.
[0042] Step 3.2: Since in the hyperbolic space, the standard Euclidean optimization method is not applicable, the Riemannian gradient descent method is used for optimization to ensure that the updated embedding is still within the hyperbolic space. In each optimization step, first multiply the Euclidean gradient by the inverse of the Poincaré metric tensor to obtain the Riemannian gradient of the loss function with respect to the embedding: where g(θ) -1 is the inverse of the Poincaré metric tensor, is the Euclidean gradient. For the D-dimensional Poincaré ball model, its metric tensor is Where ||θ|| is the norm of the embedding vector θ in Euclidean space, is a d×d identity matrix. Then update along the gradient direction, and calculate the gradient to update the embedding vector: exponential mapping from tangent space to manifold, η is the learning rate.
[0043] Embodiment 3:
[0044] Based on Example 2, Figure 1 As shown, the step 4 specifically includes:
[0045] Step 4.1: The social media user self-report text encoding module described in step 4. First, the social media user self-report is represented by text serialization to obtain a text sequence X = [x1, x2, ..., x n ]. Use the pre-trained model BERT-Base-Chinese as the text encoder to encode the text sequence X: Get text embedding vector
[0046] Step 4.2: The knowledge embedding module described in step 4. Use the knowledge graph embedding model described in step 3 to embed each entity e in the triple (h, r, t) of the social media user's self-reported knowledge graph. i and the relationship j Encode and get the embedding vector θ = [e h , r j , e t ], use linear transformation to transform the knowledge graph embedding vector θ into the text embedding vector H x Same dimension: θ′=W1θ+b1, where W1 is a linear transformation weight matrix and b1 is a bias vector.
[0047] Step 4.3: The span representation module in step 4. The text embedding vector H obtained in steps 4.1 and 4.2 is x The embedding vector θ′ of the knowledge graph is input into the LSTM model. The LSTM model dynamically adjusts the representation of each word based on the contextual information in the user's self-report text, and dynamically captures the temporal characteristics of the relationship between entities in the graph to understand the semantics and structure of the text and graph. The specific process includes: calculating the cell state of each time step through the input gate, forget gate, and output gate of the LSTM model and the hidden state h t=o t ⊙tanh(c t ), where i t 、f t , o t They represent input gate, forget gate, and output gate respectively. Indicates the candidate status of the current cell.
[0048] Step 4.4: The joint attention knowledge fusion module described in step 4. Use the channel attention mechanism and the spatial attention mechanism to perform knowledge fusion, and input the hidden layer h of the text embedding vector obtained in step 4.3 as the input feature. text and the hidden layer h of the knowledge graph embedding vector kg , B is the batch size, C1 is the number of channels of text embedding, C2 is the number of channels of knowledge graph embedding, and H and W are the height and width of the feature map respectively.
[0049] The channel attention mechanism uses two convolutional layers and a batch normalization layer to generate channel attention weights. First, the number of channels of the input feature is reduced by a convolution operation: h′=Conv1(h), and the vectors h′ are obtained respectively. text and h′ kg , the number of channels is reduced from C1 to C′1, and C2 to C′2. Then the ReLU activation function is used to introduce nonlinearity and make adjustments: h″=ReLU(h′), and the vector h″ is obtained respectively. text and h″ kg The activated feature map is then restored to its original number of channels through a convolutional layer, h final =Conv2(h″), and we get vectors and The average activation value of each channel is calculated by global average pooling, and then the attention weight of each channel is obtained by Sigmoid activation function: α c =σ(GAP(h final )),in, is the Sigmoid activation function, is the global average pooling. The obtained channel attention weight and Multiply the corresponding elements of the original feature map of the text and knowledge graph: h att1 =α c h, get the weighted channel attention feature map and
[0050] The spatial attention mechanism further processes the feature map after applying channel attention, generates attention weights for emphasizing or suppressing spatial positions, and enhances the spatial feature expression ability of the model. and Perform global maximum pooling and global average pooling in the channel dimension to obtain two H×W×1 feature maps respectively. and The obtained feature map is concatenated according to the channel to obtain a feature map of size H×W×2. A 7×7 convolution operation is performed on the concatenated result to obtain a feature map of size H×W×1. Then, the Sigmoid activation function is used to obtain the spatial attention weight: α s =σ(f 7×7 [h catt_max ;h catt_max ]), the spatial attention weight matrix is then multiplied by the original feature map h satt =α s ·h catt , we can get the weighted spatial attention feature map and Finally, additive fusion is used to fuse the user self-report text feature map and the knowledge graph feature map processed by the channel-spatial attention mechanism:
[0051] Step 4.3: Identify the prediction module described in step 4. The feature h after channel-spatial attention knowledge fusion fused The classifier is composed of a fully connected layer, which maps the input feature vector to a single value, that is, a linear transformation: z = W2·h fused +b2, where W2 is the weight matrix and b2 is the bias term. Since the present invention is aimed at a binary classification task, that is, predicting whether a user has a specific disease, the sigmoid activation function is used for prediction to obtain the probability that the user has the disease. If P(y|z)>=0.5, it means that the model predicts that the user has the disease; if P(y|z)<0.5, it means that the model predicts that the user does not have the disease. During the training process, the classifier uses a binary cross entropy loss function to evaluate the difference between the model output and the actual label. The loss function Loss=-[ylog(p)+(1-y)log(1-p)], where y is the true label. If the user has the disease, y=1; if the user does not have the disease, y=0.
[0052] Embodiment 4:
[0053] On the basis of Example 3, in order to verify the effectiveness of the solution of the present invention, further explanation is given below in combination with experimental data:
[0054] The present invention uses COVID-19 sequelae as a specific disease for experimentation, and uses the joint extraction model of previous research results to conduct a knowledge graph embedding experiment on the triple data set WLE extracted from the self-reported text of the users of COVID-19 sequelae on Weibo. The data set includes 27,096 triples, including 4 entities: symptoms, diseases, drugs, and time; 5 relationships: accompanying, symptom period, illness period, medication period, and treatment. The WLP data set was constructed by data cleaning and manual annotation of the self-reported text of the users of COVID-19 sequelae on Weibo, and the triple vector data obtained by knowledge graph embedding was used to train and test the disease recognition prediction model. The data set contains 260 million Chinese texts.
[0055] The size of the dataset used in the experiment is shown in Table 1.
[0056] Table 1 Dataset size
[0057] Number of training sets Number of test sets Number of validation sets WLE 21660 2926 2510 WLP 20800 2600 2600
[0058] For the knowledge graph embedding experiment, the minimum loss Minloss, mean rank MR (Mean Rank) and Hits@10 are used as evaluation indicators, where Minloss represents the minimum loss value obtained by the model during the training process; MR represents the average rank of the correct triple among all candidate triples; Hits@n represents the proportion of the model correctly predicting the top 10 triples in the prediction results. The calculation formula is as follows:
[0059]
[0060] in is the loss function calculated in each epoch, θ is the model parameter; N is the total number of triplets, Rank i is the ranking of the i-th triple among all candidate triples; Is the indicator function (1 if the condition is true, 0 otherwise).
[0061] The accuracy, recall, and F1 score are used as evaluation indicators for the COVID-19 sequelae recognition and prediction experiment. The calculation formula is as follows:
[0062]
[0063] Among them, TP represents the number of both actual and predicted positive examples; FP represents the number of predicted positive examples but actual negative examples; FN represents the number of predicted negative examples but actual positive examples.
[0064] The performance of the hyperbolic space-based knowledge graph embedding optimization algorithm proposed in the present invention is compared with four other methods on the WLE dataset, and the experimental results are shown in Table 2. The results show that the overall performance of the KG2E algorithm proposed in the present invention in the task of embedding the self-reported knowledge graph of users with COVID-19 sequelae is significantly better than that of the traditional model. Its lower minimum loss, best average ranking, and HITS@10 index of up to 84.15% indicate that KG2E has stronger expression and prediction capabilities when processing knowledge graph data with hierarchical structures and complex relationships.
[0065] Table 2. Comparative experiment on knowledge graph embedding
[0066] Model Minloss MR HITS@10 TransA 0.4234 78.64 0.5536 TransD 0.4102 70.85 0.6486 TransE 0.3850 66.23 0.7224 TransH 0.3645 66.01 0.7356 KG2ER 0.3260 54.68 0.8415
[0067] In order to demonstrate the performance of the disease recognition prediction model KGBertCS based on social media knowledge graph embedding proposed in the present invention, a comparative experiment was conducted with other baseline models on the WLP dataset, and the experimental results are shown in Table 3. The results show that the performance of KGBertCS is significantly better than that of traditional BERT models (such as Bert and BertLSTM) and other variant models (such as BertDPCNN and BertHAN, etc.), indicating that such models have limitations in feature extraction of colloquial social media texts. Although convolutional neural network variants (such as BertCNN and BertDPCNN) have improved certain performance, they still cannot match the overall performance of KGBertCS. The attention mechanism variant BertATT performs well, but still lags behind KGBertCS, mainly because the present invention combines the semantic embedding of the knowledge graph and provides more external knowledge about the sequelae of the new crown and its related symptoms, which enables the model to better understand and process complex information related to user self-reports in social media. Through the introduction of the knowledge graph, the model can not only mine deep entity information, but also convert unstructured data in the text into structured knowledge, thereby improving the accuracy and recall of the prediction.
[0068] Table 3 Comparative experiment on the recognition and prediction of COVID-19 sequelae
[0069] Model Precision Recall F1 Bert 84.56 87.02 85.77 BertCNN 88.25 90.12 89.18 BertCNNPlus 88.68 90.56 89.61 BertDPCNN 89.63 91.68 90.64 BertHAN 90.23 93.05 91.62 BertLSTM 87.35 89.02 88.18 BertRCNN 88.80 90.36 89.57 BertATT 90.89 91.26 91.07 KGBertCS 94.85 96.23 95.54
[0070] The embodiments of the present invention are described in detail above, and the above embodiments are only preferred embodiments for fully illustrating the present invention, and the protection scope of the present invention is not limited thereto. Those skilled in the art can design many other modifications and implementation methods based on the present invention, and these modifications and implementation methods will fall within the scope and spirit of the principles disclosed in this application.
Claims
1. A disease recognition and prediction method based on social media knowledge graph embedding, characterized in that: The following steps are involved: The following steps are involved: Step 1: Obtain self-reported data from social media users and construct a social media disease recognition dataset, which includes basic information of social media users and their self-reported disease posts, and divide the social media disease recognition dataset into a training set, a test set, and a validation set; Step 2: Construct a knowledge graph of social media user self-reports, extract user-related information from the social media user self-reports as entities, and the corresponding relationships are formed by discretization of each entity; Step 3: Construct a knowledge graph embedding optimization model, which extends the ComplEx model embedding into the hyperbolic space, performs representation learning on the triples in the social media user self-reported knowledge graph, and uses the Riemannian gradient descent method to optimize the Poincaré sphere model to ensure that the embedding point is always in the valid area, thereby realizing the embedding of the knowledge graph; Step 4: construct a disease recognition prediction model based on social media knowledge graph embedding, the disease recognition prediction model based on social media knowledge graph embedding includes: a social media user self-report text encoding module, a knowledge embedding module, a span representation module, a joint attention knowledge fusion module and a recognition prediction module, and use the data in the training set to train the disease recognition prediction model based on social media knowledge graph embedding to obtain a trained disease recognition prediction model based on social media knowledge graph embedding; Step 5: Use the disease recognition prediction model based on social media knowledge graph embedding trained in step 4 to identify the disease and evaluate the prediction indicators.
2. The disease identification and prediction method based on social media knowledge graph embedding according to claim 1 is characterized in that: The step 3 specifically includes: Step 3.1: Extend the ComplEx model embedding to the hyperbolic space, and get the score of the triple (h, r, t) in the hyperbolic space by defining the score function; use the Poincaré ball model to represent entities and relations in the hyperbolic space; the embedding vectors h, r, t of the entity and relationship are all located in the unit sphere of the hyperbolic space In the hyperbolic space, the scoring function is redefined to replace the scoring method based on the complex inner product in the original ComplEx model; the hyperbolic cosine distance is used to calculate the score of the triple (h, r, t) to measure the matching degree of the triple in the embedding space. The specific formula is: in, represents the inner product of the embedded vectors h, r, t in the hyperbolic space, -cosh -1 () is the inverse hyperbolic cosine function; in the Poincaré ball model, the inner product in the hyperbolic space is defined as: <h, r·t> represents the similarity between the head entity h and relation r and the tail entity t; the probability of each triple being true is calculated using the logistic sigmoid function according to the above score function. Specifically, for each triple (h, r, t), the probability of it being true is For each positive and negative sample pair, the cross entropy loss function of the Poincaré sphere model is calculated to quantify the difference between the model's prediction and the actual label: Among them, y i is the label of the positive sample, y i =1 indicates a positive sample, y i =0 indicates a negative sample; Step 3.2: Use Riemannian gradient descent to optimize, ensuring that the updated embedding is still in the hyperbolic space; in each optimization step, first multiply the Euclidean gradient with the inverse of the Poincare metric tensor to get the Riemannian gradient of the loss function with respect to the embedding: Where g(θ) -1 is the inverse of the Poincare metric tensor, is the Euclidean gradient; for the D-dimensional Poincaré ball model, its metric tensor Where ||θ|| is the norm of the embedding vector θ in Euclidean space, is a d×d identity matrix; then update along the gradient direction, and calculate the gradient to update the embedding vector: exponential mapping from tangent space to manifold, η is the learning rate.
3. The disease identification and prediction method based on social media knowledge graph embedding according to claim 1 is characterized in that: The social media user self-report text encoding module in step 4 is specifically implemented as follows: First, the social media user self-report is represented by text serialization, and the text sequence X = [x1, x2, ..., x n ]; Use the pre-trained model BERT-Base-Chinese as the text encoder to encode the text sequence X: Get text embedding vector 4. The disease identification and prediction method based on social media knowledge graph embedding according to claim 1 is characterized in that: The knowledge embedding module in step 4 is specifically implemented as follows: Use the knowledge graph embedding model described in step 3 to embed each entity e in the triple (h, r, t) of the social media user’s self-reported knowledge graph i and the relationship j Encode and get the embedding vector θ = [e h , r j , e t ], use linear transformation to transform the knowledge graph embedding vector θ into the text embedding vector H x Same dimension: θ′=W1θ+b1, where W1 is a linear transformation weight matrix and b1 is a bias vector.
5. The disease recognition and prediction method based on social media knowledge graph embedding according to claim 1 is characterized in that: The span representation module in step 4 is specifically implemented as follows: Embed the text into vector H x and the knowledge graph embedding vector θ′ are respectively input into the LSTM model. The specific process includes: calculating the cell state of each time step through the input gate, forget gate, and output gate of the LSTM model and the hidden state h t =o t ⊙tanh(c t ), where i t 、f t , o t They represent input gate, forget gate, and output gate respectively. Indicates the candidate status of the current cell.
6. The disease identification and prediction method based on social media knowledge graph embedding according to claim 1 is characterized in that: The joint attention knowledge fusion module in step 4 is specifically implemented as follows: The channel attention mechanism and the spatial attention mechanism are used to perform knowledge fusion, and the input feature is the hidden layer h of the text embedding vector obtained in step 4.
3. text and the hidden layer h of the knowledge graph embedding vector kg , B is the batch size, C1 is the number of channels for text embedding, C2 is the number of channels for knowledge graph embedding, H and W are the height and width of the feature map respectively; The channel attention mechanism uses two convolutional layers and a batch normalization layer to generate channel attention weights. First, the number of channels of the input feature is reduced by a convolution operation: h′=Conv1(h), and the vectors h′ are obtained respectively. text and h′ kg , the number of channels is reduced from C1 to C′1, and C2 to C′2; then the ReLU activation function is used to introduce nonlinearity and make adjustments: h″=ReLU(h′), and the vector h″ is obtained respectively text and h″ kg ; The activated feature map is then restored to the original number of channels through a convolutional layer, h final =Conv2(h″), and we get vectors and The average activation value of each channel is calculated by global average pooling, and then the attention weight of each channel is obtained by Sigmoid activation function: α c =σ(GAP(h finai )),in, is the Sigmoid activation function, is the global average pooling; the obtained channel attention weight and Multiply the corresponding elements of the original feature map of the text and knowledge graph: h att1 =α c h, get the weighted channel attention feature map and The spatial attention mechanism further processes the feature map after applying channel attention, generates attention weights for emphasizing or suppressing spatial positions, and enhances the spatial feature expression ability of the model; first, the input feature map and Perform global maximum pooling and global average pooling in the channel dimension to obtain two H×W×1 feature maps respectively. and The obtained feature map is concatenated according to the channel to obtain a feature map of size H×W×2. A 7×7 convolution operation is performed on the concatenated result to obtain a feature map of size H×W×1. Then, the Sigmoid activation function is used to obtain the spatial attention weight: α s =σ(f 7×7 [h catt_max ;h catt_max ]), the spatial attention weight matrix is then multiplied by the original feature map h satt =α s ·h catt , we can get the weighted spatial attention feature map and Finally, additive fusion is used to fuse the user self-report text feature map and the knowledge graph feature map processed by the channel-spatial attention mechanism:
7. The disease identification and prediction method based on social media knowledge graph embedding according to claim 1 is characterized in that: The identification prediction module in step 4 is specifically implemented as follows: The feature h after the channel-spatial attention knowledge fusion fused The classifier is composed of a fully connected layer, which maps the input feature vector to a single value, that is, performs a linear transformation: z = W2·h fused +b2, where W2 is the weight matrix and b2 is the bias term; for binary classification tasks, the sigmoid activation function is used for prediction.
Citation Information
Patent Citations
Film and television entity identification method based on Bilstm-crf and knowledge graph
CN110298042A
Medical inquiry recommendation method based on knowledge graph and social media
CN111897967A
Intelligent disease prediction system based on medical knowledge graph
CN112151188A
Depression detection system based on social media
CN113139062A
Disease prediction device and equipment based on traditional Chinese medicine diagnosis atlas and readable storage medium
CN114155961A
Cited By
Large model data grading method and device
CN121681711A