A privacy-preserving embedding method for multimodal knowledge graphs of medical data
Through the multimodal graph self-attention layer and convolutional embedding model to process multimodal diagnosis and treatment data, the problems of insufficient interaction processing and lack of privacy protection in the prior art are solved, efficient embedding and deep fusion are achieved, and the security and reliability of diagnosis and treatment data are enhanced.
Patent Information
- Application Number
- CN202411592957.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-08
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2044-11-08
AI Technical Summary
The existing multimodal knowledge graph embedding methods are not accurate enough when processing complex interactions of multimodal data, making it difficult to achieve end-to-end training, and lack an effective privacy protection mechanism, resulting in leakage of sensitive information.
The multimodal graph self-attention layer and convolutional embedding model are used to process text, visual and numerical information by initializing embedding, and structured information of the multimodal knowledge graph is fused with multimodal knowledge graph, and scored through the convolutional embedding model to enhance privacy protection.
Effectively process multimodal diagnosis and treatment data, enhance data privacy, improve the model's inference ability and data security in a multimodal environment, and ensure the safety and reliability of diagnosis and treatment data in the analysis and reasoning process.
Smart Images

Figure CN119475429B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of data processing technology, and in particular relates to a privacy-preserving embedding method for multimodal knowledge graphs of medical data. Background Art
[0002] A knowledge graph is a model that organizes and displays data through a graph structure, with nodes representing entities and edges representing relationships between entities. As data complexity increases, knowledge graphs have gradually evolved into multimodal knowledge graphs, integrating data from different modalities, such as text and images, enriching entity descriptions and enhancing reasoning capabilities. To better process this structured data, knowledge graph embedding technology has emerged. By mapping entities and relationships into a low-dimensional vector space, knowledge graphs can be computationally processed and applied to tasks such as knowledge completion and relationship prediction. Multimodal knowledge graph embedding further embeds modal data such as text and images into a unified vector space, improving the model's reasoning and prediction capabilities in multimodal environments.
[0003] In the medical field, medical data is highly multimodal, encompassing text, images, and numerical data. Medical medical data knowledge graphs, through structured modeling of patient, disease, and treatment data, help improve intelligent diagnosis and personalized healthcare. However, medical data contains sensitive personal information, such as patient records and images, which, if leaked, could pose a serious threat to patient privacy. Therefore, maintaining data analysis capabilities while ensuring privacy is a key challenge in the application of medical knowledge graphs.
[0004] Existing technologies combine latent, relational, and numerical features to perform end-to-end knowledge graph embedding learning to improve the performance of knowledge graph completion tasks. Existing technologies also use different neural encoders for different data types and combine them with traditional relational models to enhance the embedding of entities and multimodal data. Existing technologies also propose adversarial feature learning methods that effectively map the text and image data of entities into a unified vector space through a multi-relational feature aggregation network to capture multimodal information. Existing technologies also introduce a hierarchical attention network for multimodal medical knowledge graphs, providing a solution for explanatory medical question-answering.
[0005] However, due to the complex intra- and inter-modal interactions exhibited by entities containing multimodal data, as well as the heterogeneity of multimodal knowledge graphs, existing multimodal knowledge graph embedding methods are not precise enough in handling the complex interactions of multimodal data, cannot fully capture the relationships between different modalities, and are difficult to achieve end-to-end training. Furthermore, they are limited in processing heterogeneous graph structures and multi-relationship chains, making it difficult to fully integrate the characteristics of multimodal data, resulting in limited model performance. Furthermore, existing multimodal knowledge graph embedding technologies have shortcomings in terms of privacy protection, failing to effectively prevent potential reconstruction of data, which could lead to the leakage of sensitive information. There are no specifically designed privacy protection mechanisms during multimodal data processing, making it difficult to ensure data security during analysis and sharing. Summary of the Invention
[0006] In response to the above-mentioned deficiencies in the prior art, the present invention provides a privacy-preserving embedding method for multimodal knowledge graphs of medical data to address the technical problems of existing multimodal knowledge graph embedding methods, such as insufficient processing of complex interactions of multimodal data, limited heterogeneous graph processing capabilities, difficulty in achieving end-to-end training, lack of privacy protection mechanisms, and difficulty in preventing leakage of sensitive information.
[0007] To achieve the above objectives, the present invention adopts the following technical solution: a privacy-preserving embedding method for multimodal knowledge graphs of medical data, comprising the following steps:
[0008] S1. Initialize and embed text information and numerical information based on diagnosis and treatment data, and initialize and embed visual information based on medical image visual data;
[0009] S2. Based on the initial embedding result, the multimodal graph self-attention layer is used to obtain the multimodal attention coefficient of the neighboring entity;
[0010] S3, based on the multimodal attention coefficient, fuses the structured information of the multimodal knowledge graph by stacking multiple encoder layers;
[0011] S4. Based on the structured information of the fused multimodal knowledge graph, the convolutional embedding model is used to score the privacy-preserving embedding of the multimodal knowledge graph.
[0012] The beneficial effects of the present invention are: the present invention can efficiently process multimodal information such as text, images and numerical values in medical data, and further deepen it through multi-layer encoders, collect multi-hop neighbor information and its related multimodal medical data, and reduce the risk of sensitive information exposure as a whole through the embedding process, thereby enhancing data privacy. In order to solve the technical problems of the existing multimodal knowledge graph embedding method, such as insufficient complex interactive processing of multimodal data, limited heterogeneous graph processing capabilities, difficulty in achieving end-to-end training, lack of privacy protection mechanism, and difficulty in preventing sensitive information leakage, the present invention effectively realizes efficient embedding and deep fusion of multimodal information, enhances the processing capabilities of heterogeneous data, and prevents sensitive information leakage through the embedding process, provides an end-to-end privacy protection mechanism, ensures the security and reliability of medical data during analysis and reasoning, enables multimodal data and structural information to be jointly learned under a unified framework, simplifies the training process and improves model performance.
[0013] Furthermore, the step S1 includes the following steps:
[0014] S101. Use the pre-trained language model to obtain the contextual information of the sentence-level description of entities and relationships in the diagnosis and treatment data, and generate text information embedding t ;
[0015] S102, extract visual information and embed it into medical image visual data v ;
[0016] S103, encode the entity numerical features in the diagnosis and treatment data and process them in the fully connected layer in turn to generate numerical data embedding e n , complete the initialization embedding of text information, visual information and numerical information.
[0017] The beneficial effects of this further solution are as follows: by converting medical data into a low-dimensional vector representation, the present invention reduces the risk of direct exposure of sensitive information, effectively preventing the reverse derivation of original information during analysis, transmission, and storage, and improving privacy protection. It can also simultaneously process multimodal data such as text, images, and numerical values in the medical data knowledge graph, providing a more comprehensive embedded representation of entities.
[0018] Furthermore, step S2 includes the following steps:
[0019] S201, using text information to embed e t and numerical data embedded in e n , construct an embedding matrix, and use visual information to embed e v , construct the modality matrix, where each modality graph attention layer processes two embedding matrices and one modality matrix;
[0020] S202. Utilize the multimodal graph self-attention layer to obtain a multimodal attention coefficient.
[0021] The beneficial effect of the above further solution is that the present invention improves the comprehensiveness of feature representation and data understanding ability by fusing multimodal information through a multimodal embedding matrix and a self-attention layer.
[0022] Furthermore, the expression of the multimodal attention coefficient is as follows:
[0023]
[0024] c jk =W1[h j ||g k ||m k ]
[0025] Among them, b ijk and Both represent multimodal attention coefficients, c i Represents the central entity e i The eigenvector of jk Represents the central entity e i In comparison, the tail entity e j and relationship r k Composite entity vector, d k Indicates the dimension of the vector, used to normalize the result of the vector dot product, exp() represents the natural exponential function, used to convert b ijk The value of b is mapped to the positive range, inr Represents entity e i The absolute attention parameter between the neighbor entity n through the relation r, n represents the entity e i A neighbor node, N i Represents entity e i The first-order neighbor set, r represents the connection center entity e i A relation with a neighbor entity, R in Represents entity e i The relationship between the connection and the neighbors, W1 represents the learnable weight matrix, || represents the connection operation, h j and g k Represent the tail entity e j The vector and relationship r k vector, m k Represents multimodal features.
[0026] The beneficial effect of the above further scheme is that the present invention can effectively distinguish the influence of different neighbors and modalities on the central entity through the calculation of the multimodal attention coefficient, thereby improving the representation accuracy of the multimodal knowledge graph.
[0027] Furthermore, step S3 includes the following steps:
[0028] S301, linearly fuse the multimodal attention coefficients to the central node to generate an embedded representation;
[0029] S302. Concatenate the embedding representations generated from M independent self-attention layers using the following formula:
[0030]
[0031] in, Represents the output after the multi-head self-attention mechanism, MultiHead() represents the multi-head self-attention mechanism, Represents the input of the mth attention head, i.e., entity e i The expression of head m represents the output of the mth attention head, W O represents the matrix of input linear layer weights, j represents entity e i The first-order neighbor, N i Represents entity e i The first-order neighbor set of Represents the neighbor node e j For the central node e i The attention weight, Represents the neighbor node e j The feature representation combined with relation k is used in the mth attention head;
[0032] S303. Based on the connection results, the following formula is used to fuse the structured information of the knowledge graph by stacking multiple encoder layers:
[0033]
[0034] in, Represents entity e i In the embedding representation of the l+1 layer, σ represents a nonlinear activation function, and R represents the set of relations, that is, all relation types in the knowledge graph. Represents entity e i The neighbor set of c i,r Represents the factor used for normalization, adjusting between different relationships, represents the attention weight from neighbor node j to center node i in layer l, Represents a combination function for embedding node i and the relationship between nodes j and i Perform feature transformation.
[0035] The beneficial effect of the above further scheme is: the present invention aggregates neighborhood information through a multi-head attention mechanism, which can better capture the diverse semantic features of nodes, and extract local features of complex relationships through convolution operations, making node embedding more expressive, thereby improving the reasoning performance of the multimodal knowledge graph and the accuracy of node representation.
[0036] Furthermore, the expression of the embedding vector is as follows:
[0037]
[0038] Among them, x i Represents entity e i The updated embedding vector, σ represents the nonlinear activation function, and j represents the entity e i The first-order neighbor, N i Represents entity e i The first-order neighbor set, k represents the relationship between neighbor entity j and another node, R jj represents the set of relationships between neighbor entity j and nodes other than neighbor entity j, α ijk represents the attention weight coefficient, which represents the contribution of neighbor node j and relationship k to the center node i, c jk Represents the feature representation of neighbor node j and relationship k.
[0039] The beneficial effect of the above further scheme is that the present invention can weightedly integrate the contributions of different neighbors to the central node by aggregating the embedded information of neighbor nodes and their relationships, thereby generating a more refined embedding representation and improving the accuracy and expressiveness of entity representation in multimodal knowledge graphs.
[0040] Furthermore, the scoring function of the convolutional embedding model is expressed as follows:
[0041] f(h,r,t)=σ(vec(W([e h ;r]*ω))e t )
[0042] Among them, f() represents the scoring function of the triple, h represents the head entity, and r represents the connection center entity e i A relationship with neighbor entities, t represents the tail entity, σ represents the sigmoid activation function, vec represents the operation of projecting the feature map into the vector space, W represents the transformation matrix, e h and e t They represent the embedding of the head entity and the tail entity respectively, r represents the embedding vector of the relationship, ω represents the weight matrix of the convolution kernel, and * represents the convolution operation.
[0043] The beneficial effect of the above further scheme is that the present invention can extract the complex interactive features of the head entity and the relationship through the fusion of convolution operation and multimodal features, and match them with the tail entity, thereby improving the accuracy of triple scoring and enhancing the reasoning ability of multimodal knowledge graph completion.
[0044] Furthermore, the loss function of the convolutional embedding model is expressed as follows:
[0045]
[0046] in, represents the loss function of the ConvE model, Indicates the number of fact triples, y i' represents the calculated score of the i'th triple, f i' Represents a label, where 1 represents true facts and 0 represents false information.
[0047] The beneficial effect of the above further scheme is that the present invention can effectively measure the model's prediction error on the authenticity of triples by adopting the binary cross-entropy loss function, thereby optimizing the training process of the multimodal knowledge graph completion model and improving the prediction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 Flow chart of the method of the present invention.
[0049] Figure 2 This is a brief flowchart of the SAGNN model.
[0050] Figure 3 A simplified schematic diagram of a multimodal graph self-attention encoder.
[0051] Figure 4 A simple diagram of graph aggregation.
[0052] Figure 5 A brief diagram of the ConvE decoder. DETAILED DESCRIPTION
[0053] The specific embodiments of the present invention are described below to facilitate understanding of the present invention by those skilled in the art. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the appended claims, these changes are obvious, and all inventions and creations utilizing the concepts of the present invention are protected.
[0054] Example
[0055] The present invention proposes a privacy-preserving embedding method for multimodal medical data knowledge graph based on self-attention graph neural network, referred to as SAGNN (Self-Attention Graph Neural Network model). The SAGNN encoder integrates neighborhood, global structure and multimodal information, and can efficiently process multimodal information such as text, images and numerical values in medical data. It is further deepened through multi-layer encoders to collect multi-hop neighbor information and its related multimodal medical data. The embedding process reduces the risk of sensitive information exposure as a whole, thereby enhancing data privacy. The overall structure of the SAGNN model is as follows: Figure 1 As shown. In the encoder, two key modules are distinguished: the multimodal graph attention module and the multimodal aggregation module:
[0056] Multimodal Graph Self-Attention Encoder: This module dynamically updates composite entities and their features, including textual, visual, and numerical information within medical data. It facilitates embedding learning within the complex multimodal medical data graph structure. The overall embedding process effectively reduces the risk of medical data exposure, thereby enhancing privacy protection.
[0057] Graph Aggregation Module: Guided by attention layer scores (which account for multimodal information in the diagnosis and treatment data), this module integrates neighbor information and its modality into central nodes. The resulting node representation is a fusion of neighborhood and multimodal data. By embedding data in a low-dimensional space, sensitive information is prevented from being directly exposed during reasoning and processing, thereby improving privacy protection.
[0058] Before describing the present invention, the following description is given:
[0059] The multimodal knowledge graph is an enhanced directed graph G = (N, R, F, M), where N represents the entity set, R represents the relationship set, F represents the set of fact triples, and M represents the multimodal data related to the entity. Each triple in the graph is represented as {h, r, t}, where are the head entity and the tail entity, r∈R as the relationship between them, and M contains multimodal information related to the entities (e.g., text, images, and numerical data).
[0060] The Graph Attention Network (GAT) circumvents the limitations of traditional graph convolutional networks that use fixed weights when weighting neighborhoods. By leveraging the subtle complexity of the attention mechanism, GAT is able to assign different weights to nodes, thereby identifying the differential contributions of different nodes to the central node. The core of the graph attention network lies in its use of attention coefficients. These coefficients basically determine how a node distributes attention among its neighbors, ensuring that different weights are assigned to different nodes. For a pair of adjacent nodes i and j in the graph, the attention coefficient e ij is calculated as follows:
[0061]
[0062] Here, h i and h j denote the feature vectors of nodes i and j, respectively. Matrix W represents the weight matrix, and a represents the learnable weight vector. The symbol || denotes concatenation, and LeakyReLU is the selected nonlinear activation function.
[0063] In order to seamlessly integrate these coefficients into the weighted information aggregation framework, they are normalized by the softmax function:
[0064]
[0065] Here, N(i) represents the neighborhood range of node i.
[0066] Each node integrates its neighborhood information by extracting information from the normalized attention coefficient:
[0067]
[0068] Here, σ represents an activation function, such as the ReLU function.
[0069] In order to enhance the expressive power of the network and introduce a richer set of parameters, the graph attention network adopts a multi-head attention mechanism. Attention calculations are performed in parallel with different parameter sets, and the results are then spliced or averaged. In the field of evolving neural network architectures, Transformer is mainly used for sequence conversion tasks, symbolizing a shift from recursive methods to a new paradigm. Its core mechanism is the self-attention mechanism. This mechanism enables the model to access the importance of different positions in the input sequence without considering their relative positions. The self-attention mechanism first projects the input X into three different subspaces: query (Query, Q), key (Key, K) and value (Value, V):
[0070] Q=XW Q
[0071] K=XW K
[0072] V=XW V
[0073] Among them, W Q 、W K and W V Represent the weight matrices associated with their respective subspaces.
[0074] Subsequently, the attention score calculation is performed, and the alignment of the query and the key produces an attention score, which is normalized by the softmax function to determine the importance of each value in the sequence:
[0075]
[0076] Among them, the normalization factor (d k is the dimension of the key vector) ensures the stability of the calculation.
[0077] The self-attention mechanism perceives information from multiple perspectives, called "heads." Each head projects the input into an independent subspace and computes an attention score. The concatenated output lays the foundation for richer representations:
[0078] MultiHead(Q,K,V)=Concat(head1,head2,...,head h )W O
[0079] For each head i:
[0080] head i =Atten(QW Qi ,KW Ki ,VW Vi )
[0081] Among them, W Qi 、W Ki and W Vi Represents the weight matrix of the i-th head, W O is the output linear layer weight matrix.
[0082] like Figure 1-2 As shown, the present invention provides a privacy-preserving embedding method for multimodal knowledge graphs of medical data, which is implemented as follows:
[0083] S1. Initialize and embed text and numerical information based on diagnosis and treatment data, and initialize and embed visual information based on medical image visual data. The implementation method is as follows:
[0084] S101. Use the pre-trained language model to obtain the contextual information of the sentence-level description of entities and relationships in the diagnosis and treatment data, and generate text information embedding t ;
[0085] S102, extract visual information and embed it into medical image visual data v ;
[0086] S103, encode the entity numerical features in the diagnosis and treatment data and process them in the fully connected layer in turn to generate numerical data embedding e n , complete the initialization embedding of text information, visual information and numerical information.
[0087] In this embodiment, each modality information is learned independently through its own modality-specific pre-trained encoder as follows:
[0088] Text information: Use pre-trained language models to capture the contextual information of sentence-level descriptions of entities and relations in medical data. Specifically, Sentence-BERT is used to generate a 768-dimensional text information embedding. t , reducing the exposure of sensitive information in the text by embedding.
[0089] Visual information: For medical image visual data, the VGG-16 convolutional neural network model is used. By discarding the final fully connected layer and softmax layer, a 4096-dimensional visual information embedding is extracted. v , ensuring that patient privacy information will not be directly reconstructed during visual feature processing.
[0090] Numerical information: Numerical features of entities in the medical data are encoded by BERT and then processed by a fully connected layer. Here, attribute keys and values are concatenated to generate a 768-dimensional numerical data embedding e n ,Effectively protect sensitive numerical information in the representation of numerical data.
[0091] S2. Based on the initial embedding result, the multimodal graph self-attention layer is used to obtain the multimodal attention coefficient of the neighboring entity. The implementation method is as follows:
[0092] S201, using text information to embed e t and numerical data embedded in e n , construct an embedding matrix, and use visual information to embed e v , construct the modality matrix, where each modality graph attention layer processes two embedding matrices and one modality matrix;
[0093] S202. Utilize the multimodal graph self-attention layer to obtain a multimodal attention coefficient.
[0094] In this embodiment, for a vertex e in the multimodal graph structure G, i , learning entity e i (A vertex in the multimodal graph is an entity) multimodal embedding vector h i ′. Construct multiple graph attention layers, each of which processes two embedding matrices and an additional modality matrix. The entity embedding matrix is represented as Among them, the i-th row corresponds to entity e i Embedding, N e represents the total number of entities, and T represents the feature dimension of each entity embedding. Meanwhile, the relation embedding is represented as modal matrix Encapsulate multimodal data. Since entities and relations are mainly composed of text information and numerical information, the embedding matrix of the two is obtained by embedding the text information into e t and numerical data embedded in e n Similarly, the modal matrix is formed by embedding the obtained visual information into e v After being processed by each attention layer, the output matrix is H′, R′, M′.
[0095] In order to i Get the updated embed h i ', learn from the multimodal information matrix of entities, relations and associations mentioned above. The learning depth of each source is determined by the multimodal attention coefficient α ijk Control, for entities (e i ,e j ) and its modality. The input features are converted into higher linear combinations based on specific triples. The local feature connection is transformed, and the multimodal aspect is mathematically defined as follows:
[0096] c jk =W1[h j ||g k ||m k ]
[0097] Among them, c jk Represents the central entity e i In comparison, the tail entity e j and relationship r k The composite entity vector h j and g k Represent the tail entity e j The vector and relationship r k vector, m k Represents multimodal features, matrix represents a learnable weight matrix and || represents a connection operation.
[0098] For fine-grained feature associations between different modalities, the dot product between the central entity embedding and the composite entity embedding is calculated to determine the absolute weight. The weights are normalized by softmax to generate the relative weights of the triplets, taking into account multimodal data. This method captures the relative weights of the triplets iteratively through each attention layer of the graph, as shown in Figure 3 shown.
[0099] The model relies primarily on direct neighbors and the nodes themselves, leveraging their multimodal properties as sources of information (although multiple iterations through multi-hop features are captured by the attention layer entering the graph). The attention parameter b ijk and α ijk Both indicate entity e j For e i The contribution of the triple (e i ,e j ,r k ), these weights are calculated as follows:
[0100]
[0101] Among them, N i Represents entity e i The first-order neighbor set, R in Represents entity e i The relationship between connections and neighbors.
[0102] S3. Based on the multimodal attention coefficient, the structured information of the multimodal knowledge graph is fused by stacking multiple encoder layers. The implementation method is as follows:
[0103] S301, linearly fuse the multimodal attention coefficients to the central node to generate an embedded representation;
[0104] S302. Concatenate the embedding representations generated from M independent self-attention layers:
[0105] S303. Based on the connection results, use the following formula to fuse the structured information of the knowledge graph by stacking multiple encoder layers.
[0106] In this embodiment, these features are linearly fused to the central node by combining the attention coefficients of the neighboring entities calculated previously to generate a more refined embedding vector. Mathematically speaking, the entity e i The updated embedding of is the weighted sum of the triplet embeddings, each of which is determined by its individual contribution. It is expressed as follows:
[0107]
[0108] Among them, x i Represents entity e iThe updated embedding vector, σ represents the nonlinear activation function, and j represents the entity e i The first-order neighbor, N i Represents entity e i The first-order neighbor set, k represents the relationship between neighbor entity j and another node, R jj represents the set of relationships between neighbor entity j and nodes other than neighbor entity j, α ijk represents the attention weight coefficient, which represents the contribution of neighbor node j and relationship k to the center node i, c jk Represents the feature representation of neighbor node j and relationship k.
[0109] This model enhances a single attention layer and uses multiple attention heads to process, so that the encoder can absorb diverse semantic nuances from the context during the embedding learning process. Specifically, it generates embedding representations from M independent self-attention layers and concatenates these embeddings to provide a single vector. Figure 4 As shown, the process can be formalized as:
[0110]
[0111] For each head m :
[0112]
[0113] By stacking multiple encoder layers, each node aggregates data from its multi-hop neighbors, effectively fusing the structured information of the entire knowledge graph and generating node representations that are essentially consistent with the surrounding context.
[0114]
[0115] in, Represents entity e i In the embedding representation of the l+1 layer, σ represents the nonlinear activation function, and R represents the set of relations, that is, all relation types in the graph. Represents entity e i The neighbor set of c i,r Represents the factor used for normalization, adjusting between different relationships, represents the attention weight from neighbor node j to center node i in layer l, Represents a combination function for embedding node i and the relationship between nodes j and i Perform feature transformation, Represents the output after the multi-head self-attention mechanism, MultiHead() represents the multi-head self-attention mechanism, Represents the input of the mth attention head, i.e., entity e i The expression of head m represents the output of the mth attention head, W O represents the matrix of input linear layer weights, j represents entity e i The first-order neighbor, N i Represents entity e i The first-order neighbor set of Represents the neighbor node e j For the central node e i The attention weight, Represents the neighbor node e j The feature representation combined with relation k is used in the mth attention head.
[0116] S4. Based on the structured information of the fused multimodal knowledge graph, the convolutional embedding model is used to score the privacy-preserving embedding of the multimodal knowledge graph.
[0117] In this embodiment, the first three steps have completed the multimodal knowledge graph embedding. The decoding step uses these embeddings to Figure 5 The convolutional embedding model ConvE is shown to evaluate the score of fact triples, thereby enhancing the ability of multimodal knowledge graph completion (KGC). The scoring function is defined as follows:
[0118] f(h,r,t)=σ(vec(W([e h ;r]*ω))e t )
[0119] Among them, f() represents the scoring function of the triple, h represents the head entity, and r represents the connection center entity e i A relationship with neighbor entities, t represents the tail entity, σ represents the sigmoid activation function, vec represents the operation of projecting the feature map into the vector space, W represents the transformation matrix, e h and e t They represent the embedding of the head entity and the tail entity respectively, r represents the embedding vector of the relationship, ω represents the weight matrix of the convolution kernel, and * represents the convolution operation.
[0120] By applying this function, the model evaluates the accuracy of facts based on specific inter-entity relationships. For model optimization, binary cross entropy is used as the loss function, which is expressed as:
[0121]
[0122] in, represents the loss function of the ConvE model, Indicates the number of fact triples, y i'represents the calculated score of the i'th triple, f i' Represents a label, where 1 represents true facts and 0 represents false information.
[0123] In summary, through the above design, the present invention effectively realizes the efficient embedding and deep fusion of multimodal information, enhances the processing capability of heterogeneous data, prevents the leakage of sensitive information through the embedding process, provides an end-to-end privacy protection mechanism, and ensures the security and reliability of medical data during the analysis and reasoning process.
Claims
1. A privacy-preserving embedding method for multimodal knowledge graphs of medical data, characterized by: The following steps are involved: S1. Initialize and embed text information and numerical information based on diagnosis and treatment data, and initialize and embed visual information based on medical image visual data; S2. Based on the initial embedding result, the multimodal graph self-attention layer is used to obtain the multimodal attention coefficient of the neighbor entity, which is specifically: S201, using text information to embed e t and numerical data embedded in e n , construct an embedding matrix, and use visual information to embed e v , construct the modality matrix, where each modality graph attention layer processes two embedding matrices and one modality matrix; S202. Utilize the multimodal graph self-attention layer to obtain a multimodal attention coefficient; S3. According to the multimodal attention coefficient, the structured information of the multimodal knowledge graph is fused by stacking multiple encoder layers. Specifically: S301, linearly fuse the multimodal attention coefficients to the central node to generate an embedded representation; S302. Concatenate the embedding representations generated from M independent self-attention layers using the following formula: in, Represents the output after the multi-head self-attention mechanism, MultiHead() represents the multi-head self-attention mechanism, Represents the input of the mth attention head, i.e., entity e i The expression of head m represents the output of the mth attention head, W O Represents the matrix of input linear layer weights, j represents the entity e i The first-order neighbor, N i Represents entity e i The first-order neighbor set of Represents the neighbor node e j For the central node e i The attention weight, Represents the neighbor node e j The feature representation combined with relation k is used in the mth attention head; S303. Based on the connection results, the following formula is used to fuse the structured information of the knowledge graph by stacking multiple encoder layers: in, Represents entity e i In the embedding representation of the l+1 layer, σ represents a nonlinear activation function, and R represents the set of relations, that is, all relation types in the knowledge graph. Represents entity e i The neighbor set of c i,r Represents the factor used for normalization, adjusting between different relationships, represents the attention weight from neighbor node j to center node i in layer l, Represents a combination function for embedding node i and the relationship between nodes j and i Perform feature transformation; S4. Based on the structured information of the fused multimodal knowledge graph, the convolutional embedding model is used to score the privacy-preserving embedding of the multimodal knowledge graph.
2. The privacy-preserving embedding method for multimodal knowledge graphs of medical data according to claim 1 is characterized in that: The step S1 comprises the following steps: S101. Use the pre-trained language model to obtain the contextual information of the sentence-level description of entities and relationships in the diagnosis and treatment data, and generate text information embedding t ; S102, extract visual information and embed it into medical image visual data v ; S103, encode the entity numerical features in the diagnosis and treatment data and process them in the fully connected layer in turn to generate numerical data embedding e n , complete the initialization embedding of text information, visual information and numerical information.
3. The privacy-preserving embedding method for multimodal knowledge graphs of medical data according to claim 2 is characterized in that: The expression of the multimodal attention coefficient is as follows: c jk =W1[h j ||g k ||m k ] Among them, b ijk and Both represent multimodal attention coefficients, c i Represents the central entity e i The eigenvector of jk Represents the central entity e i In comparison, the tail entity e j and relationship r k Composite entity vector, d k Indicates the dimension of the vector, used to normalize the result of the vector dot product, exp() represents the natural exponential function, used to convert b ijk The value of b is mapped to the positive range, inr Represents entity e i The absolute attention parameter between the neighbor entity n through the relation r, n represents the entity e i A neighbor node, N i Represents entity e i The first-order neighbor set, r represents the connection center entity e i A relation with a neighbor entity, R in Represents entity e i The relationship between the connection and the neighbors, W1 represents the learnable weight matrix, ‖ represents the connection operation, h j and g k Represent the tail entity e j The vector and relationship r k vector, m k Represents multimodal features.
4. The privacy-preserving embedding method for multimodal knowledge graphs of medical data according to claim 3 is characterized in that: The expression of the embedding vector is as follows: Among them, x i Represents entity e i The updated embedding vector, σ represents the nonlinear activation function, and j represents the entity e i The first-order neighbor, N i Represents entity e i The first-order neighbor set, k represents the relationship between neighbor entity j and another node, R jj represents the set of relationships between neighbor entity j and nodes other than neighbor entity j, α ijk represents the attention weight coefficient, which represents the contribution of neighbor node j and relationship k to the center node i, c jk Represents the feature representation of neighbor node j and relationship k.
5. The privacy-preserving embedding method for multimodal knowledge graphs of medical data according to claim 1 is characterized in that: The scoring function of the convolutional embedding model is expressed as follows: f(h,r,t)=σ(vec(W([e h ;r]*ω))e t ) Among them, f() represents the scoring function of the triple, h represents the head entity, and r represents the connection center entity e i A relationship with neighbor entities, t represents the tail entity, σ represents the sigmoid activation function, vec represents the operation of projecting the feature map into the vector space, W represents the transformation matrix, e h and e t They represent the embedding of the head entity and the tail entity respectively, r represents the embedding vector of the relationship, ω represents the weight matrix of the convolution kernel, and * represents the convolution operation.
6. The privacy-preserving embedding method for multimodal knowledge graphs of medical data according to claim 1 is characterized in that: The loss function of the convolutional embedding model is expressed as follows: Among them, L(f,y) represents the loss function of the ConvE model, |F| represents the number of fact triplets, and y i' represents the calculated score of the i'th triple, f i' Represents a label, where 1 represents true facts and 0 represents false information.
Citation Information
Patent Citations
Medical image report generation method and device based on multi-modal fusion
CN115331769A
Data processing method, device and equipment
CN117726459A