Link Prediction Method Based on the Representation of Entities and Relationships by Fusing Multimodal Information
By extracting single-modal features of images and text and fusing multimodal information, the noise problem in multimodal knowledge graph construction is solved, the accuracy and interpretability of link prediction are improved, and a more stable multimodal knowledge representation is achieved.
Patent Information
- Application Number
- CN202310641906.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-01
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2043-06-01
AI Technical Summary
The existing multimodal knowledge graph construction method has noise problems when fusing multimodal data, ignoring the internal structural information of the knowledge graph, resulting in the lack of interpretability of the inference process and the lack of full use of the relationship between the image and the entity.
The single-modal features of images and text are extracted respectively through the visual module and the text module, and visual representations and text representations are generated. The entity and relationship vector representations containing multimodal information are learned through the fusion module, and link prediction is combined with the decoding part to optimize the multimodal knowledge representation.
It improves the accuracy of link prediction and interpretability of multimodal knowledge representation, enhances the robustness of the model, and can effectively handle the noise impact caused by image data.
Smart Images

Figure CN116680343B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of knowledge graph knowledge reasoning, and more specifically, to a link prediction method based on entity and relationship representation by fusing multi-modal information. Background Art
[0002] Traditional multi-modal knowledge graph construction methods relying on manual labor cannot well contain all modal knowledge. At the same time, there is noise information between different modalities, resulting in sparse graphs and possibly incorrect triples. This requires research on multi-modal knowledge reasoning technology. The reasoning technology for multi-modal knowledge graphs can mainly assist in inferring new facts, relationships, axioms, and rules, etc. More importantly, it predicts the missing parts of triples. Link prediction is an important task in knowledge reasoning, and its goal is to predict the missing entities or potential relationships in the knowledge graph, thereby enriching and improving the knowledge graph.
[0003] With the development of knowledge graph reasoning, many methods have emerged in recent years, including reasoning based on graph structure and rules, reasoning based on knowledge graph representation learning, reasoning based on neural networks, and hybrid reasoning, etc. These reasoning methods have their own advantages and disadvantages. The construction of multi-modal knowledge graphs is a very complex process, which requires extracting entities, attributes, and relationships from a large amount of structured, semi-structured, and unstructured data. Link prediction based on entity and relationship representation by fusing multi-modal information can be regarded as a sub-problem of knowledge reasoning in multi-modal knowledge graphs. Currently, the knowledge reasoning methods for multi-modal knowledge graphs mainly use techniques based on multi-modal knowledge representation learning, but there are the following defects when using the knowledge reasoning methods based on multi-modal knowledge representation learning to solve the above problems:
[0004] (1) After introducing multi-modal data, there will inevitably be related quality problems, without considering the complexity of each modal data itself and the noise brought by information interaction between them.
[0005] (2) Existing methods focus more on how to obtain more auxiliary information from these two major types of multi-modal data, such as images and texts, while ignoring the internal structure information of the knowledge graph, resulting in a lack of certain interpretability in the reasoning process.
[0006] (3) Since the relationship between most entities and images is one-to-many, that is, one entity corresponds to multiple different images related to the entity, the existing technology does not consider how to better utilize the information in these images. Summary of the Invention
[0007] In view of this, the present invention provides a link prediction method based on the representation of entities and relationships by fusing multi-modal information, which can improve the accuracy of the link prediction task and enhance the interpretability of multi-modal knowledge representation learning.
[0008] To achieve the above object, the present invention adopts the following technical solutions:
[0009] A link prediction method based on the representation of entities and relationships by fusing multi-modal information, comprising the following steps:
[0010] Collect multi-modal data related to the topic of the knowledge graph to be constructed and perform preprocessing; the multi-modal data includes image data, text data, and triple data;
[0011] Perform entity recognition, relationship extraction, and attribute extraction on the preprocessed triple data, and align the image data and text data related to the entities in the triples;
[0012] Extract single-modal features from the image data through the visual module, learn the key features related to the entities in the image data, and generate visual representations;
[0013] Extract single-modal features from the text data and triple data through the text module to generate text representations that can reflect semantic information;
[0014] Use the generated visual representations, text representations, and triple data as inputs together to train the fusion module and learn the vector representations of entities and relationships containing multi-modal information;
[0015] Decode the feature representations learned by the fusion module through the decoding part and perform link prediction, and output the probability of predicting positive triples.
[0016] Further, the preprocessing includes: using open-source tools to perform data cleaning, data conversion, and data integration operations on the image data, text data, and triple data respectively.
[0017] Further, after aligning the image data and text data related to the entities in the triples, use the triples and the corresponding image data and text data of the entities as the multi-modal total dataset, and randomly divide the multi-modal total dataset into a training set, a validation set, and a test set.
[0018] Further, the visual module includes an input, a selection module, and a visual encoder; the process of extracting features from the image data through the visual module is as follows:
[0019] Take multiple image data corresponding to the same entity as input, send them into the optimization module to obtain the optimal one image data, and divide the optimal image data into small-sized blocks as the input of the visual encoder, and output a visual representation through the visual encoder.
[0020] Further, in the optimization module, irrelevant images and low-quality images are filtered out through two steps of similarity calculation and clarity evaluation, and a relatively optimal image is retained as the subsequent input; specifically including:
[0021] Use the perceptual hashing algorithm to calculate the similarity of multiple images corresponding to the same entity, calculate the Hamming distance to obtain the similarity between images, and filter out images with too high similarity and irrelevant images;
[0022] Use the sum of absolute values of gray-level differences function for clarity evaluation. By taking differences between adjacent pixels in the horizontal and vertical directions of the image, taking the absolute value and then accumulating, use the accumulated value as the representation of image clarity, filter out the image with the optimal clarity, and divide the image into small-sized blocks as the input vector of the visual encoder;
[0023] The visual encoder adopts the encoder structure in the Transformer architecture, and the specific encoding process is:
[0024] The input vector first passes through the multi-head attention layer, residual connection and layer normalization operation, and then passes through the feed-forward neural network, residual connection and layer normalization operation to obtain the visual representation.
[0025] Further, if there is only one corresponding image data for the same entity, directly use the image data as the input of the visual encoder;
[0026] If there are zero corresponding image data for the same entity, fill all the inputs of the visual encoder with 0.
[0027] Further, the process of text feature extraction through the text module is:
[0028] Split the text sequence in the text data into multiple sentences, add a "[CLS]" token at the beginning of the entire text sequence, and add "[SEP]" tokens in the middle of two sentences and at the end of the entire sequence;
[0029] Convert the sequential splicing of triple data into a text sequence through the "[CLS]" token and "[SEP]" tokens;
[0030] Use text embedding, position encoding and token encoding together as the input vector and send it into the text encoder;
[0031] The text encoder adopts the encoder structure in the Transformer architecture. The input vector first passes through the multi-head attention layer, residual connection, and layer normalization operation, and then through the feed-forward neural network, residual connection, and layer normalization operation to obtain the text representation.
[0032] Furthermore, the fusion module consists of three parts: a visual fusion encoder, a central encoder, and a text auxiliary encoder;
[0033] The text auxiliary encoder adopts the encoder structure of Transformer. The input text representation first passes through the multi-head attention layer, residual connection, and layer normalization operation, and then through the feed-forward neural network, residual connection, and layer normalization operation to output the feature vector of the text;
[0034] The visual fusion encoder adopts the encoder structure of Transformer. The input image representation first passes through the multi-head attention layer, residual connection, and layer normalization operation, and then through the feed-forward neural network, residual connection, and layer normalization operation to finally output the feature vector of the image;
[0035] The central encoder adopts the encoder structure in the Transformer architecture. The text feature vector output by the text auxiliary encoder and the image feature vector output by the visual fusion encoder first pass through the multi-modal attention layer, residual connection, and layer normalization operation, and then through the feed-forward neural network, residual connection, and layer normalization operation to output the entity and relationship vector representation.
[0036] Furthermore, the operation process of the multi-modal attention layer is as follows:
[0037] Calculate the text modality attention value;
[0038] Calculate the visual modality attention value;
[0039] For each triple representation input, first obtain the initial representation of the triple structure through the linear transformation matrix, and then pass through the LeakyRelu non-linear layer and then through the Softmax layer to obtain the triple attention value;
[0040] Take the text modality attention value, visual modality attention value, and triple attention value as the weights of the text feature vector, visual feature vector, and triple structure feature vector respectively, and perform weighted summation and averaging to obtain the multi-modal representation of the entity;
[0041] Perform weighted summation on the multi-modal representation of the entity and the original feature representation of the entity to obtain the final vector representation of the entity;
[0042] For the multi-modal representation of the relationship between entities, take the triple structure information as the multi-modal representation, and perform weighted summation on the original representation of the relationship and the multi-modal representation to obtain the final vector representation of the relationship;
[0043] Use the vector representations of entities and relationships as the output of the fusion module.
[0044] Furthermore, the decoding part uses the decoder in the Transformer network architecture, and the decoding process is as follows:
[0045] Input the target output sequence, and perform masked multi-head attention layer, residual connection, and layer normalization operations;
[0046] Use the output vector of the previous layer and the output of the fusion module as the input vector together, and perform multi-head attention layer, residual connection, and layer normalization operations;
[0047] Pass the output vector of the previous layer through a feed-forward neural network, residual connection, and layer normalization operations;
[0048] Pass the output vector of the previous layer through one fully connected layer and one softmax layer to obtain the final probability output result as the output of the decoding part.
[0049] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses a link prediction method based on the representation of entities and relationships by fusing multi-modal information. During the model training process, the idea of a translation model is introduced to constrain the distances between the head entity, relationship, and tail entity. By optimizing the representation of multi-modal knowledge, the accuracy of link prediction is improved, the interpretability of multi-modal knowledge representation learning can be enhanced, and it has a certain robustness in practical applications while ensuring the final effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to the provided drawings.
[0051] Figure 1 It is a structural schematic diagram provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0052] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0053] Such as Figure 1As shown in the figure, an embodiment of the present invention discloses a link prediction method based on the representation of entities and relationships by fusing multi-modal information, including the following steps:
[0054] S1. Collect multi-modal data related to the topic of the knowledge graph to be constructed and perform preprocessing; the multi-modal data includes image data, text data, and triple data;
[0055] S2. Perform entity recognition, relationship extraction, and attribute extraction on the preprocessed triple data, and align the image data and text data related to the entities in the triple with it;
[0056] S3. Extract single-modal features from the image data through the visual module, learn the key features related to the entities in the image data, and generate a visual representation;
[0057] S4. Extract single-modal features from the text data and triple data through the text module to generate a text representation that can reflect semantic information;
[0058] S5. Use the generated visual representation, text representation, and triple data as inputs together to train the fusion module to learn the entity and relationship vector representations containing multi-modal information;
[0059] S6. Decode the feature representation learned by the fusion module through the decoding part and perform link prediction to output the probability of predicting a positive triple.
[0060] The above steps will be further described below.
[0061] In S1, collect a large amount of multi-modal data related to a certain news event, including text data, triple data, and image data. Specifically, multi-modal data related to the required news event can be found from public websites and the GDELT database, including but not limited to news image data, news title data, image description data, CSV data (tabular data), etc.
[0062] After collecting a large amount of multi-modal data, perform preprocessing on it, including: using open-source tools to perform data cleaning, data conversion, and data integration operations on the image data, text data, and triple data respectively to obtain high-quality multi-modal data.
[0063] Specifically: The image data is preprocessed by directly compressing it to the input size of the network, 224*224, without maintaining the original aspect ratio of the image. The re.sub() method in Python is used to remove invalid information from the text data, and invalid fields and completely identical duplicate data are deleted from the structured data. For special symbols and invalid characters in the text data and triple data, regular expressions and string replacement methods are used to replace or delete them. At the same time, the categorical information existing in numerical form is converted into the corresponding text form. Finally, the preprocessed data information is integrated into a CSV file.
[0064] In a specific embodiment, in S2, the DeepKE tool can be used to perform named entity recognition, relation extraction, and attribute extraction on the triple-structured data in the preprocessed high-quality news event multi-modal dataset, and align the relevant image data and text data with it.
[0065] After aligning the image data and text data related to the entities in the triple with it, the triple and the corresponding image data and text data of the entities in it are jointly used as the multi-modal total dataset, and the multi-modal total dataset is randomly divided into a training set, a validation set, and a test set. Among them, the training set accounts for 80% of the total dataset, the validation set accounts for approximately 10% of the total dataset, and the test set accounts for approximately 10% of the total dataset.
[0066] In a specific embodiment, in S3, the visual module includes an input, a selection module, and a visual encoder; the process of extracting features from the image data through the visual module is as follows:
[0067] Multiple image data corresponding to the same entity are used as inputs and sent to the selection module to obtain the optimal one image data, and the optimal image data is divided into small-sized blocks as the input of the visual encoder, and the visual representation is output through the visual encoder.
[0068] Specifically, since each entity corresponds to more than one image and the image quality is uneven, in the selection module, irrelevant images and low-quality images are filtered out through two steps of similarity calculation and clarity evaluation, and a relatively optimal image is retained as the subsequent input, which can reduce the noise impact brought by low-quality image data; specifically, it includes:
[0069] First, the perceptual hashing algorithm is used to calculate the similarity of multiple images corresponding to the same entity. The image is transformed from the pixel domain to the frequency domain through the discrete cosine transform. Since the human eye is not sensitive to high-frequency detail information, the image content can usually be recognized only based on the low-frequency information. Therefore, only the 8*8 low-frequency region in the upper left corner, a total of 64 pixels, is selected. The discrete cosine transform values of each pixel are calculated respectively, and the average value is obtained by averaging all 64 values. Then, the hash value is obtained by comparing the size of each discrete cosine transform value with the average value. The calculation method of the hash value is to set a 64-bit hash value of 0 or 1, where those greater than the average value are set to 1 and those less than the average value are set to 0. Then, the similarity between images is obtained by calculating the Hamming distance, that is, calculating the number of different characters at the corresponding positions of the two hash values. If the proportion of the same number of bits is greater than 0.92, it means that the two images are very similar. If it is less than 0.84, it means that these are two different images, so as to filter out images with too high similarity and completely irrelevant images.
[0070] The specific calculation method of the discrete cosine transform value is as shown in the formula:
[0071]
[0072]
[0073]
[0074] Where N refers to the total number of one-dimensional data elements, which takes the value of 8 in this example. f(i, j) is the element at the (i, j) position of the input data. The coefficients c(u) and c(v) are used to convert the discrete cosine transform matrix into an orthogonal matrix; u and v represent the coordinates in the frequency domain.
[0075] The sum of absolute differences of grayscale function is used for clarity evaluation. By taking the differences of adjacent pixels in the horizontal and vertical directions of the image, taking the absolute values and then accumulating them, the accumulated value is used as the representation of the image clarity, and the image with the optimal clarity is selected. This image is divided into small-sized blocks (such as 14*14 small blocks) as the input vector of the visual encoder;
[0076] If there is only one corresponding image data for the same entity, then this image data is directly used as the input of the visual encoder;
[0077] If there are zero corresponding image data for the same entity, then all the inputs of the visual encoder are filled with 0.
[0078] The visual encoder adopts the encoder structure in the Transformer architecture. The specific encoding process is as follows:
[0079] In the visual encoder, the input vector first passes through a multi-head attention layer, a residual connection, and a layer normalization operation, and then passes through a feed-forward neural network, a residual connection, and a layer normalization operation to obtain an output vector. Through six identical visual encoders like this, a visual representation is finally obtained.
[0080] In the residual connection and layer normalization operation: The residual connection means adding the input and output of the network, that is, the output of the network is F(x)+x. When the network structure is relatively deep, during the backpropagation of the network gradient to update the parameters, it is easy to cause the problem of gradient disappearance. However, if x is added to the output of each layer, it becomes F(x)+x, and the derivative of x is 1. So it is equivalent to adding a constant term '1' to the derivative of each layer, effectively solving the problem of gradient disappearance. Layer normalization means calculating the average value of the features of each token separately and normalizing the output into a standard normal distribution to keep the input of the next layer relatively stable.
[0081] Among them, the visual encoder uses the cross-entropy loss function as the objective function for training, and continuously adjusts and optimizes the parameters of the visual encoder during the training process until the objective function converges.
[0082] In one embodiment, in S4, the process of text feature extraction by the text module is as follows:
[0083] The method in BERT is used to split the text sequence in the text data into multiple sentences, add a "[CLS]" token at the beginning of the entire text sequence, and add "[SEP]" tokens in the middle of two sentences and at the end of the entire sequence;
[0084] For triple data, the triple data is sequentially concatenated and converted into a text sequence through the "[CLS]" token and the "[SEP]" token; "[SEP]" tokens are added in the middle of the head entity, relation, and tail entity. At this time, the token encoding calculation method is expanded from the original "the values of all tokens in the first sentence (including the "[CLS]" token and the "[SEP]" token following the first sentence) are 0, and the values of all tokens in the second sentence (including the "[SEP]" token following the second sentence) are 1" to: "the values of all tokens in the first sentence are 0, the values of all tokens in the second sentence are 1, and the values of all tokens in the third sentence are 0".
[0085] The text embedding, position encoding, and token encoding together serve as the input vector and are fed into the text encoder;
[0086] The text encoder adopts the encoder structure in the Transformer architecture. In the encoder, the input vector first passes through the multi-head attention layer, residual connection, and layer normalization operation, and then through the feed-forward neural network, residual connection, and layer normalization operation. After passing through 6 identical text encoders, the text representation is finally obtained.
[0087] Among them, the text encoder uses the cross-entropy loss function as the objective function for training, and continuously adjusts and optimizes the text encoder parameters during the training process until the objective function converges.
[0088] In one embodiment, in S5,
[0089] The fusion module consists of three parts: a visual fusion encoder, a central encoder, and a text auxiliary encoder. S3 and S4 are for coarse-grained feature extraction, and the representations of the obtained images, texts, and triples are used as inputs to train the fusion module, and the low-dimensional vector representations of entities and relationships are used as outputs; the visual fusion encoder and text auxiliary encoder in the fusion module further extract the fusion features of the previously extracted coarse-grained features.
[0090] The input text representation passes through the text auxiliary encoder. The text auxiliary encoder adopts the encoder structure of Transformer. First, it passes through the multi-head attention layer, residual connection, and layer normalization operation, and then through the feed-forward neural network, residual connection, and layer normalization operation to output the feature vector of the text; the text auxiliary attention is as follows:
[0091]
[0092] Among them, x t is the input text representation vector; and are weight matrices.
[0093] The input image representation passes through the visual fusion encoder. The visual fusion encoder adopts the encoder structure of Transformer. First, it passes through the multi-head attention layer, residual connection, and layer normalization operation. When calculating the multi-head attention layer, the PGI method in the MKGformer model is used to transfer the matrix K and matrix V of the standard attention in the text auxiliary encoder to the visual fusion encoder, and then through the feed-forward neural network, residual connection, and layer normalization operation, and finally outputs the feature vector of the image. The visual attention is:
[0094]
[0095] Among them, x v is the input image representation vector; is and are weight matrices.
[0096] The central encoder also uses the encoder structure in the Transformer architecture. First, through the multi-modal attention layer, residual connection, and layer normalization operation, and then through the feed-forward neural network, residual connection, and layer normalization operation, the maximum margin loss function is used as the objective function for training. During the training process, the encoder parameters are continuously adjusted and optimized until the objective function converges, and the representation vectors of entities and relationships are output.
[0097] Among them, the maximum margin loss function learns to generate higher-quality representations by maximizing the margin between positive and negative samples. Its basic idea is: for a triple <h, l, t> in a knowledge graph, it is trained and optimized to make its score higher, while making the score of triples outside the knowledge graph lower, maximizing the distance between them to achieve the purpose of optimizing knowledge representation. Specifically:
[0098]
[0099] S′ h,l,t ={(h′, l, t | h′ ∈ E)} ∪ {(h, l, t′ | t′ ∈ E)}
[0100] Among them, S is the correct triple, S′ is the incorrect triple obtained by replacing h or t, γ is the margin distance hyperparameter, and [x]+ is the positive value function, that is, when x > 0, [x]+ = x; when x ≤ 0, [x]+ = 0.
[0101] The operation process of the multi-modal attention layer is as follows:
[0102] Calculate the text-modal attention head t , taking the output of the text auxiliary encoder as the input;
[0103]
[0104] Among them, Qv, Kv, and Vv are the matrices corresponding to the query vector, key vector, and value vector respectively, and W Q , WK, and VV are all weight matrices.
[0105] Calculate the visual-modal attention head v , taking the output of the visual fusion encoder as the input;
[0106]
[0107] Among them, Qt, Kt, and Vt are the matrices corresponding to the query vector, key vector, and value vector respectively, and W Q , WK, and VV are all weight matrices.
[0108] For each triple representation <t headentityi , t relation j , t tailentity k (>, first obtain the initial representation head of the triple structure through the linear transformation matrix 0 ijk , and obtain head after passing through the LeakyRelu non-linear layer 1 ijk , then pass through the Softmax layer to obtain the triple attention value head ijk ;
[0109]
[0110]
[0111]
[0112] Among them, W1 and W2 are linear transformation matrices.
[0113] Take the text modality attention value, visual modality attention value, and triple attention value as the weights of the text feature vector, visual feature vector, and triple structure feature vector respectively, and perform weighted summation and averaging to obtain the multi-modal representation E of the entity Multi ;
[0114] Perform weighted summation on the multi-modal representation of the entity and the original feature representation of the entity to obtain the final vector representation E of the entity;
[0115]
[0116] E = αt entity +(1 - α)E Multi , 0 < α < 1
[0117] For the multi-modal representation of the relationship between entities, take the triple structure information as the multi-modal representation R Multi , perform weighted summation on the original representation of the relationship and the multi-modal representation to obtain the final vector representation R of the relationship;
[0118]
[0119] R = βt relation +(1 - β)R Multi , 0 < β < 1
[0120] Among them, where M is the number of modalities, and the value is 3; is the visual representation vector; is the text representation vector; t entity is the vector representation of the entity; t relationThe vector representation of the relationship; α and β are weight parameters, randomly initialized at the beginning of training, flexibly allocating the weights of multi-modal information in the representations of entities and relationships, and achieving optimization during the training process.
[0121] Use the vector representations of entities and relationships as the output of the fusion module.
[0122] In one embodiment, the decoding part uses the decoder in the Transformer network architecture, and the decoding process is as follows:
[0123] Input the target output sequence, and perform operations through the masked multi-head attention layer, residual connection, and layer normalization;
[0124] Use the output vector of the previous layer and the output of the fusion module as the input vector together, and perform operations through the multi-head attention layer, residual connection, and layer normalization;
[0125] Pass the output vector of the previous layer through the feed-forward neural network, residual connection, and layer normalization;
[0126] Pass the output vector of the previous layer through 1 fully connected layer and 1 softmax layer to obtain the final probability output result as the output of the decoding part.
[0127] The embodiment of the present invention also verifies the above method as follows.
[0128] The hardware used in the present invention is CPU: Intel(R) Xeon(R) Gold 5218 CPU @ 2.30GHz, GPU: GeForce RTX 3090 Ti, video memory capacity 24GB, memory: 128576MB. The software is: operating system: Linux Ubuntu 64-bit, CUDA (11.1), cudnn (8.0), Python (3.6), Python (3.8); use Hits@10, Mean Reciprocal Rank (MRR), and Mean Rank (MR) as evaluation metrics for link prediction of the entity and relationship representations that fuse multi-modal information. Hits@10 is the average proportion of correct triples among the top 10 prediction results, and the larger this indicator, the better; MRR is the average reciprocal rank of correct triples, and the larger this indicator, the better; MR is the average rank of correct triples, and the smaller this indicator, the better.
[0129] The prediction method of the present invention and the existing methods are used to test a set of test data, and the test results of each method are shown in Table 1 below.
[0130] Table 1 Comparison of test results of various methods
[0131]
[0132] As can be seen from the above table, compared with the existing methods, under the condition that the final effects are not much different, the method of the present invention is superior to the existing methods in all evaluation indexes of link prediction. Therefore, the method of the present invention can improve the accuracy of the link prediction task. Compared with some existing methods, in the method of the present invention, the central encoder adopts the maximum margin loss function, by minimizing the distance between the head entity vector + relation vector and the tail entity vector of the triples in the knowledge graph, while for the triples outside the knowledge graph, the distance between them is maximized, and the vector representations of the head entity, the tail entity and the relationship between them are constrained to satisfy the idea of the translation model, that is, regarding the relationship as the translation from the head entity to the tail entity, so as to improve the interpretability of multi-modal knowledge representation learning. Compared with some existing methods, when there is too much image data, it is screened by the optimization module, and when there is missing image data, the subsequent steps are carried out by filling 0. Therefore, the method of the present invention can well solve the problem that the model effect is poor due to the noise influence brought by the image data, so the method of the present invention can have a certain robustness on the premise of ensuring the final effect.
[0133] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same and similar parts among the embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and the relevant parts can be referred to the description of the method part.
[0134] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A link prediction method based on entity and relationship representation by fusing multi-modal information, characterized in that It includes the following steps: Collect multi-modal data related to the topic of the knowledge graph to be constructed and perform preprocessing; the multi-modal data includes image data, text data, and triple data; Perform entity recognition, relation extraction, and attribute extraction on the preprocessed triple data, and align the image data and text data related to the entities in the triples; Extract single-modal features from the image data through the visual module, learn the key features related to the entities in the image data, and generate visual representations; Extract single-modal features from the text data and triple data through the text module to generate text representations that can reflect semantic information; Use the generated visual representations, text representations, and triple data as inputs together to train the fusion module and learn the entity and relation vector representations containing multi-modal information; Decode the feature representations learned by the fusion module through the decoding part and perform link prediction, and output the probability of predicting positive triples; The visual module includes an input, a selection module, and a visual encoder; the process of extracting features from the image data through the visual module is as follows: Take multiple image data corresponding to the same entity as inputs, send them into the selection module to obtain the optimal one image data, and divide the optimal image data into small-sized blocks as the inputs of the visual encoder, and output visual representations through the visual encoder; In the selection module, irrelevant images and low-quality images are filtered out through two steps of similarity calculation and clarity evaluation, and a relatively optimal image is retained as the subsequent input; specifically including: Use the perceptual hashing algorithm to calculate the similarity of multiple images corresponding to the same entity, obtain the similarity between images by calculating the Hamming distance, and filter out images with too high similarity and irrelevant images; Use the sum of absolute differences of grayscale function for clarity evaluation. Make differences between adjacent pixels in the horizontal and vertical directions of the image, take the absolute values and then accumulate them. Use the accumulated value as the characterization of the image clarity, select the image with the optimal clarity, and divide the image into small-sized blocks as the input vectors of the visual encoder; The visual encoder adopts the encoder structure in the Transformer architecture, and the specific encoding process is as follows: The input vectors first go through the multi-head attention layer and the residual connection and layer operation, and then go through the feed-forward neural network and the residual connection and layer normalization operation to obtain visual representations.
2. The link prediction method based on the representation of entities and relationships by fusing multi-modal information according to claim 1, wherein The preprocessing includes: using open-source tools to perform data cleaning, data transformation, and data integration operations on the image data, text data, and triple data respectively.
3. The link prediction method for entity and relationship representation based on fused multi-modal information according to claim 1, wherein After aligning the image data and text data related to the entities in the triples, take the triples and the image data and text data corresponding to the entities in them as the multi-modal total dataset, and randomly divide the multi-modal total dataset into a training set, a validation set, and a test set.
4. The link prediction method based on entity and relationship representation by fusing multi-modal information according to claim 1, characterized in that, If there is only one corresponding image data for the same entity, directly use the image data as the input of the visual encoder; If there are zero corresponding image data for the same entity, fill all the inputs of the visual encoder with 0.
5. The link prediction method for entity and relationship representation based on fused multi-modal information according to claim 1, characterized in that, The process of extracting text features through the text module is as follows: Split the text sequence in the text data into multiple sentences, add a "[CLS]" token at the beginning of the entire text sequence, and add "[SEP]" tokens in the middle of two sentences and at the end of the entire sequence; Convert and splice the triple data in order into a text sequence through the "[CLS]" token and the "[SEP]" token; Use the text embedding, position encoding, and token encoding together as the input vector and feed it into the text encoder; The text encoder adopts the encoder structure in the Transformer architecture. The input vector first passes through the multi-head attention layer, residual connection, and layer normalization operation, and then passes through the feed-forward neural network, residual connection, and layer normalization operation to obtain the text representation.
6. The link prediction method based on entity and relationship representation by fusing multi-modal information according to claim 1, characterized in that: The fusion module consists of three parts: a visual fusion encoder, a central encoder, and a text auxiliary encoder; The text auxiliary encoder adopts the encoder structure of Transformer. The input text representation first passes through the multi-head attention layer, residual connection, and layer normalization operation, and then passes through the feed-forward neural network, residual connection, and layer normalization operation to output the feature vector of the text; The visual fusion encoder adopts the encoder structure of Transformer. The input image representation first passes through the multi-head attention layer, residual connection, and layer normalization operation, and then passes through the feed-forward neural network, residual connection, and layer normalization operation to finally output the feature vector of the image; The central encoder adopts the encoder structure in the Transformer architecture. The text feature vector output by the text auxiliary encoder and the image feature vector output by the visual fusion encoder first pass through the multi-modal attention layer, residual connection, and layer normalization operation, and then pass through the feed-forward neural network, residual connection, and layer normalization operation to output the entity and relationship vector representation.
7. The link prediction method for entity and relationship representation based on fused multi-modal information according to claim 1, characterized in that: The operation process of the multi-modal attention layer is as follows: Calculate the text modality attention value; Calculate the visual modality attention value; For each triple representation input, first obtain the initial representation of the triple structure through the linear transformation matrix, and then pass through the LeakyRelu non-linear layer and then through the Softmax layer to obtain the triple attention value; Take the text modality attention value, visual modality attention value, and triple attention value as the weights of the text feature vector, visual feature vector, and triple structure feature vector respectively, and perform weighted summation and averaging to obtain the multi-modal representation of the entity; Perform weighted summation on the multi-modal representation of the entity and the original feature representation of the entity to obtain the final vector representation of the entity; For the multi-modal representation of the relationship between entities, take the triple structure information as the multi-modal representation, and perform weighted summation on the original representation of the relationship and the multi-modal representation to obtain the final vector representation of the relationship; Take the vector representations of the entity and the relationship as the output of the fusion module.
8. The link prediction method based on entity and relationship representation by fusing multi-modal information according to claim 1, characterized in that The decoding part uses the decoder in the Transformer network architecture, and the decoding process is as follows: Input the target output sequence and perform the masked multi-head attention layer, residual connection, and layer normalization operation; Take the output vector of the previous layer and the output of the fusion module together as the input vector and perform the multi-head attention layer, residual connection, and layer normalization operation; Pass the output vector of the previous layer through a feed-forward neural network, a residual connection, and a layer normalization operation; Pass the output vector of the previous layer through a fully-connected layer and a softmax layer to obtain the final probability output result as the output of the decoding part.
Citation Information
Patent Citations
Multimodal knowledge representation method fusing entity image information and entity category information
CN113486190A
Multi-source information fusion enhanced knowledge graph representation learning method
CN115563314A