Text-image enhanced multi-modal knowledge graph embedding method
Through image filtering and deep learning technology, the quality of solid image data sets is improved, and combined with BERT and 2D modal neural network models, the problems of insufficient single-modal representation and poor data quality in knowledge graph embedding are solved, and more comprehensive entity representation and stronger model expression capabilities are achieved.
Patent Information
- Application Number
- CN202510269332.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-03-07
AI Technical Summary
There are problems in the existing knowledge graph embedding methods such as insufficient single-modal representation, limited expression ability of shallow model, and poor multimodal data quality.
The quality of the solid image data set is improved through the image filtering algorithm, the deep residual network ResNet is used for image feature extraction, text features are obtained in combination with the BERT model, and the 2D modal neural network model is used to fuse multimodal features for knowledge graph embedding.
It improves the comprehensiveness and accuracy of entity representation, enhances the expression ability of the model, solves the problem of data sparseness, and improves the representation ability of the knowledge graph.
Smart Images

Figure CN120235225A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of knowledge graph embedding, and in particular to a multi-modal knowledge graph embedding method with text-image enhancement. Background Art
[0002] With the rapid development of artificial intelligence and big data technologies, knowledge graphs, as a way of constructing knowledge representation, have received extensive attention. Knowledge Representation Learning (KRL) technology provides important support for the application of knowledge graphs by transforming the triples in the knowledge graph into an empowered form and embedding entities into a low-dimensional operation space.
[0003] However, the existing knowledge relation representation learning methods have limitations in multiple aspects: First, although traditional translation models such as TransE and TransH can handle one-to-one relations, their performance significantly degrades when dealing with complex relations such as one-to-many and many-to-one, which seriously affects the representation ability of knowledge graphs; Second, currently mainstream knowledge graph embedding models mainly use a single mode for feature learning, such as only using text information for entity representation. This method ignores the rich semantic information contained in other entity modalities (such as images and audio), resulting in incomplete and inaccurate entity representations; Third, although the existing shallow neural network models have advantages in efficiency, their limited expressive ability is difficult to capture the complex mathematical associations between entities and relations, especially when dealing with hierarchical and progressive knowledge; In addition, in practical applications, knowledge graphs usually face the problem of data sparsity, that is, there is a lack of direct relationship connections between a large number of entities, which makes it difficult for traditional embedding methods relying solely on structural information to learn accurate entity representations.
[0004] For existing multi-modal knowledge graphs, the quality of entity image datasets generally has problems such as low correlation between images and entities and inaccurate extraction, which directly affects the learning effect of knowledge representation based on multi-modal information. Summary of the Invention
[0005] The purpose of this part is to outline some aspects of the embodiments of the present invention and briefly introduce some preferred embodiments. Simplifications or omissions may be made in this part, as well as in the abstract and title of the present application, to avoid obscuring the purpose of this part, the abstract, and the title. However, such simplifications or omissions shall not be used to limit the scope of the present invention.
[0006] In view of the above existing problems, the present invention is proposed.
[0007] Therefore, the technical problems solved by the present invention are: the problems of insufficient single-modal representation, limited expressive ability of shallow models, and poor quality of multi-modal data existing in the existing knowledge graph embedding methods.
[0008] To solve the above technical problems, the present invention provides the following technical solutions: Image preprocessing is performed on all entity image sets using an image filtering algorithm to obtain a high-quality image set corresponding to each entity;
[0009] Based on the high-quality image set, image feature extraction is performed using the deep residual network ResNet, and entity image feature representations are obtained through global average pooling;
[0010] The text descriptions of all entities are input into the BERT model to obtain text feature vectors, and entity text feature representations are obtained through a fully connected layer;
[0011] The entity structure representations, text feature representations, and image feature representations of all entities are input into a 2D modality neural network model. The 2D modality neural network model adopts the ConvE model structure for multimodal fusion and knowledge graph embedding to obtain entity / relationship representations that fuse text and images.
[0012] As a preferred solution of the text-image enhanced multimodal knowledge graph embedding method of the present invention, the image preprocessing includes:
[0013] Performing unification and color space conversion operations on all entity images;
[0014] Calculating the histogram data of the image and normalizing it;
[0015] Using the histogram correlation comparison method to calculate the similarity score between images;
[0016] Screening to obtain a high-quality image set related to the entity according to a set similarity threshold.
[0017] As a preferred solution of the text-image enhanced multimodal knowledge graph embedding method of the present invention, the calculation formula of the histogram correlation comparison method is:
[0018]
[0019] where N is the total number of bins in the histogram, is the histogram mean of the k-th image, H k (J) is the value of the J-th bin of the histogram of the k-th image, d(H1, H2) is the similarity score of the histograms of two images, and H1 and H2 are two histograms to be compared.
[0020] As a preferred solution of the text-image enhanced multimodal knowledge graph embedding method of the present invention, the structure of the deep residual network ResNet includes an input layer, an original surface layer, multiple residual block groups, a global average pooling layer, and a fully connected layer;
[0021] Among them, the multiple residual block groups achieve non - linear mapping of features through residual connections, and the global average pooling layer is used to compress the spatial dimension of the feature map.
[0022] As a preferred solution of the text - image enhanced multi - modal knowledge graph embedding method of the present invention, in the 2D morphological neural network model, the scoring function of the entity structure representation is:
[0023]
[0024] Among them, R r ∈R k, is a relation parameter dependent on the relation R between the head and tail entities, represents the two - dimensional structure of H s and R r Correspondingly, if H s , R r ∈R k , then At this time, k = k w k h *, where * represents the convolution operation, w is the filter of the 2D convolutional layer, W is the fully - connected projection matrix, H s represents the structure representation of the head entity of the triple in the knowledge graph, and T s represents the structure representation of the tail entity of the triple in the knowledge graph.
[0025] As a preferred solution of the text - image enhanced multi - modal knowledge graph embedding method of the present invention, in the 2D morphological neural network model, the scoring function of the entity text feature representation is:
[0026]
[0027] Among them, both the head entity and the tail entity in are entity text features, while the head entity in is an entity structure feature, while the tail entity is an entity text feature, the head entity in is an entity text feature, while the tail entity is an entity structure feature, H t represents the text representation of the head entity of the triple, and T t represents the text representation of the tail entity of the triple.
[0028] As a preferred solution of the text - image enhanced multi - modal knowledge graph embedding method of the present invention, in the 2D morphological neural network model, the scoring function of the entity image feature representation is defined as:
[0029]
[0030] Among them, both the head entity and the tail entity in the front part of are entity image features, while the head entity in i is an entity structure feature, and the tail entity is an entity image feature, i the head entity in
[0031] As a preferred solution of the text-image enhanced multi-modal knowledge graph embedding method described in the present invention, the 2D shape neural network model needs to be trained, including:
[0032] Combining a scoring function of entity structure representation, entity text feature representation, and entity image feature representation, and using the sigmoid function for scoring;
[0033] Constructing a binary cross-entropy loss function:
[0034]
[0035] where N is the number of samples, t i is the feature label, s i is the predicted score. For 1-1 scoring, t is a label vector of dimension R 1×1 and for 1-N scoring, the dimension is represented as R 1×N ;
[0036] Then, the Adam optimizer is used to optimize and train the loss function.
[0037] Advantages of the present invention:
[0038] 1. By using an image filtering algorithm, the quality of the entity image dataset is improved, images with low relevance to the entity are removed, and a higher-quality knowledge representation is obtained;
[0039] 2. By using a 2D convolutional neural network model to train the entity relationship representation, the feature interaction between entities and relationships is maximized;
[0040] 3. Using a 2D convolutional neural network model to fuse the features of each modality, the entity text features and entity image features are fused into the entity structure features with higher efficiency. Description of the Drawings
[0041] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. Among them:
[0042] Figure 1 It is a schematic flowchart of a text-image enhanced multi-modal knowledge graph embedding method shown in the present invention;
[0043] Figure 2 It is a schematic flowchart of obtaining an entity text representation shown in the present invention;
[0044] Figure 3 It is a schematic flowchart of extracting image features shown in the present invention;
[0045] Figure 4 It is a schematic diagram of the residual block group structure shown in the present invention;
[0046] Figure 5 It is a schematic flowchart of an entity image representation shown in the present invention;
[0047] Figure 6 It is a schematic diagram of the ConVE model structure shown in the present invention. Specific Embodiments
[0048] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will make a detailed description of the specific embodiments of the present invention in conjunction with the accompanying drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments.
[0049] Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0050] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0051] According to an embodiment of the present invention, in combination with Figure 1 the flowchart shown, a text-image enhanced multi-modal knowledge graph embedding method specifically includes:
[0052] S1. Use an image filtering algorithm to perform image preprocessing on all entity image sets to obtain a high-quality image set corresponding to each entity;
[0053] S2. Based on the high-quality image set, use the deep residual network ResNet to extract image features, and obtain the entity image feature representation through global average pooling;
[0054] S4. Input the text descriptions of all entities into the BERT model to obtain text feature vectors, and obtain the entity text feature representation through a fully connected layer;
[0055] S4. Input the entity structure representations, text feature representations, and image feature representations of all entities into the 2D modality neural network model. The 2D modality neural network model adopts the ConvE model structure for multimodal fusion and knowledge graph embedding to obtain the entity / relationship representations that fuse text and images.
[0056] The following combines Figures 2 to 6 the flowchart shown and some preferred or optional examples of the present invention to more specifically describe the implementation process and / or effects of certain examples of the present invention.
[0057]
Image Preprocessing
[0058] It should be noted that the image filtering algorithm adopted in this embodiment is used to filter out images irrelevant to the entity. Specifically, by using the image similarity algorithm (histogram comparison method), find the image set associated with the entity, that is, obtain the histogram data of the two pictures to be compared, then normalize the histogram data, and finally obtain a similarity index. By setting the boundary of the similarity index, it can be determined whether the pictures are of the same entity;
[0059] In an alternative embodiment, using the histogram comparison method to find the image set associated with the entity includes:
[0060] Preprocess all images, including resizing to a unified size, converting to the same color space (such as HSV), and denoising, to ensure the consistency of the comparison results;
[0061] Calculate the histogram of the image, select appropriate color channels (such as the H and S channels in HSV), divide the pixel values into a certain number of bins, count the frequency of pixels in each bin, and normalize the histogram, and normalize the result to the [0,1] interval to eliminate the influence of image size on the comparison;
[0062] Adopt the histogram comparison method to calculate the similarity between the reference image and the candidate image;
[0063] Correlation comparison process: For multiple input images, first obtain their histograms {H1, H2,..., H n}, then use the correlation comparison formula to pairwise compare the histograms of each image of the same entity with the histograms of all the other images to obtain the histogram similarity score d. If the histogram similarity of an image with the other images is less than the set similarity threshold threshold, it is considered that the correlation between the image and the entity is low, and the image is removed;
[0064] Exemplarily, define the histograms of two pictures to be compared as {H1, H2}, then the correlation comparison formula for the two histograms is:
[0065]
[0066] where N is the total number of bins in the histogram, is the mean value of the histogram of the k-th picture, and H k (J) is the value of the J-th bin of the histogram of the k-th picture, d(H1, H2) is the similarity score of the histograms of the two pictures, and H1 and H2 are the two histograms to be compared;
[0067] After obtaining the similarity scores of all the images, filter out the images that meet the conditions according to the set threshold, and form an image set related to the entity with these images;
[0068] Obtain the image set {img1, img2,..., img n} corresponding to the entity, where img i represents the image set composed of the images related to the i-th entity.
[0069]
Obtain Image Feature Vectors
[0070] After removing the images with low relevance, extract the image feature vectors. Refer to the Figure 3 shown flowchart (entity images 1-4 represent 4 different images of the same entity), and the implementation steps of the ResNet network specifically include:
[0071] Input layer: Input the images corresponding to each entity where, represents the m-th image of the i-th entity;
[0072] Normalize the input image, scale the pixel values to the [0, 1] interval, and adjust the input image to the corresponding size according to the input size of ResNet. The processed image has the shape (C, H, W), where C is the number of channels, H is the height of the image, and W is the width of the image;
[0073] Initial convolutional layer: Perform spatial downsampling and feature extraction on the input image (with the shape (C, H, W)). This layer consists of a kernel of size k h ×kw Composed of convolutional kernels, including a batch normalization layer and a ReLU activation function, to obtain an image with a shape of (C, H1, W1);
[0074] Residual block group: The structure of the basic residual block is as Figure 4 shown. The residual block group is composed of multiple residual blocks. Each residual block group is responsible for extracting more advanced features from the features extracted from the previous group. Through the residual connection, each residual block group can learn the complex non-linear mapping between the input and the output;
[0075] The operation of a residual block group is as follows:
[0076] Convolution: The first convolutional layer uses a convolutional kernel, and the output shape is C×H×W;
[0077] Batch normalization: Perform batch normalization on the output after convolution;
[0078] ReLU activation function: Apply ReLU activation to the output;
[0079] The second convolutional layer: The second convolutional layer, and the output shape is C×H×W;
[0080] Residual connection: Add the input to the output after convolution to form a residual connection. The output obtained through the residual connection still has a shape of C×H×W;
[0081] Average pooling layer: After feature extraction and non-linear mapping through a series of residual block groups, the ResNet network uses a global average pooling layer as the last convolutional layer of the network. In this layer, each feature map is reduced to a single value, thereby significantly reducing the model parameters and the amount of computation. The input comes from the output of the last residual block, with a shape of C f ×H f ×W f , apply global average pooling to the feature maps of each channel, and calculate the average value of each channel. Compress the spatial dimension (H f ×W f ) to 1, and the output shape is C f ×1;
[0082] Fully connected layer: Input the pooled feature vector into the fully connected layer to map it to the space of the number of categories. The output of the fully connected layer has a shape of output size ×1 (consistent with the size of the entity text feature vector);
[0083] Through the processing of the above layer structure, the feature vector of each image of the entity is obtained. Refer to Figure 5 , and then concatenate the vectors of multiple images of the entity into an m×output sizeThe feature matrix of the i-th entity's entity images (where m is the total number of entity images of the i-th entity), and then average pool the feature matrix (with a shape of m×output size two-dimensional vector) along the m dimension to obtain a one-dimensional vector with a dimension of 1×output size , and this feature vector fuses the features of all images of each entity.
[0084]
Obtain text feature vector
[0085] Furthermore, in this embodiment, the text feature vector is obtained by using the Bert model in pytorch. The text description of the entity is obtained from the annotation file of the WordNet corpus and input into the Bert model to obtain all word vectors {x1, x2, …, x n} corresponding to the entity text description. Then, the output of the Bert model is input into a fully connected layer to obtain a one-dimensional embedding vector representing the entity text feature;
[0086] Refer to Figure 2 , which specifically includes the following steps:
[0087] Extract the definition of the entity from WordNet: Use the wn.synsets method in the NLTK library to find the synsets of the entity and obtain the definition text from them. For example, for the entity "dog", find the synsets of the entity {dog.n.01, dog.n.02,...}, where n.01 and n.02 are two different synsets, which respectively correspond to the meanings of "dog" in different contexts. The text description of the synset is provided by synset.definition(), and the most appropriate text description is selected as the text description of the entity through professional knowledge (such as "A dog is a domesticated carnivorous mammal.");
[0088] Load the pre-trained BERT model and tokenizer: Through the Transformers library, load the BERT model and the corresponding BertTokenizer, input the obtained entity description text into the BERT model for tokenization processing, and obtain word vectors;
[0089] Pass the tokenized input to the BERT model: The BERT model returns the word vectors of each token (one-dimensional vectors of the same length);
[0090] Concatenate the n word vectors into an n×hidden size two-dimensional vector as the entity text description vector (assuming there are n words in the sentence, and the word vector dimension of each word is hidden size);
[0091] Then, the entity text description two-dimensional vector is input into the fully connected layer to obtain a one-dimensional entity text feature vector; the calculation formula of the fully connected layer is:
[0092] Y = XW T + b
[0093] where X is the input matrix, with a shape of (n, hidden size ), W T is the transpose of the weight matrix, with a shape of (hidden size , output size ), b is the bias matrix, with a shape of (1, output size ), and Y is the output of the fully connected layer, with a shape of (n, output size );
[0094] The output result of the fully connected layer (a two-dimensional vector with a shape of (n, output size )) is subjected to average pooling along the n dimension to obtain a one-dimensional vector with a dimension of 1×output size , and this one-dimensional vector is used as the entity text feature vector.
[0095] 【2D Convolution Knowledge Graph Embedding】
[0096] It should be noted that the structure-based representation of the knowledge graph refers to representing the entities and relationships in the knowledge graph by utilizing the structural features of the graph. The structure-based representation method focuses on the positions, connection methods, and mutual relationships of entities and relationships in the graph, rather than directly relying on external semantic information (such as text, images);
[0097] The 2D convolution knowledge graph embedding proposed in this embodiment makes up for the deficiencies of the knowledge graph representation based solely on structure, fuses the text representation and image representation of entities into the knowledge graph embedding for knowledge graph learning, and improves the expressiveness and generalization ability of the embedding model;
[0098] Among them, the triple representation of the entity structure feature in the knowledge graph is (H s , R, T s ), the triple representation of the entity text feature is (H t , R, T t ), the triple representation of the entity image feature is (H i , R, T i ), H s represents the structural representation of the head entity of the triple in the knowledge graph, T sRepresents the structural representation of the tail entity of a triple in the knowledge graph, H t Represents the text representation of the head entity of the triple obtained from the text feature vector, T t Represents the text representation of the tail entity of the triple obtained from the text feature vector, H i Represents the image representation of the head entity of the triple obtained from the image feature vector, T i Represents the image representation of the tail entity of the triple obtained from the image feature vector, R represents the relationship between the head and tail entities, and all different entity representations share the corresponding relationship vector;
[0099] Preferably, the 2D convolutional model adopted in this embodiment is the ConvE model, and the structure of the ConvE model is as Figure 6 shown. According to the architecture diagram of the model, the vectors of the head entity and the relationship are stacked and then reshaped into a two-dimensional tensor (similar to the representation of an image). After convolution, a feature map is obtained, and then it is projected into a k-dimensional space through a fully connected layer and matched with the embedding of the candidate target in the inner product layer (implemented using a scoring function), and finally, the entity / relationship representation integrating text and image is obtained.
[0100] As an example, in the 2D morphological neural network model, the scoring function of the entity structure representation is:
[0101]
[0102] where, R r ∈R k, is a relationship parameter depending on the relationship R between the head and tail entities, represents H s and R r 's two-dimensional structure. Correspondingly, if H s , R r ∈R k , then At this time, k = k w k h , * represents the convolution operation, w is the filter of the 2D convolutional layer, W is the fully connected projection matrix, H s represents the structural representation of the head entity of the triple in the knowledge graph, T s represents the structural representation of the tail entity of the triple in the knowledge graph;
[0103] To fuse the entity text features, in the 2D morphological neural network model, the scoring function of the entity text feature representation is:
[0104]
[0105] where, both the head entity and the tail entity in are entity text features, while The head entity in it is the entity structure feature, while the tail entity is the entity text feature, The head entity in it is the entity text feature, while the tail entity is the entity structure feature, H t Represents the text representation of the head entity of the triple, T t Represents the text representation of the tail entity of the triple.
[0106] As an example, in order to fuse the entity image features, in the 2D morphological neural network model, the scoring function of the entity image feature representation is defined as:
[0107]
[0108] Among them, Both the head entity and the tail entity in the front part of it are entity image features, while The head entity in it is the entity structure feature, while the tail entity is the entity image feature, The head entity in it is the entity image feature, while the tail entity is the entity structure feature, H i Represents the image representation of the head entity of the triple, T i Represents the image representation of the tail entity of the triple.
[0109] In an alternative embodiment, the 2D shape neural network model needs to be trained, including:
[0110] Combining the scoring functions of the entity structure representation, entity text feature representation, and entity image feature representation, and using the sigmoid function to score;
[0111] Construct a binary cross-entropy loss function:
[0112]
[0113] Among them, N is the number of samples, t i Is the feature label, s i Is the predicted score. For 1-1 scoring, t is a label vector of dimension R 1×1 And for 1-N scoring, the dimension is represented as R 1×N ;
[0114] Then use the Adam optimizer to optimize and train the loss function.
[0115] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A text-image enhanced multimodal knowledge graph embedding method, characterized in that: include: Use image filtering algorithms to preprocess all entity image sets to obtain high-quality image sets corresponding to each entity; Based on the high-quality image set, a deep residual network ResNet is used to extract image features, and a feature representation of the entity image is obtained by global average pooling; Input the text descriptions of all entities into the BERT model to obtain the text feature vector, and obtain the entity text feature representation through the fully connected layer; The entity structure representation, text feature representation and image feature representation of all entities are input into the 2D modal neural network model. The 2D modal neural network model adopts the ConvE model structure for multimodal fusion and knowledge graph embedding to obtain the entity / relationship representation of the fused text and image.
2. The text-image enhanced multimodal knowledge graph embedding method according to claim 1, characterized in that: The image preprocessing comprises: Perform unification and color space conversion operations on all entity images; Calculate the histogram data of the image and normalize it; The histogram correlation comparison method is used to calculate the similarity score between images; A set of high-quality images related to the entity is obtained by screening according to the set similarity threshold.
3. The text-image enhanced multimodal knowledge graph embedding method according to claim 2, characterized in that: The calculation formula of the histogram correlation comparison method is: Where N is the total number of bins in the histogram, is the histogram mean of the kth image, H k (J) is the value of the Jth bin of the histogram of the kth image, d(H1,H2) is the similarity score of the histograms of the two images, and H1 and H2 are the two histograms to be compared.
4. The text-image enhanced multimodal knowledge graph embedding method according to claim 1, characterized in that: The structure of the deep residual network ResNet includes an input layer, an original surface layer, multiple residual block groups, a global average pooling layer and a fully connected layer; The plurality of residual block groups realize nonlinear mapping of features through residual connections, and the global average pooling layer is used to compress the spatial dimension of the feature map.
5. The text-image enhanced multimodal knowledge graph embedding method according to claim 1, characterized in that: In the 2D morphological neural network model, the scoring function of the entity structure representation is: Among them, R r ∈R k, is a relation parameter that depends on the relation R between the head and tail entities. Indicates H s and R r The two-dimensional structure, accordingly, if H s , R r ∈R k ,So At this time k = k w k h , * represents the convolution operation, w is the filter of the 2D convolution layer, W is the fully connected projection matrix, H s Represents the structural representation of the head entity of the triple in the knowledge graph, T s Structural representation of the tail entity of a triple in a knowledge graph.
6. The text-image enhanced multimodal knowledge graph embedding method according to claim 1, characterized in that: In the 2D morphological neural network model, the scoring function of the entity text feature representation is: in, The head entity and tail entity in are both entity text features, and The head entity in is the entity structure feature, while the tail entity is the entity text feature. The head entity in H is the entity text feature, while the tail entity is the entity structure feature. t The text representation of the head entity of the triple, T t The textual representation of the tail entity of a triple.
7. The text-image enhanced multimodal knowledge graph embedding method according to claim 1, characterized in that: In the 2D morphological neural network model, the scoring function for entity image feature representation is defined as: in, The head entity and the tail entity in the front part of are both entity image features, while The head entity in is the entity structure feature, and the tail entity is the entity image feature. The head entity in H is the entity image feature, while the tail entity is the entity structure feature. i The image representation of the head entity of the triple, T i Image representation of the tail entity of a triple.
8. The text-image enhanced multimodal knowledge graph embedding method according to any one of claims 5 to 7, characterized in that: The 2D shape neural network model needs to be trained, including: The scoring function combines the entity structure representation, entity text feature representation and entity image feature representation, and uses the sigmoid function for scoring; Construct a binary cross entropy loss function: Where N is the number of samples, t i is the feature label, s i is the predicted score. For a 1-1 score, t is the dimension R 1×1 The label vector of , and for 1-N scoring, the dimension is expressed as R 1×N ; The Adam optimizer is then used to optimize the loss function.
Citation Information
Patent Citations
Element reverse detection method and system
CN106204602A
Aluminum electrolytic capacitor defect detection method based on image processing
CN113658092A
Text-image enhanced multi-modal knowledge graph embedding method
CN115099409A
Multi-modal knowledge graph establishment method and application
CN117131933A
Multi-modal knowledge graph privacy protection embedding method for diagnosis and treatment data
CN119475429A