A multimodal knowledge graph completion method based on dynamic prompt learning and multi-granularity aggregation

By introducing dynamic prompt learning templates and multi-grained cross-modal aggregation into the multimodal knowledge graph completion method, the problem of modal contradictions and insufficient feature extraction in the existing methods is solved, and the generalization ability and reasoning accuracy of the model are improved.

CN119783799BActive Publication Date: 2025-05-13NANJING UNIV OF INFORMATION SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510296663.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-05-13
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

The existing multimodal knowledge graph completion method ignores the visual noise independent of the entity when fusing visual information, resulting in modal contradictions and affects the accuracy of entity representation; at the same time, the prompt learning of fixed templates and coarse-grained feature extraction limits the generalization ability and reasoning accuracy of the model.

Method used

Dynamic prompt learning templates and multi-grained cross-modal aggregation methods are adopted to dynamically adjust the template structure through an adaptive guidance mechanism, and multi-grained cross-modal aggregation is introduced in the Transformer model to fuse visual features and text features of coarse and fine-grained size.

Benefits of technology

It improves the generalization ability and inference accuracy of the model, enhances the ability to capture local information of the image, reduces the sensitivity to irrelevant visual noise, and improves the robustness and adaptability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119783799B_ABST
    Figure CN119783799B_ABST
Patent Text Reader

Abstract

The present invention provides a multimodal knowledge graph completion method based on dynamic prompt learning and multi-granularity aggregation, comprising the following steps: Step 1, generating a dynamic prompt template: selecting a suitable template structure according to task requirements and dynamically adjusting the template structure using an adaptive guidance mechanism; Step 2, establishing a multimodal feature encoder to convert text and image data into feature vectors for training this model; Step 3, multi-granularity cross-modal aggregation MCA: fusing features of different modalities and different granularities; Step 4, designing a multi-task joint loss function, optimizing the performance of the model in more than two multimodal tasks and training. The method of the present invention can more comprehensively understand the image content, thereby improving the accuracy of reasoning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a multimodal knowledge graph completion method based on dynamic prompt learning and multi-granularity aggregation. Background Art

[0002] With the popularity of multimedia data on social media platforms, multimodal knowledge graphs (MKGs) have attracted more and more attention. MKGs contain not only traditional text information but also visual information, which greatly enriches the expressive power of knowledge graphs and enables them to more comprehensively capture and express information of the complex real world. Multimodal knowledge graph completion (MKGC) technology aims to fill in missing facts in knowledge graphs by analyzing visual and textual data, thereby improving the completeness and reasoning accuracy of knowledge graphs.

[0003] Existing MKGC methods mainly face the following challenges:

[0004] Modal inconsistency: When fusing visual information, existing methods often ignore visual noise that is irrelevant to the entity, resulting in modal inconsistency and affecting the accuracy of entity representation.

[0005] Universality of model architecture: Different MKGC tasks and modality representations require changes to the model architecture, increasing the complexity of model development and maintenance.

[0006] Limitations of cue learning: Existing cue learning methods are mainly based on fixed templates, which lack flexibility and are difficult to adapt to the needs of different tasks.

[0007] To address these challenges, some MKGC methods have been proposed:

[0008] Methods based on early fusion: Aggregation methods learn the features of different modalities separately and then aggregate them directly, lacking the interaction within the modalities. Alignment methods: Align the distributions of different modalities through normalization, but ignore the intrinsic characteristics of the modalities.

[0009] Late fusion based methods: DRAKE: uses visual information to supplement the missing visual context and represents textual information as object-aware knowledge representation. DSDCAN: reconstructs data of a unified modality by reconstructing the extracted modality style features and the content features of another modality.

[0010] Disadvantages of existing technology:

[0011] Lack of dynamic prompt learning templates: Most existing multimodal knowledge graph completion (MKGC) methods use fixed templates to fuse visual and textual information. Since different tasks have different information requirements, fixed templates are difficult to adapt to all tasks, resulting in limited model performance.

[0012] Lack of multi-granularity feature extraction: Existing MKGC methods mainly focus on coarse-grained image features and ignore fine-grained image features, which results in the model being unable to effectively capture local information in the image and affects the reasoning accuracy.

[0013] Modal contradiction: Existing MKGC methods often ignore visual noise irrelevant to entities when fusing visual information, resulting in modal contradiction and affecting the accuracy of entity representation.

[0014] Problems caused by shortcomings of existing technologies:

[0015] Due to the above shortcomings, the existing MKGC method has deficiencies in the following aspects:

[0016] Poor generalization ability: The model is difficult to adapt to the needs of different tasks and its generalization ability is limited.

[0017] Low inference accuracy: The model cannot effectively capture local information in the image, resulting in low inference accuracy.

[0018] Poor robustness: The model is easily affected by irrelevant visual noise and has poor robustness. Summary of the invention

[0019] Purpose of the invention: The technical problem to be solved by the present invention is to provide a multimodal knowledge graph completion method based on dynamic prompt learning and multi-granularity aggregation in view of the deficiencies of the prior art. The method optimizes the Transformer architecture, adds dynamic prompt templates and multi-granularity cross-modal aggregation, and obtains an improved Transformer model for multimodal knowledge graph completion, which specifically includes the following steps:

[0020] Step 1, generate dynamic prompt template: select the appropriate template structure according to the task requirements and use the adaptive guidance mechanism to dynamically adjust the template structure;

[0021] Step 2: Establish a multimodal feature encoder to convert text and image data into feature vectors for training this model;

[0022] Step 3: Multi-granularity Cross-modal Aggregation (MCA): Fuse features of different modalities and granularities to capture multimodal information more comprehensively.

[0023] Step 4: Design a multi-task joint loss function to optimize the performance of the model in more than two multimodal tasks and perform training.

[0024] Step 1 includes: given an entity description T = {V1, V2, ..., Vn} and its associated image I = {I1, I2, ..., In}, design the following dynamic prompt learning template ti :

[0025] t i =[CLS][Entiry1][SEP][Relation][SEP][Entity2][SEP]

[0026] [V1][V2]...[Vn][SEP][MASK][SEP],

[0027] Among them, [CLS] is usually a special token in the field of natural language processing (NLP) to indicate the classification information of the entire input sequence. It is usually added to the beginning of the input sequence to capture the semantic features of the entire text; [Entity1] represents the first entity, which is suitable for tasks such as named entity recognition (NER) for entity recognition, relation extraction (RE) for head entities, and link prediction for head entities; [SEP] is used to separate different parts in a sentence. [Relation] represents the relationship between entities, which is used for relation recognition in RE tasks and for relation recognition in link prediction; [Entity2] represents the second entity as the tail entity in RE tasks and as the tail entity in link prediction; Vn represents the nth element in the entity description, and In represents the nth element in the associated image;

[0028] The token [MASK] is used to represent the information that the improved Transformer model needs to predict, such as completing the tail entity in the link prediction task. Through this template, the improved Transformer model can dynamically adjust the prompt according to the specific task requirements, making the integration of different modes in tasks such as entity alignment, relationship reasoning, and link prediction more flexible and effective, and improving the accuracy and adaptability of the task.

[0029] At the same time, an adaptive task guidance mechanism is used to dynamically adjust the prompt design of the improved Transformer model. This mechanism modifies the structure of the prompt template according to specific task requirements (such as named entity recognition (NER), relation extraction (RE), or link prediction). In the NER task, the structure of the [Entity1] part is reconstructed according to the context details, and the entity attributes are reordered or extended. This may involve reordering or extending the description of the entity attributes, allowing the model to more accurately locate and identify the target entity. This dynamic adjustment ensures that the prompt design is tailored to the complexity of the task, thereby enhancing the accuracy of entity recognition. Similarly, in the relation extraction task, the [relation] part is reorganized by adjusting the representation and positioning to ensure a closer alignment between the relation type and context features.

[0030] In step 1, the dynamic prompt template includes a multimodal link prediction MLP template, a multimodal named entity recognition MNER template and a multimodal relationship extraction MRE template:

[0031] The multimodal link prediction MLP template is expressed as:

[0032] [CLS][Entity1][SEP][Re1][SEP]

[0033] [MASK][SEP][V1][V2]...[Vn][SEP],

[0034] where the token [MASK] is used to predict the missing entity associated with the first entity [Entity1] through the relation [Rel];

[0035] The multimodal named entity recognition MNER template is expressed as:

[0036] [CLS][MASK][SEP][Re1][SEP]

[0037] [MASK][SEP][V1][V2]...[Vn][SEP],

[0038] For the named entity recognition task, the token [MASK] is used to identify and classify entities in textual and visual contexts;

[0039] The multimodal relation extraction MRE template is expressed as:

[0040] [CLS][Entity1][SEP][Re1][SEP]

[0041] [V1][V2]...[Vn][SEP].

[0042] After the model automatically recognizes the current task type, it will select the corresponding template structure.

[0043] In step 2, the multimodal feature encoder includes a visual encoder and a text encoder; the visual encoder is responsible for encoding the synchronized granular visual features, and the text encoder is responsible for encoding the text features;

[0044] The text encoder is used to extract text features Q, including:

[0045] Given a dynamic prompt learning template t i , using the transformer-based bidirectional encoder BERT (BidirectionalEncoder Representations from Transformers) to transform the dynamic prompt learning template t iEncoded into a high-dimensional representation E(t i ), the formula is:

[0046]

[0047] Where L is the length of the input sequence and d is the embedding dimension;

[0048] Then the coding sequence E(t i ) Apply the multi-head attention mechanism and query the matrix Q (l) , key matrix K (l) Sum value matrix V (l) The calculation formula is:

[0049]

[0050] in are the weight matrices of query, key and value respectively. d h is the dimension of each attention head, each attention head in the multi-head attention mechanism l It is expressed as:

[0051] head l =Attention(Q (l) ,K (l) ,V (l) ),

[0052] in Attention is usually implemented as a scaled dot-product attention, calculated as:

[0053]

[0054] Among them, QK T is the dot product of the query and the key, d k is the dimension of the key, the softmax function is used to convert the result of the dot product into a probability distribution, and finally multiply it by the value matrix V to get the weighted output;

[0055] The outputs of all attention heads are concatenated and linearly transformed to obtain:

[0056] T0=E(t i )+pos,

[0057]

[0058] Among them, pos is positional encoding, which is used to provide the model with the position information of each element in the sequence, because Transformer itself does not have the ability to process sequence order; MHA is multi-head attention, LN is layer normalization, FFN is feedforward neural network, and T0 is the initial hidden state of the input sequence; is the intermediate hidden state of the lth layer, which is calculated by multi-head attention MHA and layer normalization LN, and then combined with the hidden state T of the previous layer l-1 Addition; is the intermediate hidden state; T l is the final hidden state of the lth layer, which is calculated by the feed-forward neural network FFN and layer normalization LN, and then combined with the intermediate hidden state Addition;

[0059] In the visual encoder, the residual network encoder ResNet (Residual Network) and the visual transformer encoder ViT (Vision Transformer Encoder) are used to extract coarse-grained features K and fine-grained features V from the image respectively:

[0060] The residual network encoder ResNet is responsible for extracting global coarse-grained features F from the input image I of size C×H×W coarse , expressed as:

[0061] F coaRse =ResNet(I),

[0062] in C′ is the number of channels of the final convolutional layer, H′ and W′ represent the height and width of the obtained feature map respectively; C, H, W represent the channel input, height and width of the input image respectively;

[0063] In order to convert the feature map into a set of feature vectors, F coarse Reshape into a two-dimensional matrix Z of size (H′×W′)×C′ coarse :

[0064] Z coarse =Reshape(F coarse ),

[0065] Reshape means reshaping a tensor or matrix, that is, transforming the original data structure into another shape or dimension, but the total number of elements remains unchanged;

[0066] r = H′×W′; The Reshape operation here changes the shape of F_coarse from (C, H, W) to (H×W, C), which means that each spatial position (height and width) of the feature map is flattened into a vector, and all these vectors are stacked together according to the number of channels C′; r represents the number of feature vectors after reshaping, that is, the total number of spatial positions of the feature map.

[0067] The visual transformer encoder ViT is responsible for extracting local coarse-grained features from an input image I of size C×H×W: Given an input image I, the visual transformer encoder ViT first divides the image into non-overlapping patches of size P×P, where P represents the height and width of the patches into which the input image is divided. The total number of patches v is given by:

[0068]

[0069] Then flatten each patch into a length of P 2 ×C vector and projected to a fixed embedding dimension d using a linear layer v , the obtained patch embedding representation for:

[0070]

[0071] in represents the embedding of the vth small block, represents the position encoding corresponding to the vth small block; this formula represents the final embedded representation of each image block after linear layer projection and position encoding. This embedded representation can be further processed by the Transformer model;

[0072] In the multi-head attention mechanism, each attention head hl processes the input data independently and generates its own output. Subsequently, these outputs from different attention heads are concatenated and integrated through a linear transformation layer. This integrated result is then fed into a feed-forward neural network (FFN) to further process the patch embedding. The whole process follows the standard Transformer architecture for optimization;

[0073] In the lth layer, the calculation formula for each small block intermediate output is:

[0074]

[0075] in is the output of the multi-head attention mechanism of the vth small block in the lth layer, l = 1,...,L v ; L v Indicates the total number of layers; is the output of the l-1th layer, that is, the processing result of the l-1th layer on the vth patch; Express Apply layer normalization to stabilize and speed up the training process;

[0076] is the output of the multi-head self-attention mechanism, which considers the interrelationships between different patches to capture global dependencies; is the original input;

[0077] is the output after multi-head self-attention processing, which is obtained by adding the original input (residual connection) to preserve the original information, which helps alleviate the gradient vanishing problem in deep networks;

[0078] Subsequently, a feed-forward neural network FFN is used to refine the patch embedding as:

[0079]

[0080] in, is the output of the feedforward neural network, which further processes and refines the feature representation. FFN usually consists of two linear transformations, possibly with an activation function in between;

[0081] is the final output, which is obtained by adding the output of MHA (residual connection) to retain information, which also helps alleviate the gradient vanishing problem;

[0082] After processing all layers, the final representations of all small patches are concatenated to obtain a comprehensive fine-grained feature representation:

[0083]

[0084] Where Z fine represents the final fine-grained feature representation, Concat is a concatenation operation that concatenates the final representations of all small blocks together; represents the final representation of the i-th small block after being processed by the Lv layer Transformer; n is the total number of small blocks.

[0085] In step 3, the multi-granularity cross-modal aggregation MCA includes:

[0086] Two linear layers are applied to map the coarse-grained and fine-grained visual features and text features into a unified dimensional space. For the original coarse-grained visual feature Z coarse , fine-grained visual features Zfine and text features T l , apply the following formula to convert:

[0087] z′ coarse =f1(z coarse ),

[0088] z′ fine =f2(z fine ),

[0089] x′ T =f3(T l ),

[0090] Where f1 is the first linear transformation function, usually a linear layer (e.g., a fully connected layer), which maps zcoarse to the target dimensional space; z′ coarse is the transformed coarse-grained visual feature, now in a unified dimensional space; f2 is the second linear transformation function, usually a linear layer (e.g., a fully connected layer), which maps zcoarse to the target dimensional space; z′ fine is the transformed fine-grained visual feature, now in a unified dimensional space; f3 is the third linear transformation function, usually a linear layer (e.g., a fully connected layer), which maps zcoarse to the target dimensional space; x′ T are the transformed text features, which can now be compared and combined with visual features in the same space;

[0091] Then, through the global pooling operation, from the text feature x′ T , coarse-grained feature z′ coarse and fine-grained features z′ fine Extract the global context GT from:

[0092] GT = GlobalPool(x′ T ,z′ coaRSe ,z′ fine ),

[0093] where x′ T ∈R m×d , z′ coarse ∈R n×d , z′ fine ∈R keyN×d They are query vector, key vector and value vector, m is the number of queries, keyN is the number of keys, d is the dimension of the feature; GlobalPool represents the global pooling operation;

[0094] Add the global context to the query vector x′ T and the key vector z′ coarse In the feature, the attention weight A is obtained:

[0095]

[0096] Where T represents transpose;

[0097] Calculate the value vector z′ fine The weighted sum Z of:

[0098] Z=A·z′ fine ,

[0099] where Z is a new representation that integrates local feature interactions and global context;

[0100] A learnable parameter α is used to adjust the influence of the global context and obtain the final attention weight A f :

[0101]

[0102] Among them, A f Represents attention weights, which are used to measure the importance of different parts of the input sequence to the current task (such as prediction, classification, etc.).

[0103] In step 4, the loss function is described as follows:

[0104] For a given triple (h, r, t), the improved Transformer model generates embeddings of the head entity h, relation r, and tail entity t based on multimodal information, and then calculates the prediction score. The loss function ζ link is defined as:

[0105] ζ link =-∑ (h,r,t)∈PS logσ(f(h,r,t))-∑ (h′,r′,t′)∈PS logσ(-f(h′,r′,t′)),

[0106] Where PS represents the positive sample set in the training data, (h′, r′, t′) represents the negative sample set, σ is the sigmoid function, and f() is the scoring function based on triples.

[0107] For the multimodal relation extraction task, the cross entropy loss function is used to optimize the ability of the improved Transformer model to predict relation categories. Given an input entity pair (e1, e2) and the corresponding relation label r, the cross entropy loss ζ relation Defined as:

[0108]

[0109] Where RN represents the number of relation categories, y i is the one-hot encoded label of the target relation, p i is the probability distribution of the predicted relation category;

[0110] In the named entity recognition task, cross entropy loss is used to optimize entity category prediction. Given an input sequence X = {X1, X2, ..., X IE} and the corresponding entity tag Y = {Y1, Y2, ..., Y IE}, where each label corresponds to a word or phrase in the sequence, then the loss function ζ NER Defined as:

[0111]

[0112] Where IE represents the length of the input sequence, EN is the number of entity categories, and y Li,Tj is the target label of category Tj at the Li-th position, p Li,Tj is the predicted probability of category Tj at the Lith position; X IE is the IEth element in the input sequence, Y IE It is the IEth element in the entity tag;

[0113] The final loss function ζ total Defined as:

[0114] ζ total =λ link ζ link +λ relation ζ relation +λ NER ζ NER ,

[0115] where λ link , relation and λ NER is a hyperparameter that controls the loss weight.

[0116] The present invention also provides an electronic device, comprising a processor and a memory, wherein the memory stores program code, and when the program code is executed by the processor, the processor executes the steps of the described method.

[0117] The present invention also provides a storage medium storing a computer program or instruction. When the computer program or instruction is run on a computer, the steps of the method described are executed.

[0118] The present invention has the following beneficial effects: 1. The present invention introduces a dynamic prompt learning template:

[0119] Results: The model can dynamically adjust the prompt structure according to the needs of different tasks;

[0120] Advantages: Improved generalization ability, the model can adapt to more tasks, such as MLP, MRE and MNER, thereby improving the practicality and application scope of the model;

[0121] Improve flexibility: The model can flexibly adjust the information fusion method according to task requirements, thereby improving the adaptability and robustness of the model;

[0122] 2. The present invention introduces a multi-granularity cross-modal aggregation method:

[0123] Results: The model can effectively capture local information in the image;

[0124] Advantages: Improved reasoning accuracy. The model can understand the image content more comprehensively, thereby improving reasoning accuracy, such as predicting missing entities in MLP tasks, extracting entity relations in MRE tasks, and identifying and classifying entities in MNER tasks.

[0125] Improved robustness: The model can effectively filter out irrelevant visual noise, thereby improving robustness and enabling it to better handle real-world data;

[0126] 3. The present invention introduces an adaptive task guidance mechanism:

[0127] Results: The model can dynamically adjust the prompt template structure according to task requirements;

[0128] Advantages: Improved task adaptability. The model can be optimized according to the characteristics of different tasks, thereby improving the performance of the model on specific tasks.

[0129] Improve efficiency: The model can avoid using template structures that are not suitable for the current task, thereby improving the training and reasoning efficiency of the model.

[0130] The advantages of DM-MKGC lie in its dynamic prompt learning template, multi-granularity cross-modal aggregation method and adaptive task guidance mechanism, which effectively improve the generalization ability, reasoning accuracy and robustness of the model, giving it a significant advantage in the field of MKGC. BRIEF DESCRIPTION OF THE DRAWINGS

[0131] Figure 1 It is a flow chart of the method of the present invention. DETAILED DESCRIPTION

[0132] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments, and the above and / or other advantages of the present invention will become more clear.

[0133] like Figure 1As shown, an embodiment of the present invention provides a multimodal knowledge graph completion method based on dynamic prompt learning and multi-granularity aggregation, comprising the following steps:

[0134] Step 1: Define the problem;

[0135] Given a multimodal knowledge graph G = {(E, R, T)}, where E represents the set of entities (nodes) and R represents the set of relations (edges). A triple in the knowledge graph is represented as (h, r, t), where h, t∈E are the head and tail entities and r∈R is the relationship between them. Each entity ei∈E is also associated with a set of images, denoted as I(ei). The multimodal knowledge graph completion (MKGC) method aims to learn the embeddings of entities and relations by jointly encoding textual and visual information. Let h, r, t∈Rd denote the low-dimensional vector embeddings of the head entity, relation, and tail entity, respectively, where d is the embedding dimension. The goal of MKGC is to predict the missing entity in the triple, either the tail entity t given (h, r, ?), or the head entity h given the triple (?, r, t). "?" represents the missing entity to be predicted.

[0136] Definition 1 (Multimodal Link Prediction): In the Multimodal Knowledge Graph MKG = (E, R, I), find the triple query (e q ,r q ,? ) or (?,r q ,e q ) in the missing entity.

[0137] Definition 2 (Multimodal Relation Extraction): In MKGG = (E, R, I), the relationship R between two entities E1 and E2 in a given context is extracted by simultaneously utilizing textual and visual information to improve accuracy and completeness.

[0138] Definition 3 (Multimodal Named Entity Recognition): In MKGG = (E, R, I), given a text sequence X = {X1, X2, ..., X n} and the corresponding image I = {I1, I2, ..., I m}, predict entity label Y = {Y1, Y2, ..., Y n By utilizing text features f T (X) and visual features f V (I) to improve recognition accuracy and resolve ambiguity issues.

[0139] In the Multimodal Knowledge Graph (MKG), the notation MKG=(E,R,I) is used to make it clear that the graph consists of three main parts:

[0140] E (Entity): represents a collection of entities (nodes). These entities can be people, places, objects, concepts, etc. They are the basic elements in the knowledge graph.

[0141] R (Relation): represents a set of relations (edges). These relations define the connection between entities, such as "is", "belongs to", "is located in", etc.

[0142] I (Image): represents the set of images associated with each entity. In a multimodal knowledge graph, in addition to text information, it also contains image information related to the entity. These images can provide additional contextual information to help better understand and represent the entity.

[0143] exist Figure 1 In the , ResNet encoder is the abbreviation of Residual Network encoder.

[0144] ViT encoder is the abbreviation of Vision Transformer Encoder.

[0145] BERT encoder is the abbreviation of Bidirectional Encoder Representations from Transformers.

[0146] h [CLS] : In the DM-MKGC model, h [CLS] It is also used to capture the overall semantic features of the image.

[0147] h r : Represents a relationship description, which is used for relationship identification in relationship extraction tasks and for relationship identification in link prediction.

[0148] h [SEP] : represents a separator, which is used to separate different text information or image information.

[0149] h vn : Represents the feature vector of the nth patch in the image. The value range of n depends on the number of patches into which the image is divided.

[0150] h mask Represents the feature vector after the mask operation in the image.

[0151] When defining multimodal link prediction, MKG = (E, R, I) is used to emphasize that when making link predictions, not only entities and relations are considered, but also image information related to the entities. The purpose of this is to utilize multimodal information (text and vision) to improve the accuracy and robustness of predictions.

[0152] In the Multimodal Knowledge Graph Completion (MKGC) method, embeddings of entities and relations are learned by jointly encoding textual and visual information to predict missing entities in triples. This approach can more comprehensively utilize the information in the knowledge graph and improve the performance of link prediction.

[0153] Step 2: Establish an improved Transformer model based on dynamic prompt learning template and multi-granularity cross-modal aggregation. The model mainly consists of three parts:

[0154] Dynamic Hint Learning Template: The Dynamic Hint Learning Template is designed to generate dynamic hints for multimodal knowledge graph completion tasks and their subtasks, effectively guiding the model to perform knowledge reasoning and enhance its performance in tasks such as entity alignment, relation reasoning, and link prediction.

[0155] Multimodal Feature Encoder: The multimodal feature encoder consists of two components: a visual encoder and a text encoder. The visual encoder is responsible for encoding synchronized granular visual features, while the text encoder is responsible for encoding text features.

[0156] Multi-granularity cross-modal aggregation (MCA): MCA aims to promote the interaction and fusion of global (coarse-grained) and local (fine-grained) visual features with text features, and enhance the synergy between multimodal information. This method can not only extract richer global and local information from images, but also effectively alleviate the differences between different modal features, further improving the overall effectiveness of feature fusion.

[0157] Dynamic Prompt Learning Template: In previous work, the fusion or concatenation of visual features and text features of entities is often done in a direct way, which fails to fully utilize the deeper capabilities of multimodal modeling, especially in effectively guiding the model to complete complex tasks. Inspired by Large Language Model (LLM), especially Masked Language Model (MLM) in Transformer model, this paper proposes a new dynamic prompt learning template, which is designed to be suitable for multimodal knowledge graph completion.

[0158] Various subtasks in Multimodal Knowledge Graph Completion (MKGC).

[0159] Given an entity description T = {V1, V2, ..., Vn} and its associated image I = {I1, I2, ..., In}, the present invention designs a dynamic prompt learning template t i :

[0160] t i =[CLS][Entiry1][SEP][Relation][SEP][Entity2][SEP]

[0161] [V1][V2]...V[n][SEP][MASK][SEP]

[0162] Among them, [Entity1] represents the first entity, which is suitable for tasks such as named entity recognition (NER) of entity recognition, relationship extraction (RE) of head entities, and link prediction of head entities. [Relation] represents the relationship between entities, which is used for relationship recognition in RE tasks and for relationship recognition in link prediction. [Entity2] represents the second entity, which serves as the tail entity in RE tasks and as the tail entity in link prediction. [V1] decomposes text information into finer-grained word-level details, and [MASK] is used to represent the information that the model needs to predict, such as completing the tail entity in the link prediction task. Through this template, the model of the present invention can dynamically adjust prompts according to specific task requirements, making the integration of different modes in tasks such as entity alignment, relationship reasoning, and link prediction more flexible and effective, thereby improving the accuracy and adaptability of the task.

[0163] The present invention adopts an adaptive task guidance mechanism to dynamically adjust the prompt design of the model. The mechanism modifies the structure of the prompt template according to specific task requirements, such as named entity recognition (NER), relation extraction (RE), or link prediction. In the NER task, the structure of the [Entity1] part is reconstructed according to the context details, which may involve reordering or extending the description of the entity attributes, allowing the model to more accurately locate and identify the target entity. This dynamic adjustment ensures that the prompt design is tailored to the complexity of the task, thereby enhancing the accuracy of entity recognition. Similarly, in the relation extraction task, the [relation] part is reorganized by adjusting its representation and positioning, ensuring a closer alignment between the relationship type and context features.

[0164] After introducing the adaptive task guidance mechanism, the present invention designs specific templates tailored for different tasks in Multimodal Knowledge Graph Completion (MKGC). These templates are dynamically generated by adjusting their structures according to task requirements.

[0165] Multimodal Link Prediction (MLP) template:

[0166] [CLS][Entity1][SEP][Re1][SEP]

[0167] [MASK][SEP][V1][V2]...[Vn][SEP]

[0168] The [MASK] token is used to predict the missing entity associated with [Entity1] through the relation [Rel]. The inclusion of visual and textual features (V1…Vn) allows the model to exploit multimodal information, thereby facilitating more accurate predictions of missing links.

[0169] Multimodal Named Entity Recognition (MNER) template:

[0170] [CLS][MASK][SEP][Re1][SEP]

[0171] [MASK][SEP][V1][V2]...[Vn][SEP]

[0172] For named entity recognition tasks, the [MASK] token is used to identify and classify entities in both textual and visual contexts. This template allows the model to efficiently process multimodal information to accurately recognize entities.

[0173] Multimodal Relation Extraction (MRE) template:

[0174] [CLS][Entity1][SEP][Re1][SEP]

[0175] [V1][V2]...[Vn][SEP]

[0176] The templates are specifically designed for relation extraction tasks, focusing on extracting the relationship between two known entities [Entity1] and [Entity2]. These templates are generated through an adaptive task-guided mechanism, dynamically adjusting their structure according to task requirements, while maintaining flexibility and framework consistency. The integration of multimodal data within these templates significantly enhances the model's ability to handle complex tasks such as link prediction, named entity recognition, and relation extraction, ultimately improving performance and robustness.

[0177] Multimodal feature encoder: The multimodal feature encoder mainly consists of two parts: (1) The text encoder is implemented with the BERT encoder, which is responsible for capturing basic syntactic features and vocabulary information to ensure that the model can understand and process complex language structures; (2) The visual encoder is implemented with the residual network encoder ResNet and the visual transformer encoder ViT, which aims to capture image features of different granularities from global to local and provide a comprehensive representation of visual information. In addition, the number of layers of the text encoder is set to LT to adapt to the different requirements of language feature extraction, while the number of layers of the visual encoder is set to LV to ensure that the model can fully learn and represent multimodal information.

[0178] BERT Encoder: Learning a template t given a dynamic prompt i , and use BERT to encode it into a high-dimensional representation E(t i ), effectively mapping the input sequence into an embedding space suitable for subsequent computing tasks, the formula is:

[0179] E(t i )∈R L×d =BERT(t i )

[0180] Where L is the length of the input sequence and d is the embedding dimension. Then the encoded sequence E(t i )Apply long attention.

[0181] The query, key and value matrices are calculated in the model of the present invention as follows:

[0182]

[0183] in, The weight matrices of query, key and value, respectively. h is the dimension of each attention head, each attention head in the multi-head attention mechanism l It is expressed as:

[0184] head l =Attention(Q (l) ,K (l) ,V (l) )

[0185] in

[0186] The outputs of all attention heads are concatenated and linearly transformed to obtain:

[0187] T0=E(t i )+pos,

[0188]

[0189] Where T l is the hidden state of the lth layer of the output text sequence;

[0190] Residual Network Encoder ResNet and Visual Transformer Encoder ViT: In the visual encoder module, ResNet and ViT are used to extract coarse-grained and fine-grained features from the image, respectively. Specifically, ResNet aims to capture global or coarse-grained information, while ViT focuses on extracting local or fine-grained features. The detailed descriptions of the two encoders are as follows:

[0191] The residual network encoder ResNet is responsible for extracting global or coarse-grained features C×H×W from an input image I of size C. The residual network encoder ResNet extracts a feature map represented as:

[0192] F coarse =ResNet(I),

[0193] Where C′ is the number of channels in the final convolutional layer, H′ and W′ represent the height and width of the obtained feature map, respectively. In order to convert the feature map into a set of feature vectors, F coarse Reshape into a 2D matrix of size (H′×W′)×C′:

[0194] Z coarse =Reshape(F coarse )

[0195] Where r = H′×W′. This reshaping preserves the spatial information extracted by ResNet and represents each spatial position as a r The feature vector of .

[0196] Visual Transformer Encoder ViT for Fine-Grained Features: For fine-grained feature extraction, a visual transformer (ViT) is used. Given an input image I, ViT first divides the image into non-overlapping patches of size P×P. Assuming the size of the input image is H×W, the total number of patches is given by:

[0197]

[0198] Then each patch is flattened into a length P 2 ×C vector and projected to a fixed embedding dimension d using a linear layer v The resulting patch embedding is represented as:

[0199]

[0200] in represents the embedding of the vth patch, Represents the position encoding corresponding to the vth patch.

[0201] In the multi-head attention mechanism, each attention head hl is responsible for processing different aspects of the input data. The outputs of all attention heads are first concatenated and then integrated through a linear transformation. This integrated output is then fed into a feed-forward neural network (FFN) for further processing. The whole process follows the standard Transformer architecture to process the embedded representation of image patches.

[0202] At layer l, the intermediate output of each patch is calculated as:

[0203]

[0204] in is the output of the multi-head attention mechanism of the vth patch in the lth layer, l = 1,...,L v ;

[0205] Subsequently, a feed-forward neural network refines the patch embedding as:

[0206]

[0207] After processing all layers, the final representations of all patches are concatenated to obtain a comprehensive fine-grained feature representation:

[0208]

[0209] in

[0210] Multi-granularity cross-modal aggregation (MCA): The image and text features extracted by the feature encoder can be directly interacted and fused, but this method may introduce irrelevant visual noise. In particular, in the process of cross-modal interaction and fusion, it may be difficult for the model to effectively learn a unified representation. This oversimplified fusion method can increase the model's sensitivity to noise, thereby affecting the overall performance. In order to solve these problems, the present invention proposes a multi-granularity cross-modal aggregation method that closely links image features with text features. This method not only focuses on the global coarse-grained features of the image, but also pays special attention to precise fine-grained features. By interacting these features with the encoded dynamic prompt learning template features, the present invention enhances the synergy of information of different granularities. The method of the present invention allows for a more comprehensive capture of key information, improves the performance of the model in cross-modal tasks, and achieves better multimodal information integration.

[0211] In order to promote better interaction between features from different modalities, the present invention applies two linear layers to map coarse-grained and fine-grained visual features and text features into a unified dimensional space. This process ensures that features can effectively interact and fuse in the same space, thereby enhancing the overall performance of the model. coarse , fine-grained visual features Z fine and text features T l , apply the following formula to convert:

[0212] z′ coarse =f1(z coarse )

[0213] z′ fine =f2(z fine )

[0214] x′ T =f3(T l )

[0215] Among them, f1, f2, and f3 represent respective linear transformations that map features to the same dimensional space, thereby facilitating subsequent interaction and fusion.

[0216] Then, the present invention uses the global context enhancement mechanism to enhance the attention mechanism by integrating global context information, which enables the model to consider both local interactions and broader context information during feature fusion. Through the global pooling operation, from the text feature x′ T , coarse-grained feature z′ coarse and fine-grained features z′ fine Extract the global context GT from:

[0217] GT = GlobalPool(x′ T ,z′ coarse ,z′ fine ),

[0218] where x′ T ∈R m×d , z′ coarse ∈R n×d , z′ fine ∈R keyN×d They are query vector, key vector and value vector, m is the number of queries, keyN is the number of keys, d is the dimension of the feature; GlobalPool represents the global pooling operation; the global pooling operation (GlobalPool) can use average pooling, maximum pooling and other techniques, or more complex global feature extraction methods.

[0219] Next, during the attention mechanism, the global context C is incorporated into the calculation of the attention weights. Specifically, the global context is added to the Query Q and Key K features before calculating the attention score:

[0220]

[0221] After calculating the attention weight A, calculate the value vector z′ fine The weighted sum Z of:

[0222] Z=A·z′ fine ,

[0223] Where Z is a new representation that integrates local feature interactions and global context. To further improve the flexibility of the model, an adaptive mechanism can be introduced to allow the model to dynamically adjust the influence of the global context. A learnable parameter α can be used to adjust the influence of the global context:

[0224]

[0225] This allows the model to fine-tune the degree to which the global context influences the attention mechanism, thereby enhancing its adaptability to different tasks or input data.

[0226] Loss function: In order to optimize the performance of the model in multiple multimodal tasks, the present invention designs a multi-task joint loss function, including loss terms for multimodal link prediction, multimodal relationship extraction, and multimodal named entity recognition.

[0227] Multimodal link prediction loss:

[0228] For the multimodal link prediction task, the goal of the present invention is to predict the tail entity given the head entity and the relation, and vice versa. To achieve this goal, the present invention defines a negative log-likelihood loss function to measure the prediction accuracy. Specifically, for a given triple (h, r, t), the model generates the embedding of the head entity h, the relation r, and the tail entity t based on the multimodal information, and then calculates the prediction score. The loss function is defined as follows:

[0229] ζ link =-∑ (h,r,t)∈PS logσ(f(h,r,t))-∑ (h′,r′,t′)∈PS logσ(-f(h′,r′,t′)),

[0230] Where PS represents the positive sample set in the training data, (h′, r′, t′) represents the negative sample set, σ is the sigmoid function, and f() is the scoring function based on triples.

[0231] Loss of multimodal relation extraction: For the multimodal relation extraction task, the present invention uses the cross entropy loss function to optimize the model's ability to predict relation categories. Given an input entity pair (e1, e2) and the corresponding relation label r), the cross entropy loss is defined as:

[0232]

[0233] Where RN represents the number of relation categories, y i is the one-hot encoded label of the target relation, p i is the probability distribution of the predicted relation category.

[0234] Loss for Multimodal Named Entity Recognition:

[0235] In the named entity recognition task, cross entropy loss is used to optimize entity category prediction. Given an input sequence X = {X1, X2, ..., X IE} and the corresponding entity tag Y = {Y1, Y2, ..., Y IE}, where each label corresponds to a word or phrase in the sequence, then the loss function ζ NER Defined as:

[0236]

[0237] Where IE represents the length of the input sequence, EN is the number of entity categories, and y Li,Tj is the target label of category Tj at the Li-th position, p Li,Tj is the predicted probability of category Tj at the Lith position; X IE is the IEth element in the input sequence, Y IE It is the IEth element in the entity tag;

[0238] Overall Loss Function: This paper adopts a joint loss strategy to integrate the loss functions of multimodal link prediction, multimodal relationship extraction, and multimodal named entity recognition to enhance the overall performance of the model. The final loss function is defined as follows:

[0239] ζ total =λ link ζ limk +λ relation ζ relation +λ NER ζ NER ,

[0240] where λ link , relation and λ NER is a hyperparameter that controls the loss weight.

[0241] In the embodiments of the present invention, two widely used datasets, FB15k-237-IMG and WN18-IMG, are used for link prediction tasks. The statistical information of the dataset is shown in Table 1. Then, experiments on multimodal relationship extraction and multimodal named entity recognition tasks are performed on the MNRE and Twitter-2017 datasets, respectively. The statistical information of the dataset is shown in Tables 2 and 3. The MNRE dataset is constructed from three sources: two multimodal named entity recognition datasets (Twitter15 and Twitter17) and other data captured from Twitter. The MNRE dataset contains 15,484 samples, 9,201 images, and covers 23 relationship categories. The MNRE dataset includes text and image posts, and entities and their relationships are manually labeled by well-educated annotators.

[0242] Table 1

[0243] Dataset Number of relationships Number of entities Number of training sets Development set number Number of test sets FB15k-237-IMG 237 14541 272115 17535 20466 WN18-IMG 18 40943 141442 5000 5000

[0244] Table 2

[0245] Statistics Vocabulary Number of sentences Number of instances Number of entities Number of relationships Number of pictures MNRE 258k 9201 15485 30970 23 9201

[0246] Table 3

[0247] Entity Type Training set Development set Test Set people 2943 626 621 Place 731 173 178 organize 1674 375 395 Miscellaneous 701 150 157 total 6049 1324 1351 Training subset 3373 723 723

[0248] In order to evaluate the effectiveness of the present invention, the present invention is compared with traditional text-based models to show the improvements brought by incorporating visual information. In addition, another set of previous SOTA multimodal methods for multimodal knowledge graph completion models are further considered. For multimodal link prediction, ViL-bERT, IKRL, TransAE, RSME, and MKGformer were selected. For MRE and MNER tasks, TSPNet, MNER-qg, UMGF, MEGA, MKGformer, and TSVFN were selected. The basic framework of the model of the present invention is executed using pytorch, with a batch size of 64 and a maximum training epoch of 30. The initial learning rate of 0.001 and the Adam optimizer are used to train the model parameters and layer 1 in the set {4,5,6,7,8}. For the multimodal link prediction task, the evaluation indicators are selected as: MR and Hits@n; Hits@n represents the ratio of the ranking of the actual answer in the prediction to be equal to or smaller than K, and is specifically defined as follows:

[0249]

[0250] The higher the Hits@n value, the better the prediction performance of the model. The numerator |{q∈Q,q≤n}| in the formula represents the number of queries whose actual answers are in the top n positions, and the denominator |Q| is the total number of queries. K is a threshold used to determine the ranking range of interest. For example, if K=10 is set, then Hits@10 will calculate the proportion of queries whose correct answers are in the top 10 positions of the predicted results.

[0251] MR represents the average ranking of the actual answer of each prediction task in the prediction results, and is specifically defined as follows:

[0252]

[0253] Where Q is the set of queries and |Q| is the rank of the actual answer to query Q. The value of MR ranges from 1 to infinity, with smaller values ​​indicating better prediction performance.

[0254] For the MRE and MNER tasks, in order to evaluate the performance of the model of the present invention, the present invention uses accuracy, precision, recall and F1 score as evaluation indicators:

[0255]

[0256] Among them, tp is true positive, fp is false positive, tn is true negative, and fn is false negative.

[0257] The experimental results of the multimodal link prediction task across two datasets (FB15k-237-IMG and WN18-IMG, FB15k-237-IMG is an extended multimodal knowledge edge graph dataset based on the standard FB15k-237 dataset, and WN18-IMG is an extension of the WN18 dataset based on wordnet) are shown in Table 4, comparing the proposed model with existing unimodal and multimodal methods. Among the unimodal methods, RotatE shows superior performance on both datasets. However, compared with the multimodal methods, the unimodal methods perform poorly overall, highlighting the limitations of unimodal models in handling different data inputs. In the comparison of multimodal methods, the model of the present invention improves Hits@1, Hits@10 and MR on the FB15k-237-IMG dataset by 6.43%, 2.62% and 9.8% respectively over the existing SOTA model. In addition, on the WN18-IMG dataset, the method of the present invention also achieves significant performance improvements on multiple evaluation indicators. These results demonstrate that the proposed dynamic cue learning template combined with the multi-granularity cross-modal aggregation method effectively integrates contextual information and enhances the fusion of text and visual data, leading to superior performance in multimodal tasks.

[0258] Table 5 gives the experimental results in relation extraction and named entity recognition. In the relation extraction task, the experimental results on the MNRE dataset show that traditional unimodal text-based methods show different performances on different indicators, but overall, they fail to reach the performance level of multimodal methods. The advantage of multimodal methods is that they can make full use of information from different modalities and enhance the robustness and expressiveness of the model through cross-modal collaboration. In contrast, the model proposed in the present invention shows a significant improvement in the F1 score among multimodal methods. Specifically, compared with the most advanced multimodal relation extraction model, the model of the present invention achieves an F1 score improvement of 1.42%. This shows that the designed prompt template effectively enriches the semantic representation of the relational information between entities and significantly improves the performance of the model in the relation extraction task.

[0259] In the named entity recognition (NER) task evaluated on the Twitter-2017 dataset, although unimodal text-based methods show strong performance, there is still much room for improvement compared to multimodal methods. Compared with existing multimodal methods, the proposed model improves by about 1.19% in F1 score, surpassing the current state-of-the-art models such as MKGformer and TSVFN. This shows that the proposed multi-granularity interaction mechanism effectively integrates textual and visual information, optimizes the synergy between different modes, and significantly improves the performance of the model in the MNER task.

[0260] Table 4

[0261]

[0262]

[0263] Table 5

[0264]

[0265]

[0266] The present invention provides a multimodal knowledge graph completion method based on dynamic prompt learning and multi-granularity aggregation. The dynamic prompt learning template enables the model to adapt to various MKGC tasks with higher flexibility and consistency. At the same time, multi-granularity cross-modal aggregation effectively fuses coarse-grained and fine-grained visual features with text features, solving the challenges associated with modal conflicts. A large number of experiments have demonstrated that the model of the present invention outperforms existing state-of-the-art methods and demonstrates significant improvements in key performance indicators. This work not only promotes the development of the MKGC field, but also provides valuable insights for future research on multimodal data integration and reasoning.

[0267] The present invention provides a multimodal knowledge graph completion method based on dynamic prompt learning and multi-granularity aggregation. There are many methods and ways to implement the technical solution. The above is only a preferred implementation of the present invention. It should be pointed out that for ordinary technicians in this technical field, without departing from the principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention. All components not specified in this embodiment can be implemented by existing technologies.

Claims

1. A multimodal knowledge graph completion method based on dynamic prompt learning and multi-granularity aggregation, characterized in that: The Transformer architecture is optimized, dynamic prompt templates and multi-granularity cross-modal aggregation are added, and an improved Transformer model is obtained for multimodal knowledge graph completion. The specific steps include: Step 1, generate dynamic prompt template: select the appropriate template structure according to the task requirements and use the adaptive guidance mechanism to dynamically adjust the template structure; Step 2: Establish a multimodal feature encoder to convert text and image data into feature vectors for training this model; Step 3: Multi-granularity cross-modal aggregation (MCA): fuse features of different modalities and granularities; Step 4: Design a multi-task joint loss function to optimize the performance of the model in more than two multimodal tasks and perform training; Step 1 includes: given an entity description T = {V1, V2, ..., Vn} and an associated image I = {I1, I2, ..., In}, design the following dynamic prompt learning template t i : t i =[CLS][Entiry1][SEP][Relation][SEP][Entity2][SEP] [V1][V2]...[Vn][SEP][MASK][SEP], Among them, [CLS] is a special token used to indicate the classification information of the entire input sequence; [Entity1] indicates the first entity, which is suitable for named entity recognition (NER) of entity recognition, relation extraction (RE) of head entities, and link prediction of head entities; [SEP] is used to separate different parts in a sentence; [Relation] indicates the relationship between entities, which is used for relation identification in RE tasks and for relation identification in link prediction; [Entity2] indicates that the second entity is used as the tail entity in RE tasks and as the tail entity in link prediction; Vn indicates the nth element in the entity description, and In indicates the nth element in the associated image; The token [MASK] is used to represent the information that the improved Transformer model needs to predict; An adaptive task guidance mechanism is used to dynamically adjust the prompt design of the improved Transformer model. The mechanism modifies the structure of the prompt template according to the task requirements. In the NER task, the structure of the [Entity1] part is reconstructed according to the context details, and the entity attributes are reordered or extended. In the relation extraction task, the [relation] part is reorganized by adjusting the representation and positioning. In step 3, the multi-granularity cross-modal aggregation MCA includes: Two linear layers are applied to map the coarse-grained and fine-grained visual features and text features into a unified dimensional space. For the original coarse-grained visual feature Z coarse , fine-grained visual features Z fine and text features T l , apply the following formula to convert: z′ coarse =f1(z coarse ), from fine =f2(z fine ), x′ T =f3(T l ), Where f1 is the first linear transformation function, which maps zcoarse to the target dimension space; z′ coarse is the transformed coarse-grained visual feature; f2 is the second linear transformation function, which maps zcoarse to the target dimension space; z′ fine is the transformed fine-grained visual feature; f3 is the third linear transformation function, which maps zcoarse to the target dimension space; x′ T is the transformed text feature; Then, through the global pooling operation, from the text feature x′ T , coarse-grained feature z′ coarse and fine-grained features z′ fine Extract the global context GT from: GT=GlobalPool(x′ T ,z′ coarse ,z′ fine ), where x′ T ∈R m×d , z′ coarse ∈R n×d , z′ fine ∈R keyN×d They are query vector, key vector and value vector, m is the number of queries, keyN is the number of keys, d is the dimension of the feature; GlobalPool represents the global pooling operation; Add the global context to the query vector x′ T and the key vector z′ coarse In the feature, the attention weight A is obtained: Where T represents transpose; Calculate the value vector z′ fine The weighted sum Z of: Z=A·z′ fime , A learnable parameter α is used to adjust the influence of the global context and obtain the final attention weight A f :

2. The method according to claim 1, characterized in that In step 1, the dynamic prompt template includes a multimodal link prediction MLP template, a multimodal named entity recognition MNER template and a multimodal relationship extraction MRE template: The multimodal link prediction MLP template is expressed as: [CLS][Entity1][SEP][Re1][SEP] [MASK][SEP][V1][V2]...[Vn][SEP], where the token [MASK] is used to predict the missing entity associated with the first entity [Entity1] through the relation [Rel]; The multimodal named entity recognition MNER template is expressed as: [CLS][MASK][SEP][Re1][SEP] [MASK][SEP][V1][V2]...[Vn][SEP], For the named entity recognition task, the token [MASK] is used to identify and classify entities in textual and visual contexts; The multimodal relation extraction MRE template is expressed as: [CLS][Entity1][SEP][Re1][SEP] [V1][V2]...[Vn][SEP].

3. The method according to claim 2, characterized in that In step 2, the multimodal feature encoder includes a visual encoder and a text encoder; the visual encoder is responsible for encoding the synchronized granular visual features, and the text encoder is responsible for encoding the text features; The text encoder is used to extract text features Q, including: Given a dynamic prompt learning template t i , using a transformer-based bidirectional encoder BERT to dynamically prompt the learned template t i Encoded into a high-dimensional representation E(t i ), the formula is: Where L is the length of the input sequence and d is the embedding dimension; Then the coding sequence E(t i ) Apply the multi-head attention mechanism and query the matrix Q (l) , key matrix K (l) Sum value matrix V (l) The calculation formula is: in are the weight matrices of query, key, and value, respectively. d h For the dimension of each attention head, each attention head in the multi-head attention mechanism l It is expressed as: head l =Attention(Q (l) ,K (l) ,V (l) ), in Attention is a scaled dot product attention, calculated as: Among them, QK T is the dot product of the query and the key, d k is the dimension of the key, the softmax function is used to convert the result of the dot product into a probability distribution, and finally multiply it by the value matrix V to get the weighted output; The outputs of all attention heads are concatenated and linearly transformed to obtain: T0=E(t i )+pos, Where pos is the positional encoding; T0 is the initial hidden state of the input sequence; MHA is multi-head attention, LN is layer normalization, FFN is a feedforward neural network, is the intermediate hidden state of layer l, T l-1 is the hidden state of the previous layer; T l is the final hidden state of layer l, is the intermediate hidden state; In the visual encoder, the residual network encoder ResNet and the visual transformer encoder ViT are used to extract coarse-grained features K and fine-grained features V from the image respectively: The residual network encoder ResNet is responsible for extracting global coarse-grained features F from the input image I of size C×H×W coarse , expressed as: F coarse =ResNet(I), in C′ is the number of channels of the final convolutional layer, H′ and W′ represent the height and width of the obtained feature map respectively; C, H, W represent the channel input, height and width of the input image respectively; In order to convert the feature map into a set of feature vectors, F coarse Reshape into a two-dimensional matrix Z of size (H′×W′)×C′ coarse : Z coarse =Reshape(F coarse ), Reshape means reshaping a tensor or matrix; r = H′×W′; r represents the number of reshaped feature vectors, that is, the total number of spatial locations of the feature map; The visual transformer encoder ViT is responsible for extracting local coarse-grained features from an input image I of size C×H×W: Given an input image I, the visual transformer encoder ViT first divides the image into non-overlapping patches of size P×P, where P represents the height and width of the patches into which the input image is divided. The total number of patches v is given by: Then flatten each patch into a length of P 2 ×C vector and projected to a fixed embedding dimension d using a linear layer v , the obtained patch embedding representation for: in represents the embedding of the vth small block, Indicates the position code corresponding to the vth small block; In the lth layer, the calculation formula for each small block intermediate output is: in is the output of the multi-head attention mechanism of the vth small block in the lth layer, l = 1,...,L v ; L v Indicates the total number of layers; is the output of the l-1th layer; Express Application layer normalization; is the output of the multi-head self-attention mechanism; is the original input; is the output after multi-head self-attention processing; Subsequently, a feed-forward neural network FFN is used to refine the patch embedding as: in, is the output of the feedforward neural network; is the final output; After processing all layers, the final representations of all small patches are concatenated to obtain a comprehensive fine-grained feature representation: Where Z fine represents the final fine-grained feature representation, Concat is a concatenation operation; represents the final representation of the i-th small block after being processed by the Lv layer Transformer; n is the total number of small blocks.

4. The method according to claim 3, characterized in that In step 4, the loss function is described as follows: For a given triple (h, r, t), the improved Transformer model generates embeddings of the head entity h, relation r, and tail entity t based on multimodal information, and then calculates the prediction score. The loss function ζ link is defined as: ζ link =-∑ (h,r,t)∈PS logσ(f(h,r,t))-∑ (h′,r′,t′)∈PS logσ(-f(h′,r′,t′)), Where PS represents the positive sample set in the training data, (h′, r′, t′) represents the negative sample set, σ is the sigmoid function, and f() is the scoring function based on triples. For the multimodal relation extraction task, the cross entropy loss function is used to optimize the ability of the improved Transformer model to predict relation categories. Given an input entity pair (e1, e2) and the corresponding relation label r, the cross entropy loss ζ relation Defined as: Where RN represents the number of relation categories, y i is the one-hot encoded label of the target relation, p i is the probability distribution of the predicted relation category; In the named entity recognition task, cross entropy loss is used to optimize entity category prediction. Given an input sequence X = {X1, X2, ..., X IE } and the corresponding entity tag Y = {Y1, Y2, ..., Y IE }, where each label corresponds to a word or phrase in the sequence, then the loss function ζ NER Defined as: Where IE represents the length of the input sequence, EN is the number of entity categories, and y Li,Tj is the target label of category Tj at the Li-th position, p Li,Tj is the predicted probability of category Tj at the Lith position; X IE is the IEth element in the input sequence, Y IE It is the IEth element in the entity tag; The final loss function ζ total Defined as: g total =λ link g link +λ relation g relation +λ NERNER , where λ link , relation and λ NER is a hyperparameter that controls the loss weight.

5. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores program codes, and when the program codes are executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 4.

6. A storage medium, characterized in that: A computer program or instruction is stored, and when the computer program or instruction is run on a computer, the steps of the method according to any one of claims 1 to 4 are executed.

Citation Information

Patent Citations

  • Multi-modal aspect-level sentiment analysis method based on text and image gating fusion mechanism

    CN117131433A

  • NPC interaction method and device, storage medium and electronic device

    CN117732070A