Multi-modal Irony Intention Recognition Method Based on Knowledge Injection and Dual Attention Network
By injecting implicit context knowledge into multimodal satire detection and using dual attention networks to construct a complete semantic representation of multimodal information, the low accuracy and noise problems in existing methods are solved, and higher accuracy and interpretability of ironic intention recognition are achieved.
Patent Information
- Application Number
- CN202210863424.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-21
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-07-21
AI Technical Summary
The existing multimodal satirical detection methods fail to effectively utilize implicit context information, resulting in a decrease in recognition accuracy, and ignore the full amount of information between text and pictures, introduce noise, and affect the actual application of the model.
Using a dual attention network based on knowledge injection, implicit context knowledge is injected into multimodal input through a multi-dimensional attention module. The dual attention network is used to collaborate on the picture and text attention modules, capture shared semantics, and distinguish semantic differences through multi-dimensional cross-modal matching layer to construct a complete semantic representation of multimodal information.
提高了讽刺意图识别的准确性和可解释性,能够更精准地定位讽刺描述区域,增强模型的实际应用效果。
Smart Images

Figure CN115408517B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for identifying multi-modal ironic intentions based on a dual attention network with knowledge injection, belonging to the technical field of multi-modal information recognition. Background Art
[0002] Ironic expressions based on multi-modal information achieve the purpose of implicitly expressing strong emotions by using text that is contrary to the metaphorical scene in the picture. Currently, ironic expressions based on text and pictures are prevalent on social platforms such as Weibo and Twitter. Since ironic expressions can reverse the polarity of emotions or opinions in the text, automatically detecting multi-modal ironic intentions is of great significance in customer service, opinion mining, and various tasks that require understanding people's true emotions.
[0003] Detecting multi-modal ironic intentions in practice is very complex. The semantic expression of the user input information is affected not only by the explicit content but also by the implicit context. The explicit content refers to the observable scene content in the input text or picture, and the implicit context refers to the invisible reasoning knowledge about the scene in the input information, including the process of the scene development and the intentions of the people in the scene. Based on the complete semantic representation of text and pictures, identifying ironic intentions requires accurately locating the parts that describe irony in multi-modal information and discriminating their semantic differences. However, existing multi-modal irony detection methods only learn features from the input text and pictures, ignoring the modeling of the implicit context behind the content. At the same time, they model the semantic differences between multi-modal information based on the unprocessed full amount of information of text and pictures, which is easy to introduce noise, resulting in a decrease in the accuracy of ironic intention recognition and affecting the practical application of the model. How to inject implicit context information into multi-modal input to obtain better feature representations and accurately locate ironic description regions based on this information for semantic difference recognition is an urgent problem to be solved. Summary of the Invention
[0004] In view of the above problems, the method for identifying multi-modal ironic intentions based on a dual attention network with knowledge injection proposed by the present invention uses a knowledge-enhanced multi-dimensional attention module to inject implicit context knowledge into the multi-modal input representation. According to the human reasoning method, the implicit context knowledge is divided into two perspectives, namely the scene state and the emotional state, to construct a complete semantic representation of multi-modal information. At the same time, a dual attention network is used to collaboratively execute the picture and text attention modules based on a joint memory vector, which aggregates the previous attention results to capture the shared semantics related to irony in the text and pictures. Finally, based on the joint embedding space, a multi-dimensional cross-modal matching layer is used to distinguish the differences between multi-modal from multiple dimensions. This helps to improve the overall performance of irony recognition and provides interpretability for the prediction results, facilitating the practical application of the model.
[0005] The technical content of the present invention includes:
[0006] A multi-modal ironic intention recognition method based on a dual attention network with knowledge injection, the method comprising:
[0007] Obtain the data content to be recognized, the data content to be recognized including: a number of <text, picture> pairs, the text containing a number of words i, and the picture involving a number of objects j;
[0008] Encode the word i in the text and the object j in the picture respectively to obtain the original word representation And the original object representation
[0009] Based on the implicit context information of the data content to be recognized, expand the original word representation And the original object representation To obtain the context-aware word representation And the context-aware object representation
[0010] Use a dual attention network to calculate the attention of the original word representation The original object representation And the context-aware word representation The context-aware object representation To obtain the attention calculation results of the original representation and the context-aware representation;
[0011] For the attention calculation results of the original representation and the context-aware representation, by comparing the differences between the text and the picture, obtain the original cross-modal contrast representation and the context-aware cross-modal contrast representation;
[0012] Based on the original cross-modal contrast representation and the context-aware cross-modal contrast representation, calculate the ironic intention recognition result of the data content to be recognized.
[0013] Further, the encoding of the object j in the picture to obtain the original object representation Includes:
[0014] For each picture, use a pre-trained object detector to detect the region of the object j from the picture, and use the pooling feature before the multi-class classification layer as the visual feature representation r of the object j j ;
[0015] Project the visual feature representation r j Into the space of the text representation;
[0016] Obtain the text representation specific to object j by calculating the relevance between each word i in the text and object j
[0017] Based on the text representation And the visual feature representation r j Calculate the representation of object j with text relevance
[0018] Input the object sequence composed of the visual feature representation r j Into a bidirectional gated recurrent neural network, and use the representation As the calculation weight, so as to obtain the original object representation of each object j
[0019] Furthermore, based on the implicit context information of the data content to be recognized, for the original word representation And the original object representation Perform expansion to obtain the word context-aware representation And the object context-aware representation Include:
[0020] Generate different types of inference knowledge for each event description in the picture or text And calculate the inference knowledge Of the common sense inference representation H M,R , where, w l Indicates the word in the inference knowledge, 1 ≤ l ≤ L, L represents the length of the inference knowledge, the relationship type R ∈ {before, after, intent}, before represents the event pre-relationship type, ofter represents the event post-relationship type, intent represents the relationship type of the person's intention in the scene, and the modality M represents the text modality or the picture modality;
[0021] Based on the text feature map H Composed of the original word representation T , the picture feature map H Composed of the original object representation I , and the common sense inference representation H M,R , calculate the association matrix C Between the data content to be recognized and the inference knowledge M ;
[0022] Based on the association matrix C M , obtain the representation of the original word representation With the implicit context information of the text of the original object representation Representation with implicit context information of text and images
[0023] By calculating a relevant weight for each of the inference knowledge calculate the representation with the representation of the enhanced representation with the enhanced representation
[0024] Based on the enhanced representation with the enhanced representation calculate the word context-aware representation with the object context-aware representation wherein, the word-aware vector representation includes: the scene state context-aware representation of the word and the emotional state context-aware representation the object context-aware vector representation includes: the scene state context-aware representation of the object and the emotional state context-aware representation
[0025] Furthermore, based on the association matrix C M obtain the original word representation with the original object representation of the text with implicit context information and the representation of the image with implicit context information including:
[0026] Based on the association matrix C M with the common sense reasoning representation H M,R use the attention mechanism to form the word-level representation of the inference knowledge with the object-level representation
[0027] At the original word representation add the word-level representation and the object-level representation respectively to the original object representation to obtain the representation with the representation
[0028] Furthermore, based on the enhanced representation calculate the word context-aware representation including:
[0029] Add the enhanced representation Specifically, it is an enhanced representation of the pre - event relationship type Enhanced representation of the post - event relationship type Enhanced representation of the intention relationship type
[0030] According to the enhanced representation The original representation of the word And the enhanced representation Calculate the context - aware representation of the word scene state
[0031] According to the enhanced representation Obtain the context - aware representation of the word emotion state
[0032] Furthermore, use a dual - attention network to perform attention calculations on the original representation of the word The original representation of the object Perform attention calculations to obtain the attention calculation results of the original representation, including:
[0033] Sum up the original representations of each word To obtain the complete text representation u (0) ;
[0034] Sum up the original representations of each object in the picture To obtain the complete picture representation v (0) ;
[0035] According to the representation u (0) And the representation v (0) Calculate the joint memory vector m (0) ;
[0036] Based on Perform iterative calculations, and after the iteration ends, obtain the joint memory vector m (K) where K represents the total number of iterations, and i (k) represents the complete text representation in the k - th iteration, v (k) represents the complete picture representation in the k - th iteration, is the element - wise product;
[0037] Based on the joint memory vector m (K) And use a dual - attention network to perform attention calculations to obtain the complete text representation u (K+1) And the complete picture representation v (K+1) ;
[0038] Take the complete text representation u (K+1) And the complete picture representation v (K+1) As the attention calculation results of the original representation
[0039] Further, based on the joint memory vector m (K) , the complete text representation u is calculated (K+1) , including:
[0040] According to the joint memory vector m (K) and the original word representation calculate the output of the feed-forward neural network in the dual attention network
[0041] Substitute the output into the softmax function to obtain the attention weights
[0042] Based on the attention weights weighted sum the original word representation to obtain the complete text representation u (K+1) .
[0043] Further, the attention calculation result for the original representation is used to obtain the original cross-modal contrast representation by comparing the differences between the text and the picture, including:
[0044] The original cross-modal contrast representation wherein, represents a trainable weight matrix, || is the absolute value of the element difference, and ; represents the concatenation operation
[0045] Further, based on the original cross-modal contrast representation and the context-aware cross-modal contrast representation, calculate the sarcasm intention recognition result of the data content to be recognized, including:
[0046] Concatenate the original cross-modal contrast representation and the context-aware cross-modal contrast representation;
[0047] Input the concatenation result into a fully connected layer and use the Sigmoid function for binary sarcasm classification to obtain the sarcasm intention recognition result of the data content to be recognized.
[0048] An electronic device, the electronic device includes a memory and a processor, and a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement any one of the above methods.
[0049] The advantages of the present invention compared with the prior art are:
[0050] (1) The method uses VisualCOMET to provide implicit context information for text and image modality information, namely scene state context and emotional state context. It adopts a knowledge-enhanced multi-dimensional attention module to inject the implicit context into the multi-modal input, generating context-aware text and image representations, thereby helping the multi-modal information to construct a complete semantic context.
[0051] (2) The designed dual attention network can accurately locate the areas describing irony in multi-modal information by repeatedly using text and image attention mechanisms. At the same time, the dual attention is applied to the original representation and the context-aware representation respectively to capture multi-modal information representations from multiple angles. Based on the multi-dimensional shared representation space, a multi-dimensional cross-modal module is used to distinguish the semantic differences between text and images, thereby accurately identifying the ironic intention.
[0052] (3) The present invention has higher performance compared with existing methods. At the same time, the dual attention module combined with the injected knowledge can provide interpretability for the predicted structure. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 Flowchart of the system model of the present invention.
[0054] Figure 2 Architecture diagram of the system model of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0055] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only specific embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0056] The technical problem solved by the present invention: A multi-modal ironic intention recognition method based on a dual attention network with knowledge injection is proposed. Aiming at the problem of multi-modal ironic intention recognition, on the one hand, a knowledge-enhanced multi-dimensional attention model is used to construct a complete semantic representation of multi-modal information. On the other hand, a dual attention mechanism is used to maintain a shared vector, extract the shared semantics related to irony in multi-modal information through text and image attention modules, and model the differences in multi-modal scenarios containing implicit context through a multi-dimensional cross-modal matching layer, which helps to improve the overall performance of irony recognition and also provides a certain degree of interpretability for the prediction results.
[0057] Technical solution of the present invention: A multi-modal sarcasm intention recognition method based on a dual attention network with knowledge injection injects implicit context knowledge into the multi-modal representation to construct the complete semantics of multi-modal information, and uses a dual attention network to capture the areas describing sarcasm in the multi-modal for multi-dimensional semantic comparison. In the model, the text and the picture are first input and encoded into vector representations, and the attention mechanism is used to align the objects in the text and the picture, so as to filter out the irrelevant information in the picture; then, in order to supplement the implicit context information lacking in the text and the picture, an event knowledge graph is used to generate scene context and sentiment context for the text and the picture, and the obtained knowledge is injected into the multi-modal input through a knowledge-enhanced multi-dimensional attention module to construct the complete semantic encoding of multi-modal information; in order to focus on the areas describing sarcasm in the text and the picture, a dual attention module that collaboratively executes the text and the picture is proposed, and the shared semantics across multiple modalities are captured in the original encoding and the complete semantic encoding of multi-modal information by maintaining a joint memory vector; based on the joint embedding space, multi-dimensional cross-modal matching is adopted to distinguish the multi-modal differences in multiple dimensions; finally, multiple comparison results are concatenated and input into the classification for multi-modal sarcasm detection.
[0058] Figure 1 The flow chart of the system model of the present invention is as Figure 1 shown. The system of the present invention includes five major parts: an input encoding module, a knowledge injection module, a dual attention interaction module, a multi-dimensional cross-modal matching module, and a classification prediction module.
[0059] First, in the encoding module, for word preprocessing, the NLTK toolkit is used to separate the text sentences to obtain words, and the word embedding uses 200-dimensional vectors generated by the Glove algorithm as initialization; for picture preprocessing, Faster-RCNN is used to extract the picture objects and their features, and repair is performed in the training stage. The hidden size of the Bi-GRU is 512 dimensions.
[0060] In the knowledge injection module, for event reasoning knowledge, the VisualCOMET event reasoning generator is used to generate reasoning knowledge of three types of relationships, namely "before", "after", and "intent", for text and images. For example, the input text is "This is why I want to be a mom", and the image description is "A woman holding a broom stands inside the house". The inferences of the three types of relationships generated by VisualCOMET for the text are: before: [be in a family roo, put on her schooluniform,...], after: [play games with friends, tell them wrong,...], intent: [stay at home cozy, make her mom happy,...]; the inferences of the three types of relationships generated for the image are: before: [put on an apron, be from a school,...], after: [clean the room, finish her housework,...], intent: [Cleaning the room with a broom, play games with friends,...]. In the actual application process, 15 candidate items are maintained for the reasoning knowledge of text and images in each relationship type. Subsequently, the knowledge injection module is used to obtain a multi-modal information representation enhanced by implicit context knowledge, and the complete semantic information of text and images is constructed.
[0061] In the dual attention interaction module, the text and image attention mechanisms obtain the representations of the text and image focusing on the ironic regions through 3 interactive iterations.
[0062] In the multi-dimensional cross-modal matching module, the text and image are compared in terms of the original representation, scene state context-aware representation, and emotional state context-aware representation.
[0063] In the classification module, the comparison results of the multi-modal information in the three dimensions are concatenated and input into the output layer composed of a two-layer fully connected network and Softmax for ironic intent classification. During training, the batch size is 32, the model training learning rate is 0.0005, and Adam is used as the optimizer.
[0064] Figure 2 The system model architecture diagram of the present invention is as Figure 2 shown:
[0065] · Input encoding module
[0066] Extract the feature of the input text and image modal information and encode it into a unified vector space. Its input is a sentence containing a series of words and an image containing multiple objects; the output is a vector representation after the following operations. The input encoding module consists of the following two parts:
[0067] (1) Text encoding module:
[0068] Given a sequence of text words {w1, w2, …, w N}, to obtain the semantic information of the sentence, a bidirectional gated recurrent neural network (bi-GRU) is used to learn the sequence semantic information representation of the words in the text and encode it into a vector form:
[0069]
[0070] where, is the hidden state representation of the i-th word output by the bi-GRU cell, and N is the number of words in the sentence, that is, the sentence length. After the word sequence is encoded by the bi-GRU, the original representation is
[0071] (2) Image encoding module
[0072] There are multiple objects involved in the image. In order to filter out the irrelevant information in the image and avoid the problem of incomplete object semantics caused by equal division and cutting of the image, directly extract the objects related to the text in the image and perform feature representation.
[0073] For each image I, use the pre-trained object detector - Faseter R-CNN to detect D significant objects from the image, and use the pooling features before the multi-class classification layer as the feature representation of the objects. Subsequently, project the visual features of the extracted objects into the space of the text representation.
[0074] r j = ReLu(W v r j + b v ),
[0075] where, r j is the visual feature representation of the j-th detected object, W v is the weight matrix, and b v is the bias parameter.
[0076] The image is used as the background information of the input text, and the text only involves some of the objects in the image. In order to suppress the negative impact caused by the irrelevant information in the image, a gated attention mechanism is used to align the text and the image by calculating the correlation at the word and region levels. For each object in the image, the gated attention mechanism uses the soft attention mechanism to calculate the correlation between each word in the text and the object and form a text representation specific to the object Then is multiplied element - by - element with the visual feature representation r j to obtain a representation for each target object with text relevance by performing element - wise multiplication:
[0077]
[0078]
[0079]
[0080] Since there is no natural order in the picture area, the scattered information in the picture is concatenated into a complete semantic expression through a bidirectional gated recurrent neural network bi - GRU.
[0081]
[0082] The original representation of the objects in the picture after being encoded by bi - GRU is
[0083]
[0084] where D is the number of objects recognized by the object detector in the picture.
[0085] · Knowledge injection module
[0086] To construct a complete semantic representation of text and pictures, implicit context information is used to naturally expand multi - modal information, thereby forming a multi - view knowledge - rich multi - modal feature representation. The knowledge injection module consists of the following two parts:
[0087] (1) Knowledge acquisition module:
[0088] The Visual - Text Event Reasoner VisualCOMET is used to provide commonsense knowledge reasoning in two dimensions of scene state context and emotional state context for the input text and pictures. VisualCOMET uses the pre - trained autoregressive language model GPT - 2 as a generation model. Given a description of a picture or an event, it can generate reasoning knowledge about three relation types: before, after, and intent, that is, the context before and after the event (scene state context) and the intent of the people in the scene (emotional state context). Commonsense knowledge reasoning usually consists of short sentences of a series of words. The reasoning knowledge about different modalities is defined as where R represents three relation types, R ∈ {before, after, intent}, and M represents text and picture modalities, M ∈ {T, I}.
[0089] The bidirectional gated recurrent network (bi - GRU) is used to process the short sentences to obtain their representations. L is the sentence length of the reasoning knowledge.
[0090] (2) Multi-dimensional knowledge injection module:
[0091] Based on common sense knowledge reasoning from different perspectives, a knowledge-aware attention layer is designed to form a multi-dimensional knowledge-aware multi-modal representation. First, each element in the knowledge reasoning query text or image is utilized and its relevance is calculated to align the elements in the text or image with the knowledge reasoning. Specifically, given the multi-modal feature representation H M (Text feature mapping or image feature mapping ) and the common sense reasoning representation the relevance between the input and the reasoning knowledge is calculated, i.e., the association matrix C M :
[0092] C M = tanh(H M W M (H M,R ) T )
[0093] where W M is the weight matrix.
[0094] Subsequently, the attention mechanism is used to form a word-level representation of the reasoning knowledge regarding the input features, and it is added to the original representation of the input features to obtain the representation of the text with implicit context information and the representation of the image with implicit context information
[0095]
[0096]
[0097] Since the event reasoner VisualCOMET generates multiple candidate reasoning knowledge for the text and image. To focus on the reasoning more relevant to the input scenario, a relevant weight is learned for each knowledge reasoning, and they are weighted and summed to generate a representation enhanced with multi-modal information knowledge:
[0098]
[0099]
[0100] where WM ,R is the weight matrix and Q is the number of candidate reasoning knowledge.
[0101] The scene state context consists of before the scene, during the scene, and after the scene. Therefore, the average of the three states is taken to obtain the word vector representation of the input's scene state context awareness:
[0102]
[0103] Accordingly, the picture vector representation of the scenario state context awareness is Similarly, the present invention can obtain the text representation of the emotional state context awareness and the picture representation
[0104] · Dual attention module
[0105] To locate the parts describing irony in the text and the picture, a joint memory vector is created, and the text attention and picture attention mechanisms are executed iteratively multiple times to collect the shared information related to irony in both modalities. Based on the dual attention mechanism, we can obtain the representations of the text and the picture focused on specific regions. The dual attention mechanism is executed in three aspects of the original representation and the context awareness representation of the multi-modal. For simplicity of expression, the description of the relationship type is omitted in the following description, that is simplified to u i , simplified to v j .
[0106] The dual attention module includes the following three sub-modules:
[0107] (1) Shared vector
[0108] The key to identifying irony in multi-modal information is to find the joint space describing the same thing, that is, the region describing irony. For this purpose, a joint memory vector is designed to collect the information identified in k iterations in the text and the picture:
[0109]
[0110] where, v (k) and u (k) are the complete representations of the picture and the text, and the initial memory representation m (0) is defined as the element-wise product of v (0) and u (0) .
[0111]
[0112]
[0113] (2) Text attention mechanism
[0114] The text attention mechanism identifies the region describing irony by calculating the attention weights of each word in the text with the joint memory vector to measure the relevance of each part in the text related to irony. Specifically, the attention weight Calculated by a two - layer feed - forward neural network and the softmax function:
[0115]
[0116]
[0117] Among them, and are model parameters, and are bias parameters.
[0118] Finally, the complete representation of the text is obtained through weighted summation:
[0119]
[0120] (3) Image attention mechanism
[0121] The calculation process is the same as that of the text attention mechanism. First, use a two - layer feed - forward neural network and the softmax function to calculate the regions in the image related to the joint memory vector (i.e., the shared semantics related to irony identified after k iterations), and obtain the complete representation of the image through weighted summation:
[0122]
[0123]
[0124]
[0125] Among them, and are model parameters, and are bias parameters.
[0126] The dual - attention module will obtain the representations highlighting the ironic parts in the text and the image after K iterations, defined as u and v.
[0127] · Multi - dimensional cross - modal matching module
[0128] To capture the semantic differences between the text and the image, use the following deep comparison attention mechanism to compare the differences between the text and the image in terms of multi - modal original representations and context - aware representations. The implementation of this module is as follows:
[0129]
[0130] Among them, is element - wise multiplication, || is the absolute value of the element - wise difference, ; is the concatenation operation, is a trainable weight matrix.
[0131] · Prediction module
[0132] Connect the above multi-dimensional cross-modal contrast representations (z raw , z sc , z en ) and input them into the fully connected layer, and use the Sigmoid function for binary irony classification.
[0133] H = fc([z raw ; z sc ; z em ),
[0134] y = Sigmoid(H).
[0135] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Those of ordinary skill in the art can modify or equivalently replace the technical solutions of the present invention without departing from the spirit and scope of the present invention. The protection scope of the present invention shall be subject to what is described in the claims.
Claims
1. A multi-modal ironic intention recognition method based on a dual attention network with knowledge injection, the method comprising: Obtaining data content to be recognized, the data content to be recognized including: a plurality of <text, image> pairs, the text containing a plurality of words i, and the image involving a plurality of objects j; Encode the word i in the text and the object j in the picture respectively to obtain the original representation of the word and the original representation of the object Based on the implicit context information of the data content to be recognized, for the original representation of the word and the original representation of the object are extended to obtain the word context-aware representation and the object context-aware representation wherein, based on the implicit context information of the data content to be recognized, for the original representation of the word and the original representation of the object are extended to obtain the word context-aware representation and the object context-aware representation including: Generate different types of inference knowledge for each event description in the picture or text And calculate the inference knowledge The commonsense reasoning representation H of M,R , where, w l represents the words in the inference knowledge, 1 ≤ l ≤ L, L represents the length of the inference knowledge, the relation type R ∈ {before, after, intent}, before represents the relation type before the event, after represents the relation type after the event, intent represents the relation type of the person's intention in the scene, and the modality M represents the text modality or the picture modality; Based on the original representation of the word to form the text feature mapping H T and the original representation of the object to form the image feature mapping H I and the common sense reasoning representation H M,R , calculate the correlation matrix C between the data content to be recognized and the reasoning knowledge M ; Based on the association matrix C M , obtain the original representation of the word and the original representation of the object with the representation of the text carrying implicit context information and the representation of the picture carrying implicit context information By learning a relevant weight for each of the inference knowledge calculate the representation and the representation of the enhanced representation and the enhanced representation Based on the enhanced representation With the enhanced representation Calculate the word context-aware representation With the object context-aware representation Wherein, the word perception vector representation Includes: the scene state context-aware representation of the word And the emotional state context-aware representation The object context-aware vector representation Includes: the scene state context-aware representation of the object And the emotional state context-aware representation Use a dual attention network to separately process the original representation of the word the original representation of the object and the context-aware representation of the word the context-aware representation of the object perform attention calculation to obtain the attention calculation results of the original representation and the context-aware representation; For the attention calculation results of the original representation and the context-aware representation, by comparing the differences between the text and the image, obtaining an original cross-modal contrast representation and a context-aware cross-modal contrast representation; Based on the original cross-modal contrast representation and the context-aware cross-modal contrast representation, calculating the ironic intention recognition result of the data content to be recognized.
2. The method according to claim 1, wherein Encoding the object j in the picture to obtain the original representation of the object including: For each image, use a pre-trained object detector to detect the region of object j from the image, and use the pooling features before the multi-class classification layer as the visual feature representation r of object j j ; Project the visual feature representation r j into the space of the text representation; Obtain a text representation specific to object j by calculating the relevance of each word i in the text to object j Based on the text representation and the visual feature representation r j calculate the representation of object j with text relevance Input the sequence of objects composed of the visual feature representation r j into a bidirectional gated recurrent neural network, and use the said representation as the calculation weight, so as to obtain the original object representation of each object j 3. The method according to claim 1, characterized in that Based on the association matrix C M , obtain the original representation of the word and the original representation of the object with the representation of the text carrying implicit context information and the representation of the picture carrying implicit context information including: Based on the association matrix C M and the commonsense reasoning representation H M,R , use the attention mechanism to form a word-level representation of the reasoning knowledge and the object-level representation In the original representation of the word and the original representation of the object add the word-level representation respectively and the object-level representation to obtain a representation and a representation 4. The method according to claim 1, wherein Based on the enhanced representation Calculate the word context-aware representation including: The enhanced representation Specifically, the enhanced representation of the pre - event relationship type The enhanced representation of the post - event relationship type The enhanced representation of the intention relationship type Based on the enhanced representation the original representation of the word and the enhanced representation calculate the context-aware representation of the word According to the enhanced representation to which it belongs Obtain the context-aware representation of the word sentiment state 5. The method according to claim 1, wherein Use a dual attention network to separately perform attention calculations on the original representation of the word the original representation of the object to obtain the attention calculation results of the original representation, including: For the original representation of each word Sum them up to obtain the complete text representation u (0) ; The original representations of each object in the image are summed to obtain the complete representation v of the image (0) ; According to the representation u (0) and the representation v (0) , calculate the joint memory vector m (0) ; Based on perform iterative calculations, and after the iteration ends, obtain the joint memory vector m (K) , where K represents the total number of iterations, u (k) represents the complete text representation of the k-th iteration, v (k) represents the complete image representation of the k-th iteration, is the element-wise product; Based on the combined memory vector m (K) , and using a dual attention network for attention calculation, the complete text representation u (K+1) and the complete image representation v (K+1) are obtained respectively; Fully represent the text as u (K+1) Fully represent the image as v (K+1) The attention calculation result as the original representation.
6. The method according to claim 5, wherein Based on the combined memory vector m (K) , calculate the complete text representation u (K+1) , including: According to the combined memory vector m (K) and the original representation of the word calculate the output of the feedforward neural network in the dual attention network Substitute the said output into the softmax function to obtain the attention weights Based on the attention weights perform a weighted sum on the original representation of the word to obtain the complete representation u of the text (K+1) .
7. The method according to claim 5, wherein The obtaining of the original cross-modal contrast representation by comparing the differences between the text and the image for the attention calculation result of the original representation includes: Original cross-modal contrastive representation wherein represents a trainable weight matrix, || is the absolute value of the element difference, and ; represents a concatenation operation.
8. The method according to claim 1, characterized in that, The calculating of the ironic intention recognition result of the data content to be recognized based on the original cross-modal contrast representation and the context-aware cross-modal contrast representation includes: Connecting the original cross-modal contrast representation and the context-aware cross-modal contrast representation; Inputting the connection result into a fully connected layer and performing binary ironic classification using the Sigmoid function to obtain the ironic intention recognition result of the data content to be recognized.
9. An electronic device, the electronic device comprising a memory and a processor, wherein a computer program is stored in the memory and is loaded and executed by the processor to implement any one of the methods in claims 1-8.
Citation Information
Patent Citations
Attention mechanism-based intention recognition method and device, equipment and storage medium
CN111737458A
Multi-modal Mongolian sentiment analysis method based on irony recognition and fine-grained feature fusion
CN113657115A