A bidirectional adapter multimodal sarcasm detection system, detection method and device
By utilizing a bidirectional adapter multimodal satire detection system and external common sense generation and semantic modulation mechanisms, the problems of unstable information interaction and difficulty in capturing fine-grained satire in multimodal satire detection are solved, achieving efficient and accurate satire semantic recognition.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NANJING UNIV OF POSTS & TELECOMM
- Filing Date
- 2026-05-19
- Publication Date
- 2026-08-04
AI Technical Summary
Existing multimodal satire detection methods lack effective semantic association modeling in the interaction between text and visual information, which makes it difficult to fully express potential satirical cues, fails to fully utilize the semantics of a single information source, is prone to interference when external knowledge is not sufficiently associated with the sample, and is difficult to capture fine-grained satirical features.
A bidirectional adapter multimodal irony detection system is adopted. The common sense generation module generates external common sense that matches the sample, the feature encoding module performs unified encoding, the common sense modeling module constructs semantic associations, the common sense perception bidirectional adaptation fusion module performs iterative updates, and finally irony detection is performed in the modality fusion module.
It enhances the accuracy and stability of multimodal satire recognition, accurately captures the satirical contrast between text and images, improves the detection capability of fine-grained satirical expressions, and enhances the robustness and interpretability of the model.
Smart Images

Figure CN122220502B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multimodal semantic information data processing technology, specifically to a bidirectional adapter multimodal irony detection system, detection method, device, and storage medium. Background Technology
[0002] Irony is a form of linguistic expression that conveys true intentions through semantic or emotional contrast. Its surface semantics are often inconsistent with the speaker's actual intentions. Therefore, the identification of ironic semantics usually requires a comprehensive analysis that combines contextual information and relevant common sense.
[0003] With the development of social media, user-posted information no longer consists solely of text but often includes multiple information formats such as images. In such multimodal expression scenarios, irony is often manifested through inconsistencies in the semantic or emotional relationship between the text and the accompanying image. In this context, relying solely on information from a single modality is insufficient for accurate identification of irony. Therefore, the key to multimodal irony detection lies in modeling the semantic relationships between text and images and their potential inconsistencies.
[0004] In early studies of multimodal satire detection, textual and image information were typically processed separately, making it difficult to characterize potential semantic discrepancies or emotional contrasts between the two. With the gradual emergence of multimodal satire data resources, related methods have begun to analyze the relationship between textual content and visual information through joint modeling strategies, and to combine graph structures or attention mechanisms to characterize the differences between different information sources at both the holistic and local levels, thereby improving satire recognition performance.
[0005] With the development of multimodal irony detection technology, some methods have begun to incorporate external knowledge to enhance the model's ability to understand irony semantics. In recent years, multimodal large language models have been able to generate relevant common sense information in the joint context of text and images, and use it as auxiliary semantic input, thereby helping to reveal the potential contrast between text content and visual information and improving the accuracy of irony recognition.
[0006] However, existing multimodal irony detection methods still have certain shortcomings, mainly in the following aspects: First, some methods focus on a single source of information and fail to effectively characterize the potential semantic discrepancies or expressive differences between textual content and visual information; Secondly, existing methods often use direct feature concatenation or element-level combination for joint modeling, which can easily introduce interference information during the fusion process, thereby weakening the original semantic expressive power. Third, in some scenarios, a single source of information is sufficient to convey satirical meaning, and forcibly introducing other information may affect the judgment result, while the relevant methods do not effectively distinguish the correlation between information sources; Fourth, some satirical expressions rely on fine-grained semantics or local visual cues, while existing methods tend to focus on overall information modeling, making it difficult to capture the aforementioned detailed features; Fifth, some methods introduce external knowledge to enhance semantic expression, but when the introduced information is not sufficiently related to the current samples, it is easy to introduce additional interference, thereby affecting the stability and robustness of the model.
[0007] To address this, this application proposes a bidirectional adapter multimodal irony detection system, method, and device. Targeting text-image scenarios, it constructs an auxiliary common-sense representation adapted to the current semantic and emotional tendency within the context of text and visual content. Based on this, a cross-information-source semantic modulation mechanism is introduced to achieve progressive interaction between different information sources. Ultimately, by maintaining the stability of the semantic representation of each information source during the interaction process, collaborative expression of cross-information semantics can be achieved, thereby enhancing the semantic characterization capability in multimodal irony recognition and solving the aforementioned technical problems. Summary of the Invention
[0008] The main objective of this invention is to provide a bidirectional adapter multimodal irony detection system, detection method, and device to solve the following technical problems mentioned in the background art: the lack of effective semantic association modeling between different information sources makes it difficult to fully express potential irony cues; key information is easily weakened during the interaction of multi-source information; semantics with discriminative value in a single information source are not fully utilized; and interference problems may arise when the introduced external knowledge is not sufficiently associated with the semantics of the sample.
[0009] The present invention solves the above-mentioned technical problems by adopting the following technical solutions: A bidirectional adapter multimodal irony detection system, comprising: The common sense generation module is used to generate and filter external supplementary common sense in the joint context of textual and visual information to obtain external common sense representations that match the input samples in terms of semantic expression and emotional orientation. The feature encoding module is used to learn textual information, visual information, and external common sense representations obtained by the common sense generation module. Through a unified encoding mechanism, information from different sources is mapped to semantic features in the same representation space, and finally a common sense feature matrix is constructed to support subsequent information interaction and association modeling. The common sense modeling module is used to perform structured semantic modeling on the common sense feature matrix output by the feature encoding module, to characterize the contextual relationships between the various components of common sense, and to build a generalized common sense semantic representation on this basis, so as to generate a common sense semantic guidance vector that can guide the interaction of subsequent textual and visual information. The common sense perception bidirectional adaptation and fusion module is used to perform bidirectional semantic modulation and iterative updates on the text feature matrix and the image feature matrix based on the common sense semantic guidance vectors corresponding to two specified modalities. While maintaining the stability of their respective semantic structures, it realizes the collaborative semantic evolution process between different information sources and finally generates the CLS global semantic vectors of the text modality and the image modality. The modality fusion module is used to perform unified representation modeling on the text information, visual information and common sense features processed by the aforementioned common sense perception bidirectional adaptation fusion module. By constructing a joint feature sequence of multi-source information, it obtains an overall representation that includes contextual semantic associations to support subsequent irony discrimination tasks. The irony detection module is used to utilize the cross-modal joint feature matrix output by the modality fusion module. By combining the CLS global semantic vectors of the text and image modalities, irony semantic prediction is performed on the text-image pairs to generate the final irony detection result.
[0010] Preferably, the common sense generation module performs the following operations: Input the text of the sample to be detected and its corresponding image, and process the input text information based on an existing multimodal large language model. With image information Perform joint semantic parsing to generate corresponding multimodal semantic descriptions. ,have: ; Based on the multimodal semantic description, a common sense reasoning model is introduced to generate candidate external common sense based on preset common sense relationship types. The set of is: ,in Indicates the pre-defined common-sense relationship type; Then, a large language model is used to perform semantic evaluation and screening of candidate external common sense, retaining common sense content that matches the current text and image information in terms of semantic expression and emotional orientation. These include: The final output is the generated and filtered reliable external common sense. A set of.
[0011] Preferably, the feature encoding module performs the following operations: The text information is encoded based on a pre-trained encoding model to obtain the corresponding text semantic representation and construct a text feature matrix; The encoding model is used to extract features from image information, obtain visual semantic representation, and construct an image feature matrix; The external common sense representation is vectorized using the same encoding method as the text information, generating a common sense feature matrix, so that text features, image features and common sense features are comparable in a unified representation space.
[0012] Preferably, the common sense modeling module performs the following operations: Self-attention computation is performed on the commonsense feature matrix to construct the contextual dependencies between the tokens that are the components of commonsense, in order to obtain an updated commonsense feature matrix. The self-attention computation is as follows:
[0013] in, The input is the original commonsense feature matrix, which contains the features of all commonsense tokens. This is a learnable weight matrix used to map features to query vectors. This is a learnable weight matrix used to map features to key vectors. This is a learnable weight matrix used to map features to value vectors. The dimension of the feature vector (the length of each token feature). This is a normalization function used to convert attention scores into weights between 0 and 1. This is the updated common sense feature matrix after self-attention weighted fusion, and this common sense feature matrix... Contextual dependencies between tokens have been modeled; Based on the self-attention updated common sense feature representation, the semantic contributions of each common sense component token are weighted and aggregated to extract overall semantic information, resulting in a generalized common sense semantic vector for each sample. For the b-th sample, there exists a generalized common sense semantic vector. Represented as:
[0014]
[0015] in, For the first In the nth sample Each component token has its updated features calculated through self-attention. For learnable projection vectors, Indicates the first The importance weight of each common-sense component token in the current round of semantic aggregation. The total number of constituent tokens (sequence length) contained in each common sense sample; Ultimately, the generalized common sense semantic vector will be generated. By mapping the textual and visual information using trainable projection matrices respectively, common-sense semantic guidance vectors for both modalities are generated, with the following mapping formula:
[0016]
[0017] in, For the first The visual modality commonsense semantic guidance vector corresponding to each sample. For the first The text modality commonsense semantic guidance vector corresponding to each sample. and These are trainable projection matrices specific to the visual modality and the text modality, respectively.
[0018] Preferably, the common sense perception bidirectional adaptation and fusion module performs the following operations: In round 0, the tCLS vector in the text feature matrix and the vCLS vector in the image feature matrix generated by the CLIP encoder are used directly as the initial semantic representations of text information and visual information, respectively. Then, starting from round 1, bidirectional semantic modulation and iterative updates are performed on the text feature matrix and image feature matrix, where: (1) Perform linear transformation and nonlinear mapping on the common sense semantic guidance vector corresponding to the text information to generate semantic modulation components, and apply them to the common sense semantic guidance vector corresponding to the visual information through residual injection, so that the visual information obtains semantic adjustment from the text information in the current round, as follows:
[0019] in, This serves as the image modality commonsense guiding vector updated after the injection of textual semantics. The modal commonsense guide vector for the original image. The modulation coefficients of the visual image. and Two layers of linear weights on the text side are used to perform non-linear transformations on the text-guided vector. This is the original text modality commonsense guide vector. and Two layers of offset for the text side. This is an activation function used to introduce non-linearity and enhance semantic fitting ability. It is a random deactivation function; (2) Using a symmetrical approach, the common sense semantic guidance vector corresponding to the visual information is processed in the same way to generate a visual semantic modulation component, which is then applied to the common sense semantic guidance vector corresponding to the text information in the form of a residual, to obtain the updated representation of the text information in the current round, as follows:
[0020] in, This is the text modality commonsense guide vector updated after image semantics are injected. For text modulation coefficients, and For the image side, there are two layers of linear weights. and Two layers of bias are applied to the image side to perform nonlinear transformations on the image guiding vector; (3) In the current iteration, the common sense adaptation vectors corresponding to the text information and visual information are embedded into their respective original feature matrices, and an extended feature sequence is formed by concatenation. Then, self-attention modeling is performed on the extended feature sequence to realize the context association modeling between the common sense adaptation vector and the original features. After modeling is completed, the updated commonsense adaptation vector is extracted from the corresponding positions, while maintaining the structural form of the original feature matrix. This allows each component token of the commonsense to obtain a semantically enhanced representation across information sources while maintaining structural consistency, resulting in:
[0021]
[0022] in, This is an enhanced image feature matrix that incorporates common-sense semantics after BERT encoding. This represents the self-attention modeling operation of the BERT encoder to achieve token-level semantic interaction. This indicates a feature sequence concatenation operation. This refers to the global CLS vector of the image output by the CLIP encoder. This is represented as a sequence of local patch features in an image output using a CLIP encoder. For image attention mask, This is an enhanced text feature matrix that incorporates common-sense semantics after BERT encoding. This is the text global CLS vector output by the CLIP encoder. This is represented as a text token feature sequence output by the CLIP encoder. For text attention mask; Finally, the common sense adaptation vectors corresponding to the text and visual information obtained in the current round are concatenated, and their semantic information is written back into the original common sense feature matrix based on the cross-attention mechanism to update the common sense representation, thereby constructing a common sense semantic representation that evolves gradually with iteration and providing input for subsequent rounds.
[0023] in, For the updated common sense feature matrix, , and For the cross-attention projection matrix, and These are the image adaptation vector and text adaptation vector output from the last iteration, respectively. The final output includes the text modality and image modality feature matrices after semantic interaction, as well as the common sense adaptation vectors and updated common sense feature matrices for the two modalities in the final round. .
[0024] Preferably, in the process of bidirectional semantic modulation and iterative updating of the text feature matrix and image feature matrix, the common sense perception bidirectional adaptation fusion module also introduces a cross-round gating update mechanism to avoid the gradual decay of historical semantic information during multiple rounds of semantic modulation. This cross-round gating update mechanism is based on the relationship between the update result of the current iteration round and the corresponding representation of the previous round, and adaptively adjusts the combination ratio of the two through a gating function to obtain the common sense adaptation vector of the current round, wherein: In the initial iteration round (first round), since there is no representation of the previous round, that is: in the first iteration, since there is no common sense adaptation vector of the previous round, the text and image [CLS] vectors output by the CLIP encoder in the initial setting (round 0) are used as the initial reference for the calculation; During the bidirectional semantic modulation and iterative update of the text feature matrix and image feature matrix in the cross-wheel gating update mechanism:
[0025]
[0026] in, and For trainable text gating functions and images, it is used to dynamically adjust the information retention and update ratio based on the semantic relationship between the current round and the historical representation, thereby maintaining the stability of the original semantic representation while continuously introducing cross-information semantics. For the first The image commonsense adaptation vector is used as the embedded image adaptation vector. For the first The commonsense text adaptation vector is used as the embedded text adaptation vector. This represents element-wise product.
[0027] Preferably, the modal fusion module performs the following operations: The input image feature matrix is iteratively updated in the commonsense-aware bidirectional adaptation fusion module. Text feature matrix The final updated common sense feature matrix and the corresponding modal attention mask ; The image feature matrix, text feature matrix, and commonsense feature matrix are sequentially combined along the token dimension to form a unified feature sequence representation. Simultaneously, the corresponding attention masks are consistently combined to obtain the overall mask representation corresponding to the feature sequence. ,have:
[0028]
[0029] in, This refers to the multimodal fusion features resulting from the concatenation of images, text, and common knowledge. This is the overall attention mask after fusion, used as a unified mask for the joint encoder; Representing a unified feature sequence and the corresponding attention encoding Input to multi-layer BERT encoder In this process, by modeling the dependencies between different positions through self-attention, feature representations containing contextual information are obtained, such as: ; The final output is a feature matrix that integrates text, images, and external common sense information. This matrix contains semantic information between different modalities and can be directly used for irony detection.
[0030] Preferably, the irony detection module performs the following operations: The cross-modal fusion feature matrix output by the input modal fusion module And the global semantic representation vectors corresponding to the text modality and the image modality (text CLS vector and image CLS vector). .
[0031] Cross-modal fusion feature matrix The input is fed into a fully connected layer and then normalized using the Softmax function to obtain an irony prediction result based on fused features. ,have:
[0032] in, This is the weight matrix of the fully connected layer. Bias for fully connected layers; The global semantic vectors (CLS) extracted from the text modality and the image modality (representing the overall information of the entire text or image) are processed separately. In this processing, each CLS vector is input into the corresponding fully connected layer and then passed through the Softmax function to obtain separate irony prediction results for the text modality and the image modality, as follows:
[0033]
[0034] in, For the text unimodal irony prediction results, The result is a single-modal image irony prediction. Predicting fused features Text modality prediction Image modality prediction According to the preset weights , , The weighted average yields the model's final ironic prediction result: ; During the training phase, supervision is applied simultaneously to the three prediction results, and the model parameters are jointly optimized using the binary classification cross-entropy loss function:
[0035] in, To optimize the total loss during model backpropagation, Indicates branch Belonging to text ,image Integration Three branches, The true label for the corresponding branch is obtained by using a binary search result, i.e. This is not meant to be sarcastic. It is meant to be sarcastic. The corresponding branch prediction probability output by each branch model; The final output represents the final prediction of whether the input text-image pair constitutes ironic semantics. It is used to complete the task of multimodal irony detection.
[0036] On the other hand, this invention also discloses a bidirectional adapter multimodal irony detection method, implemented based on any of the aforementioned bidirectional adapter multimodal irony detection systems. It constructs an overall semantic processing framework around credible external common sense, forming a processing flow that includes external knowledge generation and filtering, modal semantic representation, common sense-guided semantic modulation, bidirectional progressive information fusion, and joint semantic reasoning and irony discrimination. Under the condition that external common sense matches the semantic expression and emotional orientation of textual and visual content, it achieves stable interaction between textual and visual information, thereby facilitating the characterization of irony cues from different information sources and improving the discrimination effect and model stability of multimodal irony recognition.
[0037] Preferably, the bidirectional adapter multimodal irony detection method specifically includes: Step S1. Perform joint semantic parsing on the input text and image, generate candidate external common sense based on the existing multimodal large language model and common sense reasoning model, and obtain credible external common sense through semantic and sentiment consistency screening. Step S2. Use a unified pre-trained encoder to encode the text, image and filtered external common sense respectively to obtain the text feature matrix, image feature matrix and common sense feature matrix in the same feature space; Step S3. Perform self-attention modeling and weighted aggregation on the common sense feature matrix to generate a generalized common sense semantic vector, and project it into a common sense semantic guidance vector that is adapted to the text modality and the image modality; Step S4. Based on the common sense semantic guidance vector, perform bidirectional semantic modulation and cross-round gating update on text features and image features to achieve multi-round progressive cross-modal interaction and maintain the stability of the original semantic structure of each modality, while writing back the interaction semantics to update the common sense features. Step S5. Concatenate the enhanced text features, image features, and updated common sense features along the token dimension, and combine them with the corresponding attention mask in the joint encoder to obtain multimodal fusion features; Step S6. Based on the multimodal fusion features, the text global semantic vector, and the image global semantic vector, perform irony prediction respectively, and weightedly fuse the three prediction results to obtain the final irony detection result. The model training is completed by using multi-branch joint cross-entropy loss.
[0038] In another aspect, the present invention also discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the method described above.
[0039] In another aspect, the present invention also discloses a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the method described above.
[0040] As can be seen from the above technical solution, the present invention provides a bidirectional adapter multimodal satire detection system, detection method, and device. Compared with the prior art, the present invention has the following advantages: 1. This invention constructs a reliable external common sense generation and screening mechanism in the image-text multimodal irony detection process, which can accurately introduce external knowledge that is highly matched with the semantics and emotions of the sample, reduce irrelevant noise interference, thereby improving the reliability of common sense utilization and detection stability, and enhancing the recognition accuracy of irony semantics in complex contexts.
[0041] 2. This invention, by setting a common-sense-guided bidirectional adaptive interaction mechanism between text and image modalities, can alleviate the cross-modal semantic gap, preserve the original key semantics of each modality in the fusion interaction, thereby enabling stable execution of cross-modal semantic collaboration without weakening the effective information of a single modality, and facilitating more accurate capture of ironic contrast clues between text and images.
[0042] 3. By employing a token-level fine-grained semantic interaction method in the feature modeling stage, this invention can fully characterize subtle local semantic differences and uncover implicit irony features, thereby improving the fine-grained discrimination capability of multimodal irony detection and enhancing the detection sensitivity for weak irony and implicit irony expressions.
[0043] 4. By using external common sense in natural language form as an intermediate representation in the model reasoning, this invention can facilitate the explicit characterization of semantic basis and intuitive display of decision-making logic, thereby improving the interpretability and readability of model decisions. This makes the multimodal irony detection process more transparent, easier to verify and implement.
[0044] 5. By setting a multi-path weighted prediction strategy that combines single-modal and multi-modal approaches in the final discrimination stage, this invention can improve the overall prediction robustness of the model while taking into account both independent sarcasm cues in single-modal mode and cross-modal interactive information. As a result, it can maintain efficient, stable and accurate sarcasm detection performance in diverse image and text scenarios.
[0045] It should be understood that the descriptions in this section are not intended to identify key or essential features of embodiments of the invention, nor are they intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Of course, implementing any product of the invention does not necessarily require achieving all of the advantages described above simultaneously. Attached Figure Description
[0046] The accompanying drawings, which form part of this application, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings: Figure 1This is a schematic diagram of the data processing of the multimodal irony detection system of the present invention; Figure 2 This is an overall flowchart of the multimodal irony detection method of the present invention. Detailed Implementation
[0047] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0048] For details in the embodiments, please refer to Figures 1 to 2 .
[0049] Existing technology description 1: In existing technologies, such as the paper "Xie Y, Zhu Z, Chen X, et al. MoBA: Mixture of bi-directional adapter for multi-modal sarcasm detection[C] / / Proceedings of the 32nd ACM International Conference on Multimedia. 2024: 4264-4272", a multimodal information modeling scheme based on a bi-directional adapter is disclosed. This scheme combines the design ideas of low-rank interaction mechanism and hybrid expert model. By introducing a bi-directional adapter module into the pre-trained multimodal model, semantic interaction between text modality and image modality is realized in a plug-in manner, thereby improving the collaborative modeling capability of multimodal features without significantly increasing the scale of model parameters. The bi-directional adapter can be seamlessly integrated into existing multimodal pre-trained models without making significant adjustments to the main structure of the original model, thereby enhancing the cross-modal semantic fusion effect of the model.
[0050] Specifically, the technical solution of this prior art mainly includes the following steps: (a) Feature encoding: Based on the pre-trained multimodal encoding model, feature encoding is performed on text information and image information respectively to obtain the corresponding text feature matrix and image feature matrix; (b) Feature compression: Low-dimensional mapping is performed on the text feature matrix and the image feature matrix to construct a compact feature representation for bidirectional adapter processing; (c) Bidirectional interaction: The low-dimensional text feature matrix and the low-dimensional image feature matrix are input into the bidirectional adapter module. Under the control of the hybrid expert mechanism, semantic information from another modality is introduced to update the current modality feature matrix, thereby realizing bidirectional semantic interaction between the text modality and the image modality. (d) Semantic prediction: Dimensional recovery is performed on the feature matrix updated by the bidirectional adapter, and the feature matrices of the two modalities achieve full semantic interaction through multiple rounds of interaction, and irony prediction is performed based on the final feature matrix.
[0051] In summary, this existing technology demonstrates that multimodal information fusion is not limited to the direct fusion of feature matrices from different modalities. Instead, it can achieve cross-modal semantic interaction within the adapter by introducing bidirectional adapter-like plugins, and then feed the fused semantic information back to each modal feature matrix. While ensuring the effectiveness of multimodal fusion, this scheme effectively controls the model parameter scale by utilizing a low-rank interaction mechanism and a hybrid expert model, enabling it to achieve both computational efficiency and performance in large-scale training and practical application scenarios.
[0052] Although the MOBA method alleviates the interference of cross-modal fusion on the original semantic representation to some extent by introducing bidirectional adapters and low-rank interaction mechanisms, and realizes inter-modal information exchange in the form of plug-ins, it still has the following technical limitations in multimodal irony detection tasks: The main purpose of introducing the low-rank interaction mechanism in this method is to reduce the model parameter size and improve training and inference efficiency. However, low-rank mapping has certain limitations in semantic expression capabilities. Its performance improvement in the irony detection task is mainly reflected in the computational efficiency level, and its effect on improving detection accuracy is limited. This method only performs semantic modeling based on the text and image information of the sample itself, without introducing external common sense or background knowledge as a supplement. In satirical scenarios that rely on implicit common sense understanding, the model's ability to discriminate satirical semantics is limited. This method uses multiple sets of low-rank matrices as adaptation experts for cross-modal interaction. However, due to the common semantic differences between text modality and image modality, relying solely on low-rank mapping is insufficient to fully characterize the complex and nonlinear cross-modal semantic relationships. This method uses overall modal feature fusion as the main modeling approach, but does not explicitly model the emotional cues and fine-grained semantic contrasts that are discriminative in multimodal irony detection, making it difficult to effectively capture subtle semantic changes in ironic expressions. When feeding the fused semantic information back to the original modal feature representation, this method mainly relies on the self-attention mechanism for semantic enhancement. It lacks explicit constraints on the cross-modal semantic absorption process, making it difficult to ensure that the original modal features fully utilize the fused semantics.
[0053] Therefore, although the MOBA method has certain advantages in terms of model parameter size control and computational efficiency, it still has significant shortcomings in improving the accuracy of multimodal irony detection and common sense-assisted semantic understanding.
[0054] Existing technology description 2: Furthermore, existing technologies, such as the paper "Xie Y, Zhu Z, Chen X, et al. MoBA: Mixture of bi-directional adapter for multi-modal sarcasm detection[C] / / Proceedings of the 32nd ACM International Conference on Multimedia. 2024: 4264-4272", disclose a method that utilizes a multimodal large language model to generate external supplementary knowledge and introduces it as a semantic mediator into the multimodal sarcasm detection process. Its core idea is to generate external supplementary semantic information related to the input sample within the text-image joint context, and use this external supplementary knowledge as an intermediary bridge to guide the text modality and image modality to semantically interact with the external knowledge, thereby alleviating the semantic inconsistency problem caused by direct cross-modal fusion and improving the accuracy of multimodal sarcasm recognition.
[0055] The typical process of this technology includes the following steps: (a) Joint understanding: In the joint context of text and images, a multimodal large language model is used to perform overall semantic understanding of the input samples; (b) Knowledge generation: Based on pre-designed prompt templates, guide the multimodal large language model to generate image descriptions, sentiment analysis results and potential sentiment contradiction information related to the samples, forming an external supplementary knowledge set; (c) Sentiment assessment: For the generated external supplementary knowledge, sentiment analysis tools for social media and short text scenarios are used to assess the sentiment intensity of each knowledge fragment; (d) Knowledge screening: Based on the results of the emotional intensity assessment, external supplementary knowledge is retained or trimmed, and content with low emotional contribution is reduced; (e) Modal interaction: Cross-attention calculation is performed between the original feature matrices of the text modality and the image modality and the external supplementary knowledge feature matrix, respectively; (f) Feature update: Based on the cross-attention results, the original features of the text modality and the image modality are updated; (g) Feature concatenation: The updated modal feature matrix is concatenated with the external supplementary knowledge feature matrix to form an extended feature representation; (h) Association reinforcement: After feature concatenation, cross-attention calculation is performed again between the modality feature matrix and the feature matrix of external supplementary knowledge; (i) Weight modeling: A gating mechanism is introduced to dynamically calculate the weights of the text modal feature matrix and the image modal feature matrix during the fusion process; (j) Fusion discrimination: Based on the gating weights, a weighted sum is performed to obtain a unified fusion feature representation, and the irony detection task is completed accordingly.
[0056] Through the above technical solution, this method introduces external supplementary knowledge generated by a multimodal large language model as a semantic mediator between text modality and image modality. In the fusion process, it combines cross-attention mechanism and gating fusion strategy, which alleviates the semantic conflict problem caused by direct cross-modal feature fusion to a certain extent, thereby improving the overall performance of multimodal irony detection.
[0057] Although the EilMoB method alleviates the multimodal semantic inconsistency problem to some extent by introducing external supplementary knowledge as a semantic bridge between the text modality and the image modality, its overall technical solution still has the following shortcomings: This method mainly relies on a single prompt template in the external supplementary knowledge generation stage. Due to the limitations of the multimodal large language model's ability to understand complex emotions and cross-modal semantics, the generated supplementary knowledge may have biases in emotion judgment and semantic expression, affecting the reliability of the knowledge.
[0058] This method uses emotional intensity as the criterion for selecting supplementary knowledge, but emotional intensity is difficult to characterize the consistency between supplementary knowledge and the original text-image pair in terms of emotional orientation, and may still introduce semantically mismatched noise information.
[0059] In the modality fusion process, directly performing multi-round cross-modal interactions based on the feature matrices of the original text and images may lead to the gradual weakening of key discriminative information in the original modality.
[0060] This method focuses on cross-modal association modeling, but pays insufficient attention to semantic inconsistencies within text and image modalities and the inherent semantic ambiguity of modalities, thus limiting its ability to characterize fine-grained satirical cues.
[0061] Therefore, the EilMoB method still has room for improvement in terms of supplementary knowledge generation and filtering accuracy, noise control capabilities, and preservation of original modal semantics.
[0062] As can be seen from the existing technologies described above, current multimodal irony detection methods typically combine textual and visual information for analysis, aiming to uncover semantic contrasts between different information sources as irony clues. However, due to the differences in expression and focus between text and images, these methods face certain difficulties in uniformly characterizing the semantic changes of both, and are prone to overlooking local or implicit irony features during information interaction, thus affecting the overall detection performance.
[0063] On the one hand, existing methods tend to focus on the semantic relationship between textual and visual information at the overall level, while paying insufficient attention to the ironic cues that exist independently in a single information source and the fine-grained cross-information differences, making it difficult to effectively mine some implicit ironic semantics. On the other hand, some methods adopt direct summarization or averaging when integrating multi-source information. Although this is beneficial for overall semantic modeling, it can easily mask the fine-grained semantic features in the text and the local entity information in the image, thereby reducing the ability to identify local or weak ironic cues.
[0064] Furthermore, some multimodal irony detection methods enhance semantic expressiveness by introducing external information. However, they lack effective constraint mechanisms in the generation and utilization of this external information, failing to fully consider its matching degree with the original text and visual content in terms of semantic expression and emotional orientation. Consequently, the introduced information may contain content with low relevance to the current sample, thus introducing interference into the modeling process and affecting the accuracy and stability of irony recognition.
[0065] Therefore, such as Figure 1 As shown in the embodiments of the present invention, the bidirectional adapter multimodal irony detection system constructs an overall semantic processing framework around credible external common sense, forming a processing flow that includes external knowledge generation and filtering, modal semantic representation, common sense-guided semantic modulation, bidirectional progressive information fusion, and joint semantic reasoning and irony discrimination. Under the condition that external common sense matches the semantic expression and emotional orientation of textual and visual content, stable interaction between textual and visual information is achieved, which is beneficial for characterizing ironic cues between different information sources and improving the discrimination effect and model stability of multimodal irony recognition. Specifically, it includes the following system modules: (1) Common Sense Generation Module: Function Description: The common sense generation module is used to generate and filter external supplementary common sense in the joint context of textual and visual information, so as to obtain external common sense representations that match the input samples in terms of semantic expression and emotional orientation. Input: The text of the sample to be detected and its corresponding image; Step 1: Perform joint semantic parsing of text and image information based on a multimodal large language model to generate corresponding multimodal semantic descriptions. ,have: ; Step 2: Based on the multimodal semantic description, a common sense reasoning model is introduced to generate candidate external common sense based on the preset common sense relationship types. The set of is: ,in Indicates the pre-defined common-sense relationship type; Step 3: Utilize a large language model to perform semantic evaluation and screening of candidate external common sense, retaining common sense content that matches the current text and image information in terms of semantic expression and emotional orientation. Examples include: ; Output: A set of credible external common sense generated and filtered, which is used as input for the subsequent feature encoding module.
[0066] By constructing a mechanism for generating and filtering external common sense, the acquired common sense can be matched with the original text and visual content in terms of semantic expression and emotional orientation. This allows for the precise introduction of external knowledge that is highly matched with the semantics and emotions of the sample, reducing the introduction of irrelevant information, minimizing irrelevant noise interference, improving the reliability of common sense utilization and detection stability, and enhancing the accuracy of identifying ironic semantics in complex contexts.
[0067] (2) Feature encoding module: Function Description: The feature encoding module is used to learn representations of textual information, visual information, and filtered external common sense. Through a unified encoding mechanism, it maps information from different sources into semantic features in the same representation space to support subsequent information interaction and association modeling. Inputs include text information, image information, and externally reliable common-sense information; Step 1: Encode the text information based on the pre-trained encoding model to obtain the corresponding text semantic representation and construct the text feature matrix; Step 2: Use the coding model to extract features from the image information, obtain visual semantic representation, and construct the image feature matrix; Step 3: Vectorize external common sense using the same encoding method as the text information to generate a common sense feature matrix, so that text features, image features and common sense features are comparable in a unified representation space; Output: text feature matrix, image feature matrix, and common sense feature matrix.
[0068] At this point, by acquiring feature representations of textual and visual information based on pre-trained models and optimizing these features using appropriate processing strategies, the semantic content from different information sources can be fully preserved in the feature representations, thus supporting subsequent information interaction and association modeling. (3) Common Sense Modeling Module: Function Description: The common sense modeling module is used to perform structured semantic modeling on the common sense feature matrix output by the feature encoding module, to characterize the contextual relationships between the various components of common sense, and on this basis, to construct a generalized common sense semantic representation to generate semantic guidance vectors that can guide the interaction of subsequent textual and visual information. Input: Common sense feature matrix obtained from the feature encoding module; Step 1: Perform self-attention computation on the common sense feature matrix to model the contextual dependencies between common sense tokens, resulting in the updated common sense feature matrix:
[0069] in, The input is the original commonsense feature matrix, which contains the features of all commonsense tokens. This is a learnable weight matrix used to map features to query vectors. This is a learnable weight matrix used to map features to key vectors. This is a learnable weight matrix used to map features to value vectors. The dimension of the feature vector (the length of each token feature). This is a normalization function used to convert attention scores into weights between 0 and 1. This is the updated common sense feature matrix after self-attention weighted fusion, and this common sense feature matrix... Contextual dependencies between tokens have been modeled; Step 2: Based on the common sense feature representation updated by self-attention, the semantic contributions of each token are weighted and aggregated to extract the overall semantic information, resulting in a generalized common sense semantic vector for each sample. Assume that the b-th sample has a generalized common sense semantic vector. Represented as:
[0070]
[0071] in, For the first In the nth sample Each component token has its updated features calculated through self-attention. For learnable projection vectors, Indicates the first The importance weight of each common-sense component token in the current round of semantic aggregation. The total number of constituent tokens (sequence length) contained in each common sense sample; Step 3: Map the generalized common sense semantic vectors to the trainable projection matrices corresponding to the textual and visual information respectively to generate common sense semantic guidance vectors for two modalities:
[0072]
[0073] in, For the first The visual modality commonsense semantic guidance vector corresponding to each sample. For the first The text modality commonsense semantic guidance vector corresponding to each sample. and These are trainable projection matrices specific to the visual modality and the text modality, respectively. Output: Commonsense semantic guidance vectors for the text modality and the image modality, providing a bridge for subsequent semantic interaction between the two modalities.
[0074] (4) Common sense perception bidirectional adaptation and fusion module: Function Description: The common sense perception bidirectional adaptation and fusion module performs bidirectional semantic modulation and iterative updates on the text feature matrix and image feature matrix based on the common sense semantic guidance vectors corresponding to the two modalities. While maintaining the stability of their respective semantic structures, it realizes the collaborative semantic evolution process between different information sources. Inputs: text feature matrix, image feature matrix, and common sense semantic guidance vector and common sense feature matrix output by the common sense modeling module and mapped to the corresponding modal space; Step 1: In round 0, the tCLS vector in the text feature matrix and the vCLS vector in the image feature matrix generated by the CLIP encoder are directly used as the initial semantic representations of textual and visual information, respectively. This round does not introduce common sense semantic guidance or cross-information interaction, but only uses them as initial reference vectors in subsequent iterations. Step two: Starting from round 1, perform linear transformation and nonlinear mapping on the common sense semantic guidance vector corresponding to the text information to generate semantic modulation components. These components are then applied to the common sense semantic guidance vector corresponding to the visual information through residual injection, enabling the visual information to receive semantic adjustment from the text information in the current round.
[0075] in, This serves as the image modality commonsense guiding vector updated after the injection of textual semantics. The modal commonsense guide vector for the original image. The modulation coefficients of the visual image. and Two layers of linear weights on the text side are used to perform non-linear transformations on the text-guided vector. This is the original text modality commonsense guide vector. and Two layers of offset for the text side. This is an activation function used to introduce non-linearity and enhance semantic fitting ability. It is a random deactivation function; Step 3: Using a symmetrical approach, the common sense semantic guidance vector corresponding to the visual information is processed in the same way to generate a visual semantic modulation component. This component is then applied as a residual to the common sense semantic guidance vector corresponding to the text information to obtain the updated representation of the text information in the current round.
[0076] in, This is the text modality commonsense guide vector updated after image semantics are injected. For text modulation coefficients, and For the image side, there are two layers of linear weights. and Two layers of bias are applied to the image side to perform nonlinear transformations on the image guiding vector; Step four: To avoid the gradual decay of historical semantic information during multi-round semantic modulation, a cross-round gating update mechanism is introduced. This mechanism is based on the relationship between the current round update result and the corresponding representation of the previous round, and adaptively adjusts the combination ratio of the two through a gating function to obtain the common sense adaptation vector of the current round. In the first round of iteration, since there is no representation of the previous round, the text and image [CLS] vectors output by the CLIP encoder in round 0 are used as initial references for calculation.
[0077]
[0078] in, and For trainable text gating functions and images, it is used to dynamically adjust the information retention and update ratio based on the semantic relationship between the current round and the historical representation, thereby maintaining the stability of the original semantic representation while continuously introducing cross-information semantics. For the first The image commonsense adaptation vector is used as the embedded image adaptation vector. For the first The commonsense text adaptation vector is used as the embedded text adaptation vector. Represents element-wise product; Step 5: In the current iteration, the common sense adaptation vectors corresponding to the textual and visual information are embedded into their respective original feature matrices, and then concatenated to form an extended feature sequence. Subsequently, self-attention modeling is performed on the extended feature sequence to model the contextual association between the common sense adaptation vectors and the original features. After modeling is completed, the updated common sense adaptation vectors are extracted from the corresponding positions, while maintaining the structural form of the original feature matrix. This allows each token to obtain a semantically enhanced representation across information sources while maintaining structural consistency.
[0079]
[0080] in, This is an enhanced image feature matrix that incorporates common-sense semantics after BERT encoding. This represents the self-attention modeling operation of the BERT encoder to achieve token-level semantic interaction. This indicates a feature sequence concatenation operation. This refers to the global CLS vector of the image output by the CLIP encoder. This is represented as a sequence of local patch features in an image output using a CLIP encoder. For image attention mask, This is an enhanced text feature matrix that incorporates common-sense semantics after BERT encoding. This is the text global CLS vector output by the CLIP encoder. This is represented as a text token feature sequence output by the CLIP encoder. For text attention mask; Step six involves concatenating the common sense adaptation vectors corresponding to the text and visual information obtained in the current round, and then writing their semantic information back into the original common sense feature matrix based on a cross-attention mechanism to update the common sense representation. This process constructs a common sense semantic representation that evolves progressively with each iteration and provides input for subsequent rounds.
[0081] in, For the updated common sense feature matrix, , and For the cross-attention projection matrix, and These are the image adaptation vector and text adaptation vector output from the last iteration, respectively. Output: The text modality and image modality feature matrices after semantic interaction in this module, as well as the common sense adaptation vectors of the two modalities and the updated common sense feature matrices of the final round output. .
[0082] (5) Modal fusion module: Function Description: The modality fusion module is used to perform unified representation modeling on the text information, visual information and common sense features processed by the common sense perception bidirectional adaptation fusion module. By constructing a joint feature sequence of multi-source information, it obtains an overall representation that includes contextual semantic associations to support subsequent irony discrimination tasks. Input: The image feature matrix obtained through iterative updates in the commonsense-aware bidirectional adaptation fusion module. Text feature matrix The final updated common sense feature matrix and the corresponding modal attention mask ; Step 1: Combine the image feature matrix, text feature matrix, and commonsense feature matrix sequentially along the token dimension to form a unified feature sequence representation. Simultaneously, perform a consistent combination of the corresponding attention masks to obtain the overall mask representation corresponding to the feature sequence.
[0083]
[0084] in, This refers to the multimodal fusion features resulting from the concatenation of images, text, and common knowledge. This is the overall attention mask after fusion, used as a unified mask for the joint encoder; Step 2, feature sequence and the corresponding attention encoding Input to multi-layer BERT encoder In this process, by modeling the dependencies between different positions through self-attention, feature representations containing contextual information are obtained, such as: ; Output: A feature matrix that integrates text, images, and external common sense information. This matrix contains semantic information between different modalities and can be directly used for irony detection.
[0085] (6) Irony detection module: Function Description: The irony detection module utilizes the cross-modal joint feature matrix output by the modality fusion module. By combining the CLS global semantic vectors of the text modality and the image modality, irony semantics prediction is performed on the text-image pair to generate the final irony detection result; Input: Cross-modal fusion feature matrix output by the modal fusion module And the global semantic representation vectors corresponding to the text modality and the image modality (text CLS vector and image CLS vector). ; Step 1: Fuse the cross-modal feature matrix The input is fed into a fully connected layer and then through a softmax function to obtain an irony prediction result based on fused features. :
[0086] in, This is the weight matrix of the fully connected layer. Bias for fully connected layers; Step two involves processing the CLS global semantic vectors (representing the overall information of the entire text or image) extracted from the text and image modalities respectively. Specifically, each CLS vector is input into the corresponding fully connected layer and then passed through the Softmax function to obtain separate semantic prediction results for the text and image modalities.
[0087]
[0088] in, For the text unimodal irony prediction results, The result is a single-modal image irony prediction. Step 3: Combine the three prediction results—feature prediction. Text modality prediction Image modality prediction —According to the pre-set , , The weighted average yields the model's final ironic prediction result: ; Step four: During the training phase, supervision is applied simultaneously to the three prediction results, and the model parameters are jointly optimized using the binary classification cross-entropy loss function.
[0089] in, To optimize the total loss during model backpropagation, Indicates branch Belonging to text ,image Integration Three branches, The true label for the corresponding branch is obtained by using a binary search result, i.e. This is not meant to be sarcastic. It is meant to be sarcastic. The corresponding branch prediction probability output by each branch model; Output: The output is the final prediction of whether the input text-image pair constitutes ironic semantics. It is used to complete the task of multimodal irony detection.
[0090] In summary, this method constructs a stable multi-source information collaboration mechanism to achieve effective interaction between textual and visual information under the guidance of external common sense, while maintaining the integrity of their original semantic features and avoiding semantic weakening or information confusion during multiple rounds of interaction. In addition, it can enhance semantic characterization capabilities by leveraging external common sense when there are complex semantic relationships or low correlation between textual and visual information, and generate feature representations with consistent semantic expression. Combined with the stable information collaboration mechanism, it achieves effective interaction between different information sources, thereby improving the discriminative performance of multimodal irony recognition and the stability of model operation.
[0091] On the other hand, such as Figure 2 As shown, this invention also discloses a bidirectional adapter multimodal irony detection method, implemented based on the bidirectional adapter multimodal irony detection system in the above embodiments. It constructs an overall semantic processing framework around credible external common sense, forming a processing flow including external knowledge generation and filtering, modal semantic representation, common sense-guided semantic modulation, bidirectional progressive information fusion, and joint semantic reasoning and irony discrimination. Under the condition that external common sense matches the semantic expression and emotional orientation of textual and visual content, it achieves stable interaction between textual and visual information, thereby facilitating the characterization of irony cues from different information sources and improving the discrimination effect and model stability of multimodal irony recognition. Specifically, it includes the following steps: Step S1. Perform joint semantic parsing on the input text and image, generate candidate external common sense based on the existing multimodal large language model and common sense reasoning model, and obtain credible external common sense through semantic and sentiment consistency screening. Step S2. Use a unified pre-trained encoder to encode the text, image and filtered external common sense respectively to obtain the text feature matrix, image feature matrix and common sense feature matrix in the same feature space; Step S3. Perform self-attention modeling and weighted aggregation on the common sense feature matrix to generate a generalized common sense semantic vector, and project it into a common sense semantic guidance vector that is adapted to the text modality and the image modality; Step S4. Based on the common sense semantic guidance vector, perform bidirectional semantic modulation and cross-round gating update on text features and image features to achieve multi-round progressive cross-modal interaction and maintain the stability of the original semantic structure of each modality, while writing back the interaction semantics to update the common sense features. Step S5. Concatenate the enhanced text features, image features, and updated common sense features along the token dimension, and combine them with the corresponding attention mask in the joint encoder to obtain multimodal fusion features; Step S6. Based on the multimodal fusion features, the text global semantic vector, and the image global semantic vector, perform irony prediction respectively, and weightedly fuse the three prediction results to obtain the final irony detection result. The model training is completed by using multi-branch joint cross-entropy loss.
[0092] In a further embodiment, the irony detection model is constructed based on the above system according to this application. The performance of this model and different existing models on MMSD is tested, and the experimental results are shown in Table 1 and Table 2 below. Table 1: Performance results of each model on MMSD
[0093] Table 2: Performance results of each model on MMSD2.0
[0094] It is evident that the multimodal satire system model constructed in this application has significantly better accuracy than existing conventional models.
[0095] In a further embodiment, a set of actual multimodal irony detection cases are used to describe in detail the execution flow of the present invention, so as to demonstrate how the invention, in a real-world scenario, sequentially achieves a complete multimodal irony detection process for text-image pairs through modules such as external common sense generation, feature encoding, common sense modeling, common sense perception bidirectional adaptation and fusion, modal fusion, and irony discrimination: Step 1: Processing of the Common Sense Generation Module This step generates believable external common sense based on the input text-image pair, and its processing includes: (1) Use a multimodal large language model to analyze the input image and generate the corresponding image semantic description, such as: "There are three people in the picture, and the scene is similar to a movie or TV series clip....".
[0096] (2) Input the image semantic description and the text “Oh, great. Now no one will know we are here.” into the common sense knowledge base COMET to generate candidate external common sense, such as: “They are trying to stay hidden and not be noticed.”
[0097] (3) Combining the generated candidate common sense, image description and original text, call the large language model to form a reasoning question: "Why, are they really trying to hide?", and generate the answer to the question: "They are clearly in an open and conspicuous environment, yet they claim to be 'invisible', which reflects the absurdity of trying to hide." Finally, the output is external common sense that matches the text-image context, serving as a reliable external common sense for subsequent multimodal irony detection.
[0098] By constructing a reliable external common sense generation and screening mechanism in the image-text multimodal satire detection process, we can accurately introduce external knowledge that highly matches the semantics and emotions of the samples, reduce irrelevant noise interference, thereby improving the reliability of common sense utilization and detection stability, and enhancing the recognition accuracy of satirical semantics in complex contexts. At the same time, by using external common sense in natural language form as an intermediate representation in model reasoning, we can more easily describe the semantic basis and intuitively display the decision-making logic, thereby improving the interpretability and readability of model decisions. This makes the multimodal satire detection process more transparent, easier to verify, and more feasible to implement.
[0099] Step 2: Feature encoding module extracts feature matrix This step encodes the input text, images, and external common knowledge separately to obtain the text feature matrix. Common sense feature matrix Image feature matrix The specific process is as follows: (1) Text feature encoding: The text encoder in the pre-trained CLIP model is used to encode the input text and generate a text feature matrix. Each text token corresponds to a feature representation, for example:
[0100] (2) External common sense feature encoding: External common sense is encoded using the same CLIP text encoder as the text to obtain the common sense feature matrix. :
[0101] (3) Image feature encoding: The image encoder in the CLIP model is used to extract features from the input image to obtain an image feature matrix composed of multiple visual patches. To characterize the visual semantic information in the image:
[0102] Through the aforementioned feature encoding, text, images, and external common sense are mapped to a unified feature space, providing input for subsequent cross-modal semantic interaction and irony feature modeling.
[0103] Step 3: The common sense modeling module enhances common sense feature information. The common sense modeling module takes an external common sense feature matrix as input and, through three steps, generates two sets of common sense semantic guidance vectors that act on the text modality and the image modality respectively, for subsequent cross-modal semantic interaction: (1) Self-attention reinforces common sense characteristics The commonsense feature matrix is updated by modeling the dependencies between commonsense tokens using a self-attention mechanism. For ease of explanation, assume the projection matrix is the identity matrix, then:
[0104] Suppose the common sense feature matrix after self-attention update is as follows:
[0105] (2) Weighted aggregation generates common sense semantic representation The common sense feature matrix updated by self-attention First, assign an initial importance coefficient to each token. This represents its contribution to overall common sense. It is then normalized into attention weights. Finally, these weights are used to sum the weights of each token in the common sense feature matrix to obtain a generalized common sense semantic vector.
[0106] Assume the importance coefficient assigned to each token is:
[0107] We calculate a weighted sum of the token features after the self-attention update to obtain the importance score for each token:
[0108] Next, these scores are normalized into attention weights:
[0109] Finally, the features of all tokens are weighted and aggregated according to normalized weights to obtain the generalized common-sense semantic vector for the current round:
[0110]
[0111] (3) Generate common sense semantic guidance vectors for text and image modalities The converged commonsense semantic vectors are mapped to the feature spaces of the text modality and the image modality using two trainable projection matrices. Let the two projection matrices be:
[0112]
[0113] After mapping, we obtain commonsense semantic guidance vectors for the text modality and the image modality:
[0114]
[0115] Through the above steps, common-sense semantic guidance vectors for text and image modalities are generated, serving as a bridge for subsequent cross-modal semantic interactions.
[0116] By employing a token-level fine-grained semantic interaction approach in the feature modeling stage, we can fully characterize subtle local semantic differences and uncover implicit irony features, thereby improving the fine-grained discrimination capability of multimodal irony detection and enhancing the detection sensitivity for weak irony and implicit irony expressions.
[0117] Step 4: Modal fusion process of the common sense perception bidirectional adaptation fusion module The common sense perception bidirectional adaptation fusion module operates through multiple rounds of iteration. In each round, it generates a common sense adaptation vector for the corresponding modality and feeds back the semantic information in the adaptation vector to the feature matrix (including the common sense feature matrix) of the corresponding modality at the end of each round.
[0118] (1) Initialize the common sense adaptation vectors of the two modes Before the iteration begins, CLS vectors are extracted from the original feature matrices of the text and image modalities to serve as the common sense adaptation vectors for round 0:
[0119]
[0120] (2) Inject the semantics in the common sense semantic guidance vector of each modality into the common sense semantic guidance vector of another modality.
[0121] The following numerical example illustrates how, in the first and subsequent rounds, the common-sense semantic guidance vectors of the text modality and the image modality are mutually injected with semantics through the symmetric bidirectional adaptation module, achieving semantic fusion between modalities. The calculation process for the image modality is as follows:
[0122]
[0123]
[0124] The calculation process for text is similar:
[0125]
[0126]
[0127] (3) Cross-wheel gating fusion To ensure that the results of cross-modal semantic fusion in the previous round are not lost and to prevent the loss of the original modal feature information injected in round 0, a cross-round gating mechanism is introduced. The common sense semantic guidance vector after mutual semantic injection in the current round is weighted and fused with the common sense adaptation vector of the previous round to generate the final common sense adaptation vector of each modality in the current round. The calculation process of the common sense adaptation vector of the image modality is as follows:
[0128]
[0129]
[0130] The calculation process for the commonsense adaptation vector of the text modality is as follows:
[0131]
[0132]
[0133] (4) Update the common sense feature matrix To avoid relying solely on static common-sense semantics in multiple iterations, the semantics of the common-sense adaptation vectors generated in the current round for both modalities are input back into the common-sense feature matrix. First, the text and image bidirectional adaptation vectors are concatenated to form a joint context representation:
[0134] Then, cross-attention calculation is performed to update the commonsense feature matrix:
[0135]
[0136] (5) Back-injection semantics achieves feature fusion To feed back the semantic information of the text and image commonsense adaptation vectors generated in this round to the original modal feature matrix, the commonsense adaptation vectors are concatenated back to the corresponding modal features, and token-level semantic interaction and fusion are achieved through a self-attention mechanism.
[0137] First, the commonsense adaptation vectors of the image and text are concatenated with the original feature matrix:
[0138]
[0139] Then, self-attention calculation is performed on the concatenated matrix:
[0140]
[0141]
[0142]
[0143] Finally, to maintain the shape of the original feature matrix while preserving high-level semantic enhancement information, the concatenated commonsense adaptation vectors are removed:
[0144]
[0145] Meanwhile, the vector corresponding to the removed position will be used as the common sense adaptation vector of the previous round in the next round of iteration calculation, and used for cross-round gating fusion to ensure the continuity and stability of semantic information in the multi-round adaptation process.
[0146] (6) Multiple rounds of iteration Returning to step (2) above, this step can be repeated several times to allow the high-level semantics of the text and image modalities to fully interact under the guidance of common sense, ultimately outputting an enhanced feature matrix for the modality fusion module. , and the updated common sense feature matrix .
[0147] At this point, by setting up a common-sense-guided bidirectional adaptive interaction mechanism between text and image modalities, the cross-modal semantic gap can be alleviated. The original key semantics of each modality are preserved in the fusion interaction, so that cross-modal semantic collaboration can be executed stably without weakening the effective information of a single modality, and it is easier to capture the ironic contrast clues between text and images more accurately.
[0148] Step 5: Modal fusion and irony detection: The modal fusion and irony detection module integrates the processed text, image, and common sense information to form a cross-modal joint semantic representation. It then combines this with the original modal information of the text and image to determine whether the text-image combination is ironic. (1) Multimodal feature splicing The final output image modal feature matrix from step four Text modality feature matrix and the updated common sense feature matrix By concatenating elements along the token dimension, a multimodal fusion input is constructed, incorporating text, images, and common-sense semantics.
[0149] (2) Cross-modal self-attention modeling The concatenated multimodal feature matrix is input into the multi-layer BERT joint encoder, and token-level semantic communication and fusion are achieved through the self-attention mechanism:
[0150]
[0151] Example output matrix is as follows:
[0152] (3) Ironic prediction of fusion features For the fusion feature matrix By performing linear mapping and softmax normalization, we obtain the ironic prediction results of the multimodal fusion features:
[0153]
[0154] (4) Single-mode prediction To preserve the discriminative information of the text and image modalities themselves, linear mapping and softmax normalization are performed on the text CLS vector and the image CLS vector, respectively:
[0155]
[0156] The image prediction is as follows:
[0157]
[0158] (5) Weighted fusion yields the final prediction result. Multimodal fusion prediction and single-modal prediction are weighted according to a preset ratio. Weighted fusion yields the final output of the model:
[0159] The example weighted calculation results are as follows:
[0160]
[0161] The results show that although the image modality semantics tends to be non-ironyous, it is ultimately determined to be ironic under the guidance of the text modality and joint semantics.
[0162] In summary, this method, by setting a multi-path weighted prediction strategy that combines single-modal and multi-modal approaches in the final discrimination stage, can improve the overall prediction robustness of the model while taking into account both independent sarcasm cues in single-modal mode and cross-modal interaction information. As a result, it can maintain efficient, stable and accurate sarcasm detection performance in diverse image and text scenarios.
[0163] In another aspect, the present invention also discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the method described above.
[0164] In another aspect, the present invention also discloses a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the method described above.
[0165] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the bidirectional adapter multimodal satire detection methods described above.
[0166] It is understood that the system provided in the embodiments of the present invention corresponds to the method provided in the embodiments of the present invention, and the explanation, examples and beneficial effects of the relevant content can be referred to the corresponding parts of the above methods.
[0167] This application also provides an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, communication interface, and memory communicate with each other via the communication bus. Memory, used to store computer programs; When the processor executes the program stored in memory, it implements the above-mentioned bidirectional adapter multimodal irony detection method.
[0168] The communication bus mentioned in the above-mentioned electronic devices can be a standard bus for interconnecting peripheral components or an extended industrial standard structure bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc.
[0169] The communication interface is used for communication between the aforementioned electronic devices and other devices.
[0170] The memory may include random access memory or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0171] The processors mentioned above can be general-purpose processors, including central processing units, network processors, etc.; they can also be digital signal processors, application-specific integrated circuits, field-programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0172] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, an optical medium, or a semiconductor medium, etc.
[0173] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
[0174] Furthermore, it should be noted that if any directional indication (such as up, down, left, right, front, back, etc.) is involved in the embodiments of the present invention, the directional indication is only used to explain the relative positional relationship and movement of each component in a specific posture. If the specific posture changes, the directional indication will also change accordingly.
[0175] Furthermore, those skilled in the art should understand that in the actual use of the embodiments of this application, there may be preset thresholds used as the basis for judging the corresponding technical solutions. These thresholds are conventional technical means commonly used in the field to implement functions such as state judgment, condition recognition, and control logic switching. The specific values, setting basis, value selection methods, determination methods, and adjustment rules of the thresholds involved in this technical solution are all conventional technical choices that can be reasonably determined by those skilled in the art based on conventional technical factors such as actual application scenarios, system working states, characteristics of the detection object, hardware performance parameters, and functional requirements, through conventional experiments, calibrations, and debugging. The specific setting and adjustment of the aforementioned thresholds will not cause this technical solution to be unimplementable as a whole, nor will it affect the realization of the core concept and the achievement of the technical effects of this technical solution.
[0176] Furthermore, if the embodiments of this invention involve descriptions such as "first" or "second," these descriptions are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first" or "second" may explicitly or implicitly include at least one of those features. Additionally, the meaning of "and / or" throughout the text includes three parallel solutions; for example, "A and / or B" includes solution A, solution B, or a solution where both A and B are satisfied simultaneously. Furthermore, in the embodiments of this invention, "multiple" refers to two or more. Moreover, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by this invention.
Claims
1. A bidirectional adapter multimodal irony detection system, characterized in that, include: The common sense generation module is used to generate and filter external supplementary common sense in the joint context of textual and visual information to obtain external common sense representations that match the input samples in terms of semantic expression and emotional orientation. The feature encoding module is used to learn textual information, visual information, and external common sense representations. Through a unified encoding mechanism, it maps information from different sources into semantic features in the same representation space, and finally constructs a common sense feature matrix. The common sense modeling module is used to perform structured semantic modeling on the common sense feature matrix, characterize the contextual relationships of each component of common sense, and construct a generalized common sense semantic representation to generate a common sense semantic guidance vector that can guide the interaction of subsequent textual and visual information. The common sense perception bidirectional adaptation and fusion module is used to perform bidirectional semantic modulation and iterative update on the text feature matrix and image feature matrix in the common sense feature matrix based on the common sense semantic guidance vector corresponding to the specified modality, and generate CLS global semantic vectors for text modality and image modality. The modality fusion module is used to perform unified representation modeling on the processed text information, visual information and common sense features. By constructing a joint feature sequence of multi-source information, it obtains an overall representation that includes contextual semantic associations. The irony detection module, by combining the joint feature matrix with the CLS global semantic vectors of the text and image modalities, performs irony semantic prediction on text-image pairs and generates the final irony detection result. The commonsense-aware bidirectional adaptation and fusion module uses a CLIP encoder to generate tCLS vectors in the text feature matrix and vCLS vectors in the image feature matrix, which serve as the initial semantic representations of textual and visual information, respectively. The bidirectional semantic modulation and iterative update steps include: (1) Perform linear transformation and nonlinear mapping on the common sense semantic guidance vector corresponding to the text information to generate semantic modulation components, and apply them to the common sense semantic guidance vector corresponding to the visual information through residual injection, so that the visual information can obtain semantic adjustment from the text information in the current round. (2) Using a symmetrical approach, the common sense semantic guidance vector corresponding to the visual information is processed in the same way to generate a visual semantic modulation component, which is then applied to the common sense semantic guidance vector corresponding to the text information in the form of a residual, so as to obtain the updated representation of the text information in the current round. (3) In the current iteration, the common sense adaptation vectors corresponding to the text information and visual information are embedded into their respective original feature matrices, and an extended feature sequence is formed by concatenation operation. Then, self-attention modeling is performed on the extended feature sequence. After completing the modeling, the updated common sense adaptation vector is extracted from the corresponding position, while keeping the structure of the original feature matrix unchanged. (4) Finally, the common sense adaptation vectors corresponding to the text information and visual information obtained in the current round are concatenated, and their semantic information is written back to the original common sense feature matrix based on the cross attention mechanism to update the common sense representation; In the modality fusion module, the joint feature sequence is defined as a matrix obtained by sequentially combining and splicing the updated image feature matrix, text feature matrix, and common sense feature matrix along the token dimension.
2. The bidirectional adapter multimodal irony detection system as described in claim 1, characterized in that, The common sense generation module performs the following operations: Input the text and corresponding image of the sample to be detected, and perform joint semantic parsing on the input text and image information based on the existing multimodal large language model to generate the corresponding multimodal semantic description; Based on the multimodal semantic description, a common sense reasoning model is introduced to generate a candidate set of external common sense based on preset common sense relationship types; Then, a large language model is used to perform semantic evaluation and screening of candidate external common sense, retaining common sense content that matches the current text information and image information in terms of semantic expression and emotional orientation, and finally outputting the generated and screened set of credible external common sense.
3. The bidirectional adapter multimodal irony detection system as described in claim 1, characterized in that, The feature encoding module performs the following operations: The text information is encoded based on a pre-trained encoding model to obtain the corresponding text semantic representation and construct a text feature matrix; The encoding model is used to extract features from image information, obtain visual semantic representation, and construct an image feature matrix; The external common sense representation is vectorized using the same encoding method as the text information, generating a common sense feature matrix, so that text features, image features and common sense features are comparable in a unified representation space.
4. The bidirectional adapter multimodal irony detection system as described in claim 1, characterized in that, The common sense modeling module performs the following operations: Self-attention computation is performed on the commonsense feature matrix to construct the contextual dependencies between the components of commonsense, in order to obtain an updated commonsense feature matrix. The self-attention computation is as follows: in, This is the original common sense feature matrix. This is a learnable weight matrix used to map features to query vectors. This is a learnable weight matrix used to map features to key vectors. This is a learnable weight matrix used to map features to value vectors. The dimension of the feature vector. For normalization function, This is the updated common sense feature matrix after self-attention weighted fusion; Based on the self-attention updated common sense feature representation, the semantic contributions of each component of common sense are weighted and aggregated to extract overall semantic information, resulting in a generalized common sense semantic vector for each sample. For the b-th sample, there exists a generalized common sense semantic vector. Represented as: in, For the first In the nth sample Each component unit calculates the updated features using self-attention. For learnable projection vectors, Indicates the first The importance weight of each component of common knowledge in the current round of semantic aggregation. The total number of constituent units contained in each common sense sample; Ultimately, the generalized common sense semantic vector will be generated. By mapping the textual and visual information using trainable projection matrices respectively, common-sense semantic guidance vectors for both modalities are generated, with the following mapping formula: in, For the first The visual modality commonsense semantic guidance vector corresponding to each sample. For the first The text modality commonsense semantic guidance vector corresponding to each sample. and These are trainable projection matrices specific to the visual modality and the text modality, respectively.
5. The bidirectional adapter multimodal irony detection system as described in claim 1, characterized in that, In the common sense perception bidirectional adaptation and fusion module: (1) Perform linear transformation and nonlinear mapping on the common sense semantic guidance vector corresponding to the text information to generate semantic modulation components, and apply them to the common sense semantic guidance vector corresponding to the visual information through residual injection, so that the visual information obtains semantic adjustment from the text information in the current round, as follows: in, This serves as the image modality commonsense guiding vector updated after the injection of textual semantics. The modal commonsense guide vector for the original image. The modulation coefficients of the visual image. and Two layers of linear weights on the text side are used to perform non-linear transformations on the text-guided vector. This is the original text modality commonsense guide vector. and Two layers of offset for the text side. For activation function, It is a random deactivation function; (2) After performing self-attention modeling on the expanded feature sequence, the updated common sense adaptation vector is extracted from the corresponding position, while keeping the structure of the original feature matrix unchanged, as follows: in, This is an enhanced image feature matrix that incorporates common-sense semantics after BERT encoding. This represents the self-attention modeling operation of the BERT encoder. This indicates a splicing operation. This refers to the global CLS vector of the image output by the CLIP encoder. This indicates the local patch features of the image output using the CLIP encoder. For image attention mask, This is an enhanced text feature matrix that incorporates common-sense semantics after BERT encoding. This is the text global CLS vector output by the CLIP encoder. This represents the text token characteristics output using the CLIP encoder. For text attention mask; The final output includes the text modality and image modality feature matrices after semantic interaction, as well as the common sense adaptation vectors and updated common sense feature matrices for the two modalities in the final round. .
6. The bidirectional adapter multimodal irony detection system as described in claim 5, characterized in that, To prevent the gradual decay of historical semantic information during multiple rounds of semantic modulation, the common sense-aware bidirectional adaptation and fusion module introduces a cross-round gating update mechanism during the bidirectional semantic modulation and iterative update of the text feature matrix and image feature matrix. In this cross-round gating update mechanism: In the initial iteration rounds, the text and image vectors output by the initially configured CLIP encoder are used as initial references for computation. During the bidirectional semantic modulation and iterative update of the text feature matrix and image feature matrix in the cross-wheel gating update mechanism: in, and For trainable text gating functions and images, it is used to dynamically adjust the information retention and update ratio based on the semantic relationship between the current round and the historical representation. For the first The image commonsense adaptation vector is used as the embedded image adaptation vector. For the first The commonsense text adaptation vector is used as the embedded text adaptation vector. This represents element-wise product.
7. The bidirectional adapter multimodal irony detection system as described in claim 1, characterized in that, The modality fusion module performs the following operations: The input image feature matrix is iteratively updated in the commonsense-aware bidirectional adaptation fusion module. Text feature matrix The final updated common sense feature matrix and the corresponding modal attention mask ; The image feature matrix, text feature matrix, and commonsense feature matrix are sequentially combined along the token dimension to form a unified feature sequence representation. Simultaneously, the corresponding attention masks are consistently combined to obtain the overall mask representation corresponding to the feature sequence. ; Representing a unified feature sequence and the corresponding attention encoding Input to multi-layer BERT encoder In this process, self-attention is used to model the dependencies between different positions, thereby obtaining feature representations that include contextual information. The final output is a feature matrix that integrates text, images, and external common sense information.
8. The bidirectional adapter multimodal irony detection system as described in claim 1, characterized in that, The irony detection module performs the following operations: The cross-modal fusion feature matrix output by the input modal fusion module , and the global semantic representation vectors corresponding to the text modality and the image modality; Cross-modal fusion feature matrix The input is fed into a fully connected layer and then normalized using the Softmax function to obtain an irony prediction result based on fused features. ; The CLS global semantic vectors extracted from the text modality and the image modality are processed separately. In the processing, each CLS vector is input into the corresponding fully connected layer and then passed through the Softmax function to obtain the individual irony prediction results for the text modality and the image modality. Predicting fused features Text modality prediction Image modality prediction According to the preset weights , , The weighted average yields the model's final ironic prediction result: ; During the training phase, supervision is applied simultaneously to the three prediction results, and the model parameters are jointly optimized using the binary classification cross-entropy loss function: in, To optimize the total loss during model backpropagation, Indicates branch Belonging to text ,image Integration Three branches, For the actual labels of the corresponding branches, The corresponding branch prediction probability output by each branch model; Finally, the final prediction result is output.
9. A method for detecting multimodal irony in a bidirectional adapter, characterized in that, The implementation of the bidirectional adapter multimodal irony detection system based on any one of claims 1-8 includes: Step S1. Perform joint semantic parsing on the input text and image, generate candidate external common sense based on the existing multimodal large language model and common sense reasoning model, and obtain credible external common sense through semantic and sentiment consistency screening. Step S2. Use a unified pre-trained encoder to encode the text, image and filtered external common sense respectively to obtain the text feature matrix, image feature matrix and common sense feature matrix in the same feature space; Step S3. Perform self-attention modeling and weighted aggregation on the common sense feature matrix to generate a generalized common sense semantic vector, and project it into a common sense semantic guidance vector that is adapted to the text modality and the image modality; Step S4. Based on the common sense semantic guidance vector, perform bidirectional semantic modulation and cross-wheel gating update on text features and image features, and at the same time write back the interactive semantics to update the common sense features; Step S5. Concatenate the enhanced text features, image features, and updated common sense features along the token dimension, and combine them with the corresponding attention mask in the joint encoder to obtain multimodal fusion features; Step S6. Based on the multimodal fusion features, the text global semantic vector, and the image global semantic vector, perform irony prediction respectively, and weightedly fuse the three prediction results to obtain the final irony detection result. The model training is completed by using multi-branch joint cross-entropy loss.
10. A computer device, characterized in that, The system includes a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform module steps of the system as described in any one of claims 1 to 8.