Bi-directional cross attention and gating mechanism fused multi-mode siphonage identification method

Through the multimodal irony recognition method that integrates two-way cross attention and gating mechanism, the problem of insufficient multimodal irony recognition ability in the existing technology is solved, and a deeper understanding of context and improvement of irony recognition accuracy is achieved.

CN120105232APending Publication Date: 2025-06-06JIANGSU OCEAN UNIV +1
View PDF 0 Cites 14 Cited by

Patent Information

Application Number
CN202510179968.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

It is difficult for the prior art to effectively identify and understand multimodal irony, especially on social media. The sparse distribution and shallow relationship between single-modal and multimodal irony features are not fully resolved, resulting in weak correlation between modals.

Method used

A multimodal ironic recognition method that integrates two-way cross attention and gating mechanism is adopted, and a multi-dimensional shallow emotional characteristics are coordinated through the two-way multi-layer cross attention mechanism and gating mechanism to build a deep emotional understanding model, and the fusion of image and text features is dynamically regulated through the gated unit system.

Benefits of technology

It improves the accuracy of irony recognition, deepens context understanding, enhances the processing ability of multimodal data, and accurately recognizes and understands the irony and humor hidden behind the literal meaning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120105232A_ABST
    Figure CN120105232A_ABST
Patent Text Reader

Abstract

In order to solve the problem that the bottleneck exists when image-text features are fused in a traditional anti-modal recognition model and deep-level interaction information between modals cannot be fully mined in an existing method, the invention provides a multi-modal anti-modal recognition model (MSCAG) fusing two-way cross attention and a gating mechanism, and a technical scheme is provided for solving the problem that the traditional anti-modal recognition model has the bottleneck when the image-text features are fused in the traditional anti-modal recognition model. Related attention features on a text layer and an image layer are respectively refined through a bidirectional cross attention mechanism, the text attention features, the images and basic features of the texts are integrated through residual connection, and the image attention features, the texts and the basic features of the images are integrated through residual connection; a gating mechanism is applied to enhance information interaction and key area focusing between the two modes. And the context features and the local features are combined to realize more reliable chiffon recognition. The invention provides an innovative method for multi-mode siphonage recognition, has important practical application value, and can be applied to siphonage recognition of netizen comments in social media.
Need to check novelty before this filing date? Find Prior Art

Description

Technical field:

[0001] The present invention proposes a multimodal irony recognition method integrating bidirectional cross-attention and gating mechanism, which provides an innovative method for multimodal irony recognition, has important practical application value, and can be applied to irony recognition in social media. Background technology:

[0002] With the growing development of social media, more and more domestic netizens are keen to share their daily experiences and participate in discussions on hot current events on social networking platforms such as Tik Tok and Weibo, and communications have also expanded from text to pictures, audio, etc. Sarcasm is a subtle and complex form of emotional expression, and its essence lies in its ability to convey hidden meanings or subtexts. Netizens frequently use sarcastic language information as a sharp tool to express their opinions, which undoubtedly brings unique challenges to sentiment analysis. Therefore, irony recognition plays a vital role in improving the effectiveness of sentiment recognition, topic detection, question-answering systems, and opinion mining, which undoubtedly makes cross-modal irony recognition a key research area.

[0003] In addition, these studies have not fully addressed the sparse distribution and shallow relationship between unimodal and multimodal ironic features, resulting in weak correlation between modalities. In order to more effectively cope with challenges in complex situations, this model innovatively adopts multimodal sentiment analysis integration technology. This strategy cleverly integrates rich information from different data types to build a more comprehensive and detailed contextual framework. In the process, the bidirectional multi-layer cross-attention mechanism and gating mechanism are used to cleverly coordinate and integrate multi-faceted shallow sentiment features to build a deep sentiment understanding model. This new technology ensures that the model can more accurately capture the subtle meanings in speech, and even the irony and humor hidden behind the literal meaning can be accurately identified and understood. Summary of the invention:

[0004] In view of the shortcomings of the prior art, the present invention proposes a multimodal irony recognition method that integrates bidirectional cross-attention and gating mechanisms. Through cross-modal feature fusion, the problem of insufficient single-modal recognition ability in existing research is effectively compensated, the context understanding is deepened, and the accuracy of irony recognition is improved. An innovative gating unit system is designed to dynamically regulate the fusion process of image and text features. An innovative bidirectional multi-layer cross-modal attention mechanism is proposed, which realizes deep interaction between image and text features by implementing a bidirectional cross-type, and integrates them into an auxiliary attribute modality, and then extracts emotional data from each modality and combines them together to promote deeper interaction between image and text information. The present invention proposes a multimodal irony recognition method that integrates bidirectional cross-attention and gating mechanisms, which consists of the following steps:

[0005] S1: Use the pre-trained BERT-base to convert the text sequence into a word vector with position and paragraph information, and process it through a multi-layer encoder to extract context-related semantic features;

[0006] S2: Using ViT pre-trained on ImageNet, the 224×224 standardized image Iproc is patched (16×16) and mapped to the vector space through linear transformation to extract the global relational features of the image;

[0007] S3: A bidirectional multi-layer cross-attention mechanism is adopted to combine the extracted text features with the image features through image-guided text attention (image to text) and text-guided image attention (text to image); in addition, it also includes reverse cross-attention from text to image and reverse cross-attention from image to text to ensure bidirectional information transmission and optimize feature fusion.

[0008] S4: Dynamically adjust the weights of image and text features through the gated unit module to achieve deep fusion when the image and text are aligned, and suppress fusion in the case of misalignment to retain the original features, ensuring that the fused features effectively retain the salient characteristics of the image and text.

[0009] S5: Input the image-text fusion features into the fully connected layer for linear transformation, and predict the category probabilities of irony and non-irony through the sigmoid activation function; combine the joint loss function to train the model to improve the accuracy of irony recognition.

[0010] 2. According to the multimodal irony recognition method integrating bidirectional cross attention and gating mechanism in claim 1, it is characterized in that the specific steps of S1 are as follows:

[0011] S1-1: Input text sequence T is converted into word vector V. Each text sequence t i Mapped to a corresponding word vector v i , and add position and paragraph to preserve the position information of the vocabulary.

[0012] S1-2: These embedding vectors are input into BERT’s multi-layer encoder to obtain the contextual information H of each word. Each layer of the encoder contains a multi-head self-attention mechanism and a feedforward neural network to capture the complex dependencies in the text. The weight w of each word is calculated using a weighting mechanism i :

[0013] w i =Attention(h i )

[0014] S1-3: By weighted averaging the information of all words, we get the shallow feature c of the entire text:

[0015]

[0016] Among them, w i It is vocabulary i The weight of h i It is vocabulary i Context-aware representation of N×D It is the shallow feature of the extracted semantic information of the text, N is the sequence length, and D is the dimension of the feature vector.

[0017] 3. According to the multimodal irony recognition method integrating bidirectional cross attention and gating mechanism as described in claim 1, it is characterized in that the specific steps of S2 are as follows:

[0018] ViT is selected as the image encoder to extract shallow features of the image. It overcomes the limitations of CNN and mainly uses the encoding module of Transformer to model the global relationship of the image, which can effectively process complex patterns in the image. First, the input image I needs to be adjusted to a size of 224×224 pixels, and the image pixel values ​​are normalized to the interval [0,1] for standardization. The preprocessed image is represented as Iproc. Subsequently, the input image is divided into a series of fixed-size patches, each of which is 16×16 in size, and the image is divided into multiple patches. Each patch is flattened into a one-dimensional vector and mapped to the vector space through a linear transformation:

[0019] x i =W·flatten(p i )+b

[0020] Among them, W and b are the weight matrix and bias term, x i is the feature vector of the ith patch after linear transformation.

[0021] S2-1: To retain the position information, each patch vector is added with a position code e i ;

[0022] S2-2: concatenate all patch vectors together and add a CLStokenc at the beginning to get the sequence Z;

[0023] Z = [c; x′ 1 ; x′ 2 ; ... ; x′ 196 ]

[0024] S2-3: After the sequence Z is sent to the Transformer encoder for processing, the output of the middle layer is selected as the shallow feature of the image and global average pooling is performed:

[0025] y=GAP(z p )

[0026] Where y∈H P×D It is the shallow image features extracted from the input image, P is the sequence length, and D is the dimension of the feature vector.

[0027] 4. According to the multimodal irony recognition method integrating bidirectional cross attention and gating mechanism in claim 1, it is characterized in that the text attention feature and the image attention feature are cross-fused by the bidirectional cross attention module in S3, which specifically includes the following contents:

[0028] S3-1: The bidirectional cross-multilayer attention mechanism is an organic fusion of self-attention and multi-layer cross-attention. Through bidirectional information transfer, the understanding of each modal context can be enhanced. The mechanism consists of image-guided text cross-attention, text-guided image cross-attention, text-to-image reverse cross-attention, and image-to-text reverse cross-attention. Thereby improving the processing ability of multimodal data;

[0029] S3-2: When image features and text features are converted into numerical vectors through the embedding layer, a self-attention module is introduced to enhance feature representation, and the weighted combination information is calculated by calculating the importance of each position;

[0030] v'=SelfAttention(y)

[0031] t'=SelfAttention(c)

[0032] S3-3: In the image-guided text cross-attention mechanism, the image feature v' is used as the query Q, the original text shallow feature c is used as the key K and the value V, and the image cross-attention output v' is obtained. 1 ; In the text-guided image cross-attention mechanism, the text feature t' is used as the query Q, the original image shallow feature y is used as the key K and value V, and the text cross-attention output t' is obtained 1 :

[0033]

[0034] Where W Q , W K and W V are all learnable parameters, and K is the dimension of the query Q vector.

[0035] S3-4: Then the image cross-attention feature v' 1 Fusion with image feature v' to obtain image enhancement feature v' 2 Similarly, we can get the text enhancement feature t' 2 :

[0036] v' 2 =concat(v' 1 ,v')

[0037] t' 2 =concat(t' 1 ,t')

[0038] S3-5: Use text feature t' as query Q, enhanced image feature v' 2 As key K and value V, calculate the reverse cross-attention features of text to image:

[0039]

[0040] Use image feature v' as query Q and enhanced text feature t' 2 As key K and value V, calculate the reverse cross-attention feature t' from image to text v :

[0041]

[0042] Where W Q , W K and W V are all learnable parameters, and K is the dimension of the query Q vector.

[0043] S3-6: v' t and t' v The final concatenation is performed to obtain the image-text joint feature Z vt :

[0044] Z vt =[v' t ; t' v ]

[0045] 5. According to the multimodal irony recognition method integrating bidirectional cross attention and gating mechanism in claim 1, it is characterized in that the S4 specifically comprises the following contents:

[0046] S4-1: Control the fusion weights between different modal data through the gating unit, dynamically adjust the weights between different modal data, enhance the multimodal fusion interaction of the model, and calculate a comprehensive gating signal g to dynamically adjust the image text feature weights:

[0047] g=σ(W g [v' t ; t' v ]+b g )

[0048] Where W gis the weight matrix used to learn the linear combination of gating signals; b g is the bias term, which is used to adjust the baseline value of the gated signal;

[0049] S4-2: The gating signal is then used to adjust the text and image features, and the tanh activation function increases the nonlinear ability of the model;

[0050] H t '=g⊙tanh(t' v )

[0051] H v '=(1-g)⊙tanh(v' t )

[0052] S4-3: In the formula, ⊙ represents element-level multiplication, and tanh is the hyperbolic tangent activation function, which is used to increase the nonlinear ability of the model and help the model better capture the complex relationship between features. Finally, H' t and H' V Get the eigenvector F:

[0053] F=[H t ';H v ']

[0054] 6. According to the multimodal sentiment analysis method based on cross-attention gating unit in claim 1, it is characterized in that the specific steps of S5 are as follows:

[0055] S5-1: Average pooling of the word-level features T of each text to obtain the contextual text features S of the entire sentence;

[0056]

[0057] S5-2: The model is trained through a joint loss function to improve the model's irony accuracy. First, the image-text fusion feature F is input into a fully connected layer for linear transformation, and then the result of the linear transformation is converted into a probability value using the sigmoid activation function, and the irony and non-irony category probabilities are output. Label prediction value y pred :

[0058] y pred =σ(W o ·F+b o )

[0059] S5-3: In the formula, W o is the weight matrix, b o is the bias term, σ is the sigmoid activation function, and y pred ∈[0,1]. Then according to the real label y true and the predicted value y predCalculate binary cross entropy loss

[0060]

[0061] S5-4: Where n is the number of samples in the training set, y i is the true label of the i-th sample, y pred is the predicted probability of the i-th sample. As above, calculate the image classification loss Combining text classification loss and image classification loss The joint loss function L is formed by weighted combination:

[0062]

[0063] S5-5: In the formula, λ is a hyperparameter used to balance and Finally, the Adam optimizer is used to perform gradient descent, update the model parameters, and minimize the joint loss function L. According to the predicted value y pred To determine whether the sample is ironic, the threshold is set to 0.5. pred >0.5, it is identified as irony; otherwise, it is identified as non-irony.

[0064] Compared with the relevant prior art, this application proposal has the following main technical advantages:

[0065] Deep mining of correlation: The MSCAG model introduces a bidirectional multi-layer cross-attention module, which effectively overcomes the problem that traditional methods only perform simple splicing at the feature level and ignore the potential correlation between images and texts. This module can accurately capture and quantify the complex consistency and expression relationship between images and texts, extract interrelated high-level features, and enhance the model's understanding of deep interactive information between images and texts. In addition, the gating mechanism used by MSCAG further coordinates and integrates multi-faceted shallow emotional features, ensuring deep fusion when images and texts are aligned and the original features are retained when they are not aligned, achieving accurate recognition and understanding of subtle emotions such as irony and humor.

[0066] Adaptive multimodal coordination: In the field of multimodal irony recognition, MSCAG has achieved dynamic regulation of the image and text sentiment feature fusion process by introducing a unique adaptive gating mechanism. This gating mechanism can dynamically adjust the contribution of each modality according to the input correlation features, ensuring that key ironic information is effectively retrieved in different contexts. This adaptive regulation not only improves the adaptability of the model in irony recognition tasks, but also significantly enhances its accuracy.

[0067] Combining contextual awareness with local features: This paper uses BERT technology to extract contextual features of text and combines local features for sentiment prediction. By integrating global context information, the model can grasp the overall trend of sarcastic sentiment; at the same time, it focuses on local details to reveal subtle fluctuations in sentiment. This fusion of global and local features provides more comprehensive irony recognition.

[0068] Practical application value: Irony recognition plays a vital role in improving the effectiveness of sentiment recognition, topic detection, question-answering systems, and opinion mining. The MSCAG model's deep understanding of graphic data and accurate irony recognition capabilities help companies, research institutions, and social organizations better understand public sentiment and social public opinion dynamics. Description of the drawings:

[0069] Figure 1 It is a structural diagram of a MSCAG model provided by the present invention;

[0070] Figure 2 It is a diagram of the bidirectional cross attention fusion mechanism provided by the present invention; Specific implementation method:

[0071] The present invention is further described below in conjunction with the accompanying drawings and examples. However, the present invention can be implemented in many different ways and should not be construed as being limited to the embodiments shown; on the contrary, these embodiments provide those skilled in the art with implementation methods that meet applicable legal requirements.

[0072] Dataset: This experiment aims to address the challenges of Chinese multimodal irony recognition, especially the shortage of Chinese datasets and the uneven quality. We built a Chinese dataset specifically for multimodal irony recognition in public emergencies. We used web crawler technology to crawl a large number of netizens' comments from Weibo and Douyin, two mainstream domestic social platforms. These comments contain both image and text modal information. We manually removed redundant data and irrelevant comments, and then normalized the text to unify the encoding format. At the same time, we standardized the size of the images so that the images and text can be better integrated.

[0073] In addition, this paper is also committed to the reconstruction of contextual content. By connecting the comments with the context, each data point carries rich information, which enhances the model's sensitivity and accuracy in identifying irony. In the labeling stage of the dataset, 8,000 ironic comments and 7,200 non-ironic comments were locked from the 18,800 data initially screened, forming a balanced binary classification sample set. Non-ironic comments are clearly marked as "0", while those containing irony are given a label of "1". This binary label system simplifies the difficulty of understanding the model. 15,200 corresponding comment images are also collected. These image materials are not only an intuitive supplement to the comments, but also provide a basis for the subsequent training and development of the multimodal irony model. The final dataset is divided into training set, validation set and test set in a ratio of 8:1:1. The specific statistical data are shown in Table 1. In order to comprehensively evaluate the experimental results, the model uses accuracy (Acc), precision (P), recall (R) and F1 as performance indicators.

[0074] Table 1 Statistics of the multimodal irony dataset

[0075]

[0076] The hardware configuration and software environment of this experiment are shown in Table 2.

[0077] Table 2 Hardware configuration and software environment

[0078]

[0079] In order to verify the effectiveness of the MSCAG method, comparative experiments were carried out on the dataset from three aspects: text, image, and graphic. The experimental results are shown in Table 4.

[0080] According to the Acc and F1 values, the MSCAG method proposed in this invention is superior to the current mainstream multimodal irony recognition methods.

[0081] Table 3. Text, image and graphic comparison experiments

[0082]

[0083] In summary, the present invention proposes a multimodal irony recognition method that integrates bidirectional cross-attention and gating mechanisms. The multimodal irony recognition model MSCAG that integrates bidirectional cross-attention and gating mechanisms proposed in the present invention can efficiently interact information between different modalities through the bidirectional cross-attention mechanism, capture cross-modal correlation features, thereby reducing the heterogeneity gap between image and text features, while the gating technology dynamically adjusts the weight of each modal information to highlight the text information content and enhance the model's adaptability to complex ironic expressions. Experiments have found that the performance of the model after incorporating the bidirectional cross-attention mechanism and the gating mechanism is significantly better than the baseline model, and the superiority of the model of the present invention in multimodal irony recognition has been verified.

[0084] The above embodiment only expresses one implementation mode of the present invention, and its description is relatively specific and detailed, but it cannot be understood as limiting the scope of the invention patent. It should be pointed out that for ordinary technicians in this field, several modifications and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention.

Claims

1. A multimodal irony recognition method integrating bidirectional cross attention and gating mechanism, characterized in that: The specific steps are as follows: S1: Use the pre-trained BERT-base to convert the text sequence into a word vector with position and paragraph information, and process it through a multi-layer encoder to extract context-related semantic features; S2: Using ViT pre-trained on ImageNet, the 224×224 standardized image Iproc is patched (16×16) and mapped to the vector space through linear transformation to extract the global relational features of the image; S3: A bidirectional multi-layer cross-attention mechanism is adopted to combine the extracted text features with the image features through image-guided text attention (image to text) and text-guided image attention (text to image); in addition, it also includes reverse cross-attention from text to image and reverse cross-attention from image to text to ensure bidirectional information transmission and optimize feature fusion. S4: Dynamically adjust the weights of image and text features through the gated unit module to achieve deep fusion when the image and text are aligned, and suppress fusion in the case of misalignment to retain the original features, ensuring that the fused features effectively retain the salient characteristics of the image and text. S5: Input the image-text fusion features into the fully connected layer for linear transformation, and predict the category probabilities of irony and non-irony through the sigmoid activation function; combine the joint loss function to train the model to improve the accuracy of irony recognition.

2. According to claim 1, a multimodal irony recognition method integrating bidirectional cross attention and gating mechanism is characterized in that: The specific steps of S1 are as follows: S1-1: Input text sequence T is converted into word vector V. Each text sequence t i Mapped to a corresponding word vector v i , and add position and paragraph to preserve the position information of the vocabulary. S1-2: These embedding vectors are input into BERT’s multi-layer encoder to obtain the contextual information H of each word. Each layer of the encoder contains a multi-head self-attention mechanism and a feedforward neural network to capture the complex dependencies in the text. The weight w of each word is calculated using a weighting mechanism i : w i =Attention(h i ) S1-3: By weighted averaging the information of all words, we get the shallow feature c of the entire text: Among them, w i It is vocabulary i The weight of h i It is vocabulary i Context-aware representation of N×D It is the shallow feature of the extracted semantic information of the text, N is the sequence length, and D is the dimension of the feature vector.

3. According to claim 1, a multimodal irony recognition method integrating bidirectional cross attention and gating mechanism is characterized in that: The specific steps of S2 are as follows: ViT is selected as the image encoder to extract shallow features of the image. It overcomes the limitations of CNN and mainly uses the encoding module of Transformer to model the global relationship of the image, which can effectively process complex patterns in the image. First, the input image I needs to be adjusted to 224×224 pixel size, and the image pixel value is normalized to the [0,1] interval for standardization. The preprocessed image is represented as I proc . Subsequently, the input image is divided into a series of fixed-size patches, each of which is 16×16 in size, and the image is divided into multiple patches. Each patch is flattened into a one-dimensional vector and mapped into the vector space through a linear transformation: x i =W·flatten(p i )+b Among them, W and b are the weight matrix and bias term, x i is the feature vector of the ith patch after linear transformation. S2-1: To retain the position information, each patch vector is added with a position code e i ; S2-2: concatenate all patch vectors together and add a CLStoken c at the beginning to get sequence Z; Z=[c;x1′;x′2;...;x1′ 96 ] S2-3: After the sequence Z is sent to the Transformer encoder for processing, the output of the middle layer is selected as the shallow feature of the image and global average pooling is performed: y=GAP(z p ) Where y∈H P×D It is the shallow image features extracted from the input image, P is the sequence length, and D is the dimension of the feature vector.

4. According to claim 1, a multimodal irony recognition method integrating bidirectional cross attention and gating mechanism is characterized in that: In S3, the text attention features and the image attention features are cross-fused through a bidirectional cross attention module, which specifically includes the following contents: S3-1: The bidirectional cross-multilayer attention mechanism is an organic fusion of self-attention and multi-layer cross-attention. Through bidirectional information transfer, the understanding of each modal context can be enhanced. The mechanism consists of image-guided text cross-attention, text-guided image cross-attention, text-to-image reverse cross-attention, and image-to-text reverse cross-attention. Thereby improving the processing ability of multimodal data; S3-2: When image features and text features are converted into numerical vectors through the embedding layer, a self-attention module is introduced to enhance feature representation, and the weighted combination information is calculated by calculating the importance of each position; v'=SelfAttention(y) t'=SelfAttention(c) S3-3: In the image-guided text cross-attention mechanism, the image feature v' is used as the query Q, the original text shallow feature c is used as the key K and the value V, and the image cross-attention output v'1 is obtained; in the text-guided image cross-attention mechanism, the text feature t' is used as the query Q, the original image shallow feature y is used as the key K and the value V, and the text cross-attention output t'1 is obtained: Where W Q , W K and W V are all learnable parameters, and K is the dimension of the query Q vector. S3-4: Then the image cross-attention feature v'1 is fused with the image feature v' to obtain the image enhancement feature v'2. Similarly, the text enhancement feature t'2 can be obtained: v'2=concat(v'1,v') t'2=concat(t'1,t') S3-5: Using the text feature t' as the query Q, the enhanced image feature v'2 as the key K and the value V, calculate the reverse cross attention feature from text to image: Using the image feature v' as the query Q, the enhanced text feature t'2 as the key K and the value V, calculate the image-to-text reverse cross-attention feature t' v : Where W Q , W K and W V are all learnable parameters, and K is the dimension of the query Q vector. S3-6: v' t and t' v The final concatenation is performed to obtain the image-text joint feature Z vt : Z vt =[v' t ;t' v ]。 5. According to claim 1, a multimodal irony recognition method integrating bidirectional cross attention and gating mechanism is characterized in that: The S4 specifically includes the following contents: S4-1: Control the fusion weights between different modal data through the gating unit, dynamically adjust the weights between different modal data, enhance the multimodal fusion interaction of the model, and calculate a comprehensive gating signal g to dynamically adjust the image text feature weights: g=σ(W g [v' t ;t' v ]+b g ) Where W g is the weight matrix used to learn the linear combination of gating signals; b g is the bias term, which is used to adjust the baseline value of the gated signal; S4-2: The gating signal is then used to adjust the text and image features, and the tanh activation function increases the nonlinear ability of the model; H t '=g⊙tanh(t' v ) H v '=(1-g)⊙tanh(v' t ) S4-3: In the formula, ⊙ represents element-level multiplication, and tanh is the hyperbolic tangent activation function, which is used to increase the nonlinear ability of the model and help the model better capture the complex relationship between features. Finally, H' t and H' V Get the eigenvector F: F=[H t ';H v ']。 6. The multimodal sentiment analysis method based on cross-attention gating unit according to claim 1, characterized in that: The specific steps of S5 are as follows: S5-1: Average pooling of the word-level features T of each text to obtain the contextual text features S of the entire sentence; S5-2: The model is trained through a joint loss function to improve the model's irony accuracy. First, the image-text fusion feature F is input into a fully connected layer for linear transformation, and then the result of the linear transformation is converted into a probability value using the sigmoid activation function, and the irony and non-irony category probabilities are output. Label prediction value y pred : y pred =σ(W o ·F+b o ) S5-3: In the formula, W o is the weight matrix, b o is the bias term, σ is the sigmoid activation function, and y pred ∈[0,1]. Then according to the real label y true and the predicted value y pred Calculate binary cross entropy loss S5-4: Where n is the number of samples in the training set, y i is the true label of the i-th sample, y pred is the predicted probability of the i-th sample. As above, calculate the image classification loss Combining text classification loss and image classification loss The joint loss function L is formed by weighted combination: S5-5: In the formula, λ is a hyperparameter used to balance and Finally, the Adam optimizer is used to perform gradient descent, update the model parameters, and minimize the joint loss function L. According to the predicted value y pred To determine whether the sample is ironic, the threshold is set to 0.

5. pred >0.5, it is identified as irony; otherwise, it is identified as non-irony.

Citation Information

Cited By

  • Asphalt concrete image segmentation system and method based on gating fusion mechanism

    CN120747151A

  • An asphalt concrete image segmentation system and method based on a gating fusion mechanism

    CN120747151B

  • Power equipment image-text fusion labeling method and system based on single-double hybrid tower

    CN121053657A

  • Social medium anti-mental semantic recognition method and system, storage medium and electronic equipment

    CN121093970A

  • Social media irony semantic recognition method and system, storage medium and electronic device

    CN121093970B