Multi-modal large language model metaphor characterization method based on edge detection

Through edge detection technology, cross-layer features are extracted from multimodal large language models and weighted fusion is performed. Metaphor coding is verified using MLP classifiers, which solves the unknown of the model coding mechanism in multimodal metaphor detection, and realizes the metaphorical representation enhancement of multimodal combinatorial input.

CN120448919APending Publication Date: 2025-08-08TIANJIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510584407.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The prior art lacks in-depth discussion on the metaphor encoding mechanism of multimodal large language model in multimodal metaphor detection, and multimodal metaphor detection faces the problems of semantic complexity, diversity and semantic alignment between modals.

Method used

Edge detection technology is used to extract cross-layer features from the hidden representation of multimodal large language models, and weighted representations of a unified dimension are generated through linear projection, adaptive attention pooling and hierarchical weighted fusion. Multi-layer perceptron classifiers are used for metaphor classification, and the freeze model parameters are only trained to explore whether the model encodes metaphorical information.

Benefits of technology

By evaluating the performance of the classifier, revealing whether the multimodal large language model encodes metaphor information, it provides theoretical support for metaphor detection of multimodal models, which is highly versatile and interpretable, and verifies that multimodal combined input can enhance metaphorical representation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120448919A_ABST
    Figure CN120448919A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal large language model metaphor characterization method based on edge detection, and the method is based on a characterization extraction module, a pooling module, a hierarchical module and a classifier, and comprises an offline stage and a training stage. The method comprises the following steps: inputting data containing metaphor and non-metaphor labels into a multi-modal large language model to extract hidden representation data h; performing linear projection on the hidden representation h according to a 256-dimensional space to obtain a hidden representation tensor matrix z; pooling and normalizing the hidden representation tensor matrix through an adaptive attention mechanism to obtain a hidden pooling representation matrix p (i); performing hierarchical weighted summation on the hidden pooling representation matrix to obtain hidden weighted representation v; for the hidden weighted representation, predicting whether input data contains a metaphor label # imgabs0 # or not, using a cross entropy loss function to predict the difference between a result # imgabs1 # and a real label, and continuously adjusting parameters of a classifier through a back propagation algorithm; according to the method, whether the multi-modal large language model encodes metaphor information or not can be explored only by training a small-scale classifier, and an effective analysis tool is provided for subsequent research on the semantic representation capability of the multi-modal model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to related fields such as multimodal large language models, metaphor detection and explainable artificial intelligence, and proposes a metaphor representation method for multimodal large language models based on edge detection. Background Art

[0002] Metaphor, as a complex linguistic phenomenon, is an indispensable component of human cognition and language expression. Within the framework of Conceptual Metaphor Theory (CMT), metaphors concretize abstract concepts by mapping the characteristics of the source domain to the target domain [1]. This mechanism is widely present in human daily language. For example, the metaphor "time is money" emphasizes the preciousness of time by conveying the scarcity and value of money. This cross-domain analogy not only enhances the expressiveness of language, but also promotes a deeper understanding of complex concepts. However, the diversity, context dependence, and high demand for background knowledge of metaphors make it a very challenging task in the field of natural language processing (NLP).

[0003] In traditional metaphor detection research, early methods mainly relied on rules and knowledge bases, using predefined linguistic rules to identify metaphors [2]. Although these methods can work effectively in certain specific scenarios, they generally lack adaptability to the diversity of metaphors, especially when faced with cross-domain and multimodal contexts. Subsequently, feature engineering-based methods introduced statistical and machine learning models to classify metaphors by extracting features such as word semantic similarity and syntactic dependencies [3]. However, these methods are highly dependent on language and tasks in feature selection, which limits their generalization ability.

[0004] In recent years, with the rapid development of deep learning technology, neural network-based methods have gradually replaced traditional methods and become the mainstream direction of metaphor detection. In particular, the rise of pre-trained language models (PLMs) has significantly improved the performance of metaphor detection [4]. For example, language models based on Transformer architectures such as BERT [5] [6] capture deeper semantic relationships through contextualized embeddings and have become an important tool for metaphor detection. However, these methods mainly focus on metaphor detection in a single modality (text) and lack attention to multimodal metaphors.

[0005] The rise of multimodal learning has opened up new research directions for metaphor detection technology. Joint modeling of multimodal data (such as images and text) enables the model to capture complex semantic relationships across modalities [7]. Some studies have attempted to combine images and text for metaphor detection tasks. For example, Xu et al. proposed the C4MMD framework and implemented multimodal metaphor detection using the Chain-of-Thought (CoT) method [8]. However, existing research mainly focuses on performance optimization, and lacks in-depth exploration of the specific mechanisms of multimodal large language models in metaphor knowledge encoding.

[0006] In addition, as the scale of models continues to expand, the interpretability of models has gradually become a research hotspot. In recent years, edge probing technology has been widely used to explain the semantic representation ability of pre-trained language models [9]. This technology reveals the model's representation ability for specific language phenomena by freezing model parameters and training additional classifiers. For example, Aghazadeh et al. successfully applied edge probing to metaphor research and proved that the intermediate layer of the pre-trained model can effectively encode metaphor knowledge

[10] . In summary, the field of multimodal metaphor detection is still in its early stages of development. Its main challenges include the semantic complexity and diversity of metaphors and the semantic alignment problem between modalities. At the same time, the metaphor encoding of large multimodal language models still needs further exploration. How to use interpretability methods to explore the metaphor representation within large multimodal language models has become an urgent problem to be solved.

[0007] [References]

[0008] [1]Lakoff G, Johnson M. Metaphors we live by[M]. University of Chicagopress,2008.

[0009] [2]Fass D.met*:A method for discriminating metonymy and metaphor bycomputer[J].Computational linguistics,1991,17(1):49-90.

[0010] [3]Turney P,Neuman Y,Assaf D,et al.Literal and metaphorical senseidentification through concrete and abstract context[C] / / Proceedings of the2011Conference on Empirical Methods in Natural Language Processing.2011:680-690.

[0011] [4]Song W,Zhou S,Fu R,et al.Verb metaphor detection via contextualrelation learning[C] / / Proceedings of the 59th Annual Meeting of theAssociation for Computational Linguistics and the 11th International JointConference on Natural Language Processing(Volume 1:Long Papers).2021:4240-4251.

[0012] [5]Devlin J.Bert:Pre-training of deep bidirectional transformers forlanguage understanding[J].arXiv preprint arXiv:1810.04805,2018.

[0013] [6]Vaswani A.Attention is all you need[J].Advances in NeuralInformation Processing Systems,2017.

[0014] [7]Liu H,Li C,Wu Q,et al.Visual instruction tuning[J].Advances inneural information processing systems,2024,36.

[0015] [8]Xu Y, Hua Y, Li S, et al. Exploring Chain-of-Thought for Multi-modalMetaphor Detection[C] / / Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024:91-101.

[0016] [9]Tenney I,Xia P,Chen B,et al.What do you learn from context? probingfor sentence structure in contextualized word representations[J].arXivpreprint arXiv:1905.06316,2019.

[0017]

[10] Aghazadeh E, Fayyaz M, Yaghoobzadeh Y. Metaphors in pre-trained language models: Probing and generalization across datasets and languages[J].arXiv preprint arXiv:2203.14139,2022. Summary of the Invention

[0018] This paper proposes a metaphor representation method for a multimodal large language model based on edge detection, aiming to explore the metaphor knowledge encoding of a multimodal large language model.

[0019] Based on edge detection technology, this paper extracts cross-layer features from the hidden representation of a large multimodal language model. Through projection, pooling, and weighted fusion operations, it generates a weighted representation with unified dimensions. This representation is used to train a multi-layer perceptron (MLP) classifier for binary classification tasks. By observing the performance of the MLP classifier, it is verified whether metaphorical information is encoded within the model. In the experiment, the paper designed two input scenarios: image and image-text pair. While exploring whether the large multimodal language model encodes metaphorical information, the paper also revealed the enhanced effect of multiple modalities on the encoding of metaphorical information through comparative analysis.

[0020] The present invention is implemented through the following technical solutions:

[0021] A metaphor representation method for a multimodal large language model based on edge detection is provided. The method is based on a representation extraction module, a pooling module, a hierarchical module, and a classifier, and includes the following steps:

[0022] Offline stage:

[0023] Input the data containing metaphorical and non-metaphorical labels into the multimodal large language model to extract the hidden representation data h;

[0024] The hidden representation h is linearly projected into 256-dimensional space according to the following formula to obtain the hidden representation tensor matrix z

[0025] z=W p h+b p

[0026] in, is the weight matrix of linear projection; is the bias vector, d is the dimension of the original hidden representation;

[0027] The hidden representation tensor matrix is pooled and normalized through the adaptive attention mechanism to obtain the hidden pooled representation matrix p (i) ;

[0028] The hidden weighted representation v is obtained by performing hierarchical weighted summation on the hidden pooling representation matrix through the following formula;

[0029]

[0030] Among them, α i Represents the weight of the i-th layer in the multimodal large language model;

[0031] The following formula is used to predict whether the hidden weighted representation contains metaphor labels in the input data

[0032]

[0033] Among them, W c and b c are the weight matrix and bias vector respectively, σ(·) is the Sigmoid activation function;

[0034] Training phase

[0035] The cross entropy loss function is used to predict the results The difference between the true label and the real label is calculated, and the parameters of the classifier are continuously adjusted through the back propagation algorithm.

[0036] Furthermore, the hidden representation tensor matrix is pooled and normalized through the adaptive attention mechanism to obtain the hidden pooled representation matrix p (i) Process, including:

[0037] Perform multi-head self-attention processing on the hidden representation tensor matrix z according to the dimension of the input tensor [B, L, T, D];

[0038] Where: B is the batch size, L is the number of layers, T is the sequence length, and D is the feature dimension;

[0039] Perform the maximum pooling operation on the representation after attention processing along the sequence length T:

[0040] p=max(z1,z2,...,z T )

[0041] The pooled representation is l2 normalized according to the following formula to eliminate the influence of feature scale:

[0042]

[0043] Furthermore, the weight α of the i-th layer in the multimodal large language model is i for:

[0044]

[0045] Where: i is a trainable parameter.

[0046] Beneficial effects

[0047] Compared with the prior art, the technical solution of the present invention has the following beneficial effects:

[0048] This paper uses edge detection technology to extract cross-layer features from the hidden representation of a large multimodal language model. Through projection, pooling, and weighted fusion, it generates a weighted representation of uniform dimensionality. This weighted representation is then fed into a multi-layer perceptron (MLP) classifier for training on metaphor classification tasks. By evaluating the classifier's performance, it is possible to determine whether metaphorical information is encoded within the model.

[0049] In the edge detection-based metaphor representation method for a multimodal large language model provided by the present invention, by freezing the model parameters, only a small-scale classifier needs to be trained to explore whether the model encodes metaphorical information. By revealing the encoding of metaphorical knowledge by the model, the present invention provides theoretical support for the study of metaphor detection in multimodal large language models. In addition, the method of the present invention can be applied to different multimodal large language models, has strong versatility, and provides an effective analysis tool for subsequent research on the semantic representation ability of multimodal models. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 is an example diagram of the method of the present invention. DETAILED DESCRIPTION

[0051] In order to more clearly understand the above-mentioned objects, features and advantages of the present invention, the scheme of the present invention will be further described below. The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to illustrate the present invention and are not intended to limit the present invention.

[0052] like Figure 1 As shown in the figure, the present invention proposes a multimodal large language model metaphor representation method based on edge detection, which includes the following modules: representation extraction module, adaptive attention pooling module, hierarchical weighted summation module and classifier module. The specific implementation of each module is as follows:

[0053] 1. Characterization Extraction Module

[0054] The purpose of the representation extraction module is to extract and project the representation of the multimodal large language model. The specific implementation is as follows:

[0055] 1.1 Extraction and characterization

[0056] Obtain the model's representation h of the input data from the multimodal large language model, which contains high-dimensional features of the input data (such as images or image-text pairs).

[0057] 1.2 Linear Projection

[0058] Since the representations of different layers in a large multimodal language model have different dimensions, in order to align the dimensions, we use a linear transformation to project the representation h to 256 dimensions. The projection calculation formula is as follows:

[0059] z=W p h+b p

[0060] in, and are the weight matrix and bias vector of the linear projection, respectively, and d is the dimension of the original representation.

[0061] 2. Adaptive Attention Pooling Module

[0062] The self-attention mechanism can focus on different parts of the input data and capture long-range dependencies. This is crucial for metaphorical representation, as metaphors often rely on long-range dependencies, such as cross-modal conceptual connections. Pooling can extract key information from the representation and generate a feature vector with global semantics. The specific implementation is as follows:

[0063] 2.1 Multi-head Self-Attention Mechanism

[0064] Perform multi-head self-attention on the projected tensor z. The input tensor has the dimension [B, L, T, D], where B is the batch size, L is the number of layers, T is the sequence length, and D is the feature dimension.

[0065] 2.2 Pooling

[0066] Perform the maximum pooling operation on the representation after attention processing along the sequence length T:

[0067] p=max(z1,z2,...,z T )

[0068] 2.3 Normalization

[0069] The pooled representation is l2 normalized to eliminate the influence of feature scale:

[0070]

[0071] After l2 normalization, all feature vectors are adjusted to the same magnitude, so that the features of different input samples are in a fair position during model training and comparison.

[0072] 3. Hierarchical Weighted Sum Module

[0073] The hierarchical weighted summation module aims to fuse feature representations from different levels to capture multi-level semantic information. The specific implementation is as follows:

[0074] 3.1 Layer weight calculation

[0075] Calculate the weight α of each layer through the softmax function i :

[0076]

[0077] By introducing a trainable parameter β i By using a softmax operation to calculate weights, the classifier can automatically assign weights based on the importance of metaphorical information representation at different layers. This process is adaptive and can be dynamically adjusted based on the characteristics of the training data and the learning progress of the model, thereby fully exploring the metaphor-related information in each layer.

[0078] 3.2 Weighted Sum

[0079] In order to integrate information from multiple layers, the pooled representations of each layer are weighted and summed according to the calculated weights to form the final weighted representation vector. Represents the representation of the i-th layer after maximum pooling and l2 normalization (i.e., p in the previous text normBy weighted fusion of these representations, the semantic information contained in different levels can be fully utilized. The final weighted representation v is calculated as follows:

[0080]

[0081] This weighted summation approach avoids the limitations of simply averaging information from each layer or relying on information from only one layer.

[0082] 4. Classifier Module

[0083] The classifier module is used to perform metaphor detection based on the weighted representations obtained in the previous step. A multi-layer perceptron (MLP) is used as the classifier, and its input is the weighted representation v. The calculation formula is as follows:

[0084]

[0085] Among them, φ(·) represents the activation function, σ(·) is the sigmoid function, and W c and b c are the parameters of the MLP layer, is the final prediction result. The activation function φ(·) introduces nonlinearity into the model, enabling it to learn complex nonlinear relationships in the input data. In this paper, the ReLU activation function was selected after evaluating and comparing the performance of different activation functions. The Sigmoid function is used to map the model output to a probability value between 0 and 1, facilitating binary classification. The classifier predicts whether the input contains metaphorical information based on the binary labels in the corresponding dataset.

[0086] During training, a cross-entropy loss function is used to measure the difference between the model's predictions and the true labels. Backpropagation is used to continuously adjust the MLP parameters, gradually reducing the loss function and optimizing the classifier's performance. Classifier performance reflects the degree to which the multimodal large language model encodes metaphorical knowledge. If the classifier's predictions match the labels, it indicates that the model has effectively encoded the relevant metaphorical information.

[0087] This example uses three datasets widely used in multimodal metaphor detection research as experimental materials: MultiCMET, MultiMET, and MetMEME. Based on the data source, MultiMET is further divided into three subsets: MultiMET_ads, MultiMET_Facebook, and MultiMET_Twitter; MetMEME is divided into two subsets: MetMEME_Chinese and MetMEME_English. Detailed information about the relevant datasets is provided in Table 1.

[0088] During the experiments, we randomly sampled 1,000 samples from each dataset, consisting of 500 metaphorical data and 500 non-metaphorical data, to ensure that the accuracy and Macro-F1 of the fair coin random baseline were 50% in all cases. To ensure the reliability of the experimental results, the dataset was split into training, validation, and test sets with a ratio of 0.8 / 0.1 / 0.1.

[0089] Table 1 Dataset information statistics

[0090] Dataset Metaphorical data volume Non-metaphorical data volume Total data volume MultiCMET 6,306 7,514 13,820 MultiMET_ads 2,537 1,653 4,190 MultiMET_Facebook 1,066 745 1,811 MultiMET_Twitter 2,555 1,943 4,498 MetMEME_Chinese 2,327 3,718 6,045 MetMEME_English 1,114 2,886 4,000

[0091] This example uses the Accuracy and Macro-F1 metrics for evaluation. Accuracy refers to the ratio of correctly classified samples to the total number of samples, intuitively reflecting the accuracy of the model's predictions across all samples. A higher value indicates better overall model prediction accuracy. Macro-F1 is a metric that comprehensively considers both recall and precision. It calculates and averages the F1 value for each category. A higher value indicates more balanced prediction performance across all categories.

[0092] This example selects three multimodal large language models, namely LLaVA-1.5-7B, Qwen-VL, and mBLIP-BLOOMZ-7B. These models support image input and text input and have been pre-trained on large-scale data. The hidden layer representations of all models are processed by the representation extraction module, adaptive attention pooling module, and hierarchical weighted summation module proposed in this invention, and the metaphor classification task is finally completed by the MLP classifier. The experiment uses Adam as the optimizer, with an initial learning rate of 5e-5, a batch size of 16, and 100 training rounds.

[0093] Table 2 Classifier effects based on training of various multimodal large language models

[0094]

[0095]

[0096] The experimental results in Table 2 show that by learning representations of metaphorical images or combining them with text descriptions and training a classifier based on these representations, the classifier achieves significantly higher accuracy in metaphor detection tasks than the fair coin random baseline (50.00). This result confirms that large multimodal language models can effectively encode metaphorical information. Furthermore, a comparison of experimental results using pure image input and image-text pair input on the same dataset reveals that the classifier trained with a combination of image and text input performs better than the classifier trained with pure image input. This result suggests that the combined input of multiple modalities can enrich the metaphorical representation within large multimodal language models.

[0097] In summary, the edge-detection-based metaphor representation method for a multimodal large language model proposed in this paper can effectively exploit the multimodal large language model's ability to encode metaphorical information through a representation extraction module, an adaptive attention pooling module, a hierarchical weighted summation module, and a classifier module. The experimental results not only verify that the multimodal large language model can effectively encode metaphorical information, but also show that multimodal combined input can enrich the metaphorical representation within the model. This method provides an effective research approach and method for the interpretability of multimodal metaphor detection, is innovative and practical, and is expected to provide important reference and inspiration for subsequent related research and applications.

[0098] Although the present invention has been described above, the present invention is not limited to the above-mentioned specific embodiments. The above-mentioned specific embodiments are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can make many variations without departing from the purpose of the present invention, and these are all protected by the present invention.

Claims

1. A metaphor representation method for a multimodal large language model based on edge detection, characterized in that: The method is based on a representation extraction module, a pooling module, a hierarchical module, and a classifier, and includes the following steps: Offline stage: Input the data containing metaphorical and non-metaphorical labels into the multimodal large language model to extract the hidden representation data h; The hidden representation h is linearly projected into 256-dimensional space according to the following formula to obtain the hidden representation tensor matrix z z=W p h+b p in, is the weight matrix of linear projection; is the bias vector, d is the dimension of the original hidden representation; The hidden representation tensor matrix is pooled and normalized through the adaptive attention mechanism to obtain the hidden pooled representation matrix p (i) ; The hidden weighted representation v is obtained by performing hierarchical weighted summation on the hidden pooling representation matrix through the following formula; Among them, α i Represents the weight of the i-th layer in the multimodal large language model; The following formula is used to predict whether the hidden weighted representation contains metaphor labels in the input data Among them, W c and b c are the weight matrix and bias vector respectively, σ(·) is the Sigmoid activation function; Training phase The cross entropy loss function is used to predict the results The difference between the true label and the real label is calculated, and the parameters of the classifier are continuously adjusted through the back propagation algorithm.

2. The edge detection-based multimodal large language model metaphor representation method according to claim 1, characterized in that: The hidden representation tensor matrix is pooled and normalized through the adaptive attention mechanism to obtain the hidden pooled representation matrix p (i) Process, including: Perform multi-head self-attention processing on the hidden representation tensor matrix z according to the dimension of the input tensor [B, L, T, D]; Where: B is the batch size, L is the number of layers, T is the sequence length, and D is the feature dimension; Perform the maximum pooling operation on the representation after attention processing along the sequence length T: p=max(z1,z2,...,z T ) The pooled representation is l2 normalized according to the following formula to eliminate the influence of feature scale:

3. The edge detection-based multimodal large language model metaphor representation method according to claim 1, characterized in that: The weight α of the i-th layer in the multimodal large language model i for: Where: β i is a trainable parameter.