Noun Metaphor Recognition Method Based on Large Language Model and Graph Neural Network
Through the method based on large language model and graph neural network, the problems of shortage of data resources and insufficient semantic representation in Chinese noun metaphor recognition are solved, and higher recognition accuracy and understanding of complex contexts are achieved.
Patent Information
- Application Number
- CN202411861877.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-12-17
AI Technical Summary
The prior art faces the problems of data resource shortage, insufficient semantic representation and difficult context modeling in the recognition of Chinese noun metaphor, resulting in low recognition accuracy.
The noun metaphor recognition method based on large language model and graph neural network is adopted to generate diversified corpus through data augmentation and distillation, and deep semantic modeling is used to construct semantic graphs based on dependent syntax and PMI information. The graph convolution network is used to update node features, and finally the classification of noun metaphors is completed.
It significantly improves the accuracy of noun metaphor recognition, reduces the dependence on large-scale annotation data, and can effectively capture metaphorical information in complex contexts.
Smart Images

Figure CN119646227B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of natural language processing, and relates to large language models and deep learning natural language processing. Specifically, it relates to a noun metaphor recognition model based on large language models and graph neural networks and the training of the model. Background Art
[0002] The conceptual metaphor theory holds that metaphor is a cognitive process of understanding unknown concepts with known concepts, and its working mechanism is the mapping from the source domain to the target domain of unknown concepts. In the process of human cognition, many complex concepts that are difficult to describe will be encountered. At this time, other known concepts will be used to understand and construct complex unknown concepts through the way of metaphor. For example, in the metaphor "Efficiency is life", the concept of "life" is used to explain the concept of "efficiency", and the attributes of "life" such as "precious" are extended to the concept of "efficiency" through mapping. The basic expression of metaphor mapping is "X is Y", where X represents the unknown concept and Y represents the known concept. When saying "X is Y", that is, when using Y to construct concept X, the conceptual structure of Y is mapped to X. In fact, what X maps is only part of the attributes of Y, and which part of the attributes of Y the mapping is related to is determined by factors such as empirical knowledge, culture, and context.
[0003] For example, the mapping contained in the conceptual metaphor "The lens of a camera is money" is shown in Table 1. The attributes of money are mapped to the lens. That is, according to common sense, the precious attribute should be mapped to the lens to describe the value of the lens.
[0004] Table 1 Examples of Noun Metaphors
[0005]
[0006] Currently, there are few available data resources for metaphor recognition: The shortage of data resources faced by existing Chinese noun metaphor recognition research is mainly manifested in the small scale of the dataset, the great difficulty of annotation, and the lack of data diversity. Due to the complex expression form of metaphor and its dependence on context and cultural background, annotating metaphor data is not only time-consuming and laborious, but also easily affected by subjective factors, resulting in difficulties in ensuring the consistency and accuracy of annotation.
[0007] Currently, the semantic representation of metaphor recognition is insufficient: Existing models generally rely on simple feature engineering, external semantic resources (such as WordNet, FrameNet) and shallow deep learning methods, resulting in overly single semantic representation and lack of in-depth understanding of the complexity, implicit expression and cross-domain mapping of metaphors. These methods are difficult to effectively capture the differences between literal meaning and metaphorical meaning, especially when dealing with fuzzy or polysemous expressions, the accuracy rate is low. For noun metaphors, there is no direct connection between the source domain and the target domain, and the mapping relationship cannot be fully represented.
[0008] Current metaphor recognition context modeling problem: Although some deep learning models (such as Bi-LSTM) attempt to capture the sequential and contextual features of metaphors, their performance is still insufficient when dealing with long-distance dependencies and complex contexts. The meaning of a metaphor is usually closely related to the context and requires an in-depth understanding of the cross-domain mapping relationship between the source domain and the target domain. However, existing models are insufficient in capturing these fine-grained information, especially in cross-domain mapping and multi-level semantic fusion. Summary of the Invention
[0009] In order to solve the problem of accurately identifying noun metaphors, in a first aspect, according to some embodiments of the present application, a method for identifying noun metaphors based on a large language model and a graph neural network is used for a metaphor recognition model. The metaphor recognition model includes a preprocessing model and a graph convolutional model. The recognition method includes the following steps:
[0010] Train the metaphor recognition model;
[0011] Input the text to be processed into the trained preprocessing model, and the preprocessing model outputs a first feature representation;
[0012] Construct a semantic graph of the text to be processed according to syntactic dependency relationships and point mutual information;
[0013] Input the first feature representation and the semantic graph into the trained graph convolutional network, and the graph convolutional network outputs a second feature representation;
[0014] Classify the second feature representation through a classifier to obtain the noun metaphor recognition result in the text to be processed.
[0015] According to some embodiments of the present application, a method for identifying noun metaphors based on a large language model and a graph neural network, wherein, in the step of training the metaphor recognition model, it includes
[0016] Perform data augmentation on a first data set through a first large language model to construct a second data set;
[0017] Perform data distillation on the second data set through a second large language model to construct a third data set;
[0018] Input the third data set into the preprocessing model, and the preprocessing model outputs a first feature representation;
[0019] Construct a semantic graph of the third data set according to syntactic dependency relationships and point mutual information;
[0020] Input the first feature representation and the semantic graph into the graph convolutional network, and the graph convolutional network outputs a second feature representation;
[0021] Classify the second feature representation through a classifier.
[0022] According to the noun metaphor recognition method based on large language models and graph neural networks in some embodiments of the present application, the first data set contains Chinese sentences with noun metaphors or Chinese sentences without noun metaphors, which is represented by the following formula:
[0023] Data Input ={Sentence}
[0024] In the formula, Sentence represents a Chinese sentence containing a noun metaphor or a Chinese sentence without a noun metaphor;
[0025] The prompt for the first large language model in the data augmentation, where the prompt for the sentence Data Input containing a noun metaphor is represented by the following formula:
[0026] Sentences metaphor =Qwen(Data Input +"Generate five similar sentences including noun metaphors?")
[0027] In the formula, Sentences metaphor represents the augmented sentence of the sentence Data Input containing a noun metaphor;
[0028] The prompt for the first large language model in the data augmentation, where the prompt for the sentence Data Input without a noun metaphor is represented by the following formula:
[0029] Sentences non-metaphor =Qwen(Data Input +"Generate five similar sentences without noun metaphors?")
[0030] In the formula, Sentences non-metaphor represents the augmented sentence of the sentence Data Input without a noun metaphor;
[0031] The second data set includes Sentences metaphor and Sentences non-metaphor ;
[0032] The prompt for the second large language model in the data distillation, where the prompt for the sentence Data Input containing a noun metaphor and its augmented sentence Sentences metaphor is expressed as:
[0033] Verificationmetaphor = ChatGLM(Sentences metaphor ,"Is it all noun metaphor sentences?")
[0034] The prompt for the second large language model in the data distillation, where for sentences that do not contain noun metaphors Data Input and its augmented sentences Sentences non-metaphor The prompt is represented by the following formula:
[0035] Verification non-metaphor = ChatGLM(Sentences non-metaphor ,"Are none of them noun metaphor sentences?")
[0036] If the answer to the prompt is yes, then the third data set is represented by the following formula:
[0037] Final data = Verification metaphor + Verification non-metaphor
[0038] If the answer to the prompt is no, then delete the sentences that are not noun metaphors from the sentences Data Input and its augmented sentences Sentences metaphor The deletion is represented by the following formula:
[0039] Filtered metaphor = ChatGLM(Sentences metaphor ,"Delete sentences that are not noun metaphors")
[0040] If the answer to the prompt is no, then delete the sentences that are noun metaphors from the sentences Data Input and its augmented sentences Sentences non-metaphor The deletion is represented by the following formula:
[0041] Filtered non-metaphor = ChatGLM(Sentences non-metaphor ,"Delete sentences that are noun metaphors")
[0042] Then the third data set is represented by the following formula:
[0043] Final data = Filtered metaphor + Filtered non-metaphor
[0044] In the formula, Finaldata The third data set;
[0045] Preferably, the first large language model includes the Qwen model, the second large language model includes the ChatGLM model, and the preprocessing model includes the ERNIE model.
[0046] According to the noun metaphor recognition method based on large language models and graph neural networks according to some embodiments of the present application, wherein, in the step of inputting the third data set into the preprocessing model, the preprocessing model outputs a first feature representation, including the following steps:
[0047] S310. Tokenize the sentences in the third data set Final data The words after tokenization are {w 1 , w 2 , …, w n}, and each word generates a pre-trained word embedding E(w i ), which is represented by the following formula:
[0048] E(w i ) = Embedding(w i )
[0049] S320. Perform positional encoding on the words, where the positional encoding is represented as P(i), which is represented by the following formula:
[0050]
[0051] In the formula, i is the position information of the word, j is the dimension of the positional encoding, and d is the embedding dimension;
[0052] S330. Obtain the weighted sum of the values of each word through the attention mechanism for the word vector H 0 (i) and perform weighted aggregation of the context features. Concatenate the outputs of each attention head and map them back to the dimension of the model through a linear transformation to obtain the output of the multi-head attention mechanism MultiHead(Q, K, V);
[0053] Among them, the word vector H 0 (i) of each word is represented by the following formula:
[0054] H 0 (i) = E(w i ) + P(i);
[0055] S340. Normalize the output of each layer through layer normalization, which is represented by the following formula:
[0056] H L = LayerNorm(H 0 (i) + MultiHead(Q, K, V))
[0057] S350. Output the first feature representation, which is represented by the following formula:
[0058] H final = H L
[0059] In the formula, H final represents the first feature representation.
[0060] According to the noun metaphor recognition method based on the large language model and the graph neural network according to some embodiments of the present application, wherein, in step S330, the output of the multi-head attention mechanism MultiHead(Q, K, V) includes the following steps:
[0061] S331. For the input word vector H 0 (i) Use different W Q , W K , W V to generate multiple Q, K, V, and perform attention calculation independently for each head to obtain Attention 1 , Attention 2 ,.....Attention h ;
[0062] S332. Concatenate the outputs of each head, and map them back to the dimension of the model through a linear transformation W O , which is represented by the following formula:
[0063] MultiHead(Q, K, V) = Concat(Attention 1 , Attention 2 ,..., Attention h )W O
[0064] In the formula, MultiHead(Q, K, V) represents, represents the linear transformation matrix, d model represents the dimension of the model, h represents the number of heads, and d v represents the dimension of a single head;
[0065] Among them, calculating Attention in step S331 includes the following steps:
[0066] S3311. Let the word vector H 0 (i) enter the attention mechanism, and generate query (Query), key (Key), and value (Value) matrices from the word vector H 0 (i) through a linear transformation, which is represented by the following formula:
[0067] Q = H 0 (i)W Q , K = H 0 (i)W K , V = H 0 (i)W V
[0068] wherein respectively represent the query projection matrix, the key projection matrix, and the value projection matrix; respectively represent the query matrix, the key matrix, and the value matrix;
[0069] S3312. Calculate the similarity score Score(i, j) between the query matrix Q and the key matrix K, which is represented by the following formula:
[0070]
[0071] wherein, Q i represents the query vector of the i-th word, K j represents the key vector of the j-th word, the similarity score Score(i, j) represents the relationship strength between positions i and j, and the superscript (t) represents the calculation of the t-th head in the multi-head attention mechanism;
[0072] S3313. Integrate the scores of all positions into a matrix form, which is represented by the following formula:
[0073]
[0074] wherein represents the similarity score of each position in the sentence to all other positions, and n represents the number of words;
[0075] S3314. Scale the similarity score, which is represented by the following formula:
[0076]
[0077] S3315. Calculate the weighted sum of the values of each word, which is represented by the following formula:
[0078]
[0079] wherein, the numerator represents converting the correlation score of the i-th position to the j-th position into a positive value, and the denominator represents the sum of all correlation scores in the i-th row for normalization;
[0080] S3316. Use the weighted sum a of the values of each word ij to weight the value matrix V j , and perform weighted aggregation of context features, which is represented by the following formula:
[0081]
[0082] In the formula, n represents the number of words.
[0083] According to some embodiments of the present application, a method for identifying noun metaphors based on a large language model and a graph neural network, wherein, in the steps, a semantic graph of the third data set is constructed according to syntactic dependency relationships and point mutual information, including the following steps:
[0084] S410. Construct a vocabulary for the sentences in the third data set Final data including
[0085] S411. Represent the set of sentences in the third data set Final 1 , s 2 , …, s N} by S = {s data}, where each sentence s i contains a set of vocabulary;
[0086] S412. Split each sentence s i into a series of vocabulary s i = {w i1 , w i2 , …, w ij}, where w ij represents the j-th word in the sentence s i ;
[0087] S420. Perform a union operation on the set of vocabulary in all sentences to obtain a vocabulary, which is represented by the following formula:
[0088]
[0089] In the formula, vocab represents the vocabulary, U represents the union operation, w ij represents the j-th word in the sentence s i , and N represents the number of sentences;
[0090] S430. For each word w k ∈ vocab, represent the number of occurrences of the word w k in all documents by the word frequency word_freq(w k ), traverse all documents, and count the frequency of each word. The frequencies of each word in all documents are accumulated and recorded in the word frequency dictionary word_freq, which is represented by the following formula:
[0091]
[0092] In the formula, is an indicator function representing the result of a conditional judgment. In the formula, if the word w k appears in the sentence S i , then otherwise it is 0;
[0093] S440. Define a context window of a fixed size. By traversing the words in the sentence, calculate the co-occurrence times of word pairs, including
[0094] S441. The sequence of words in sentence Si is Si = {w i1 , w i2 , …, w ij}, where j represents the number of words, the window size is window_size. For each word w i in the sentence S ij , define its context window as:
[0095] window(w ij ) = {w ij-k , w ij-k+1 ,..., w ij+k}
[0096] In the formula, k is an integer from -window_size to window_size;
[0097] S442. For the number of times each word pair (w p , w q ) appears in the context window, by maintaining a count dictionary word_pair_count, record the number of times each pair of words w p and w q appears in the corpus, which is represented by the following formula:
[0098] word_pair_count(w p , w q ) = word_pair_count(w p , w q ) + 1
[0099] In the formula, word_pair_count() represents the counting function, and w p and w q are any pair of different words in the window;
[0100] S450. Calculate the weight of the edge between two points, where any two points are represented by any two of the said words. If there is no syntactic dependency between the two points, use the PMI value as the weight. If there is a syntactic dependency, use the PMI value plus 10 as the weight;
[0101] Among them, the PMI is calculated as follows:
[0102]
[0103] Wherein, word_pair_count(w i , w j ) is the number of times the word pair (w i , w j ) appears in the document, and N represents the number of sentences, where:
[0104]
[0105] Among them, if there is a syntactic dependency relationship, the PMI value plus 10 is used as the weight, which is represented by the following formula:
[0106] PMI(w i , w j ) = PMI(w i , w j ) + 10, if (w i , w j ) ∈ special_word_pairs
[0107] Wherein, special_word_pairs is a set of syntactic dependency relationships between words obtained after calculating the syntactic dependency relationships by spacy.
[0108] According to the noun metaphor recognition method based on a large language model and a graph neural network according to some embodiments of the present application, wherein, in the step, the first feature representation and the semantic graph are input into a graph convolutional network, and the graph convolutional network outputs a second feature representation, including using the update formula of the graph convolutional network GCN to update and optimize the first feature representation through the semantic graph to aggregate neighbor features and output the second feature representation, wherein the update formula is represented by the following formula:
[0109]
[0110] Wherein, H (L) is the node feature representation of the L-th layer. Initially, L = 0, and the output of the preprocessing model is used as H (0) ; N(i) is the set of neighbor nodes of node i, and each node represents a word; is the adjacency matrix of the semantic graph, including self-connections, is the degree matrix of the nodes; is the Laplacian normalization operation of the semantic graph; W (I) is the trainable parameter of the I-th layer; σ() is the activation function; represents the second feature representation.
[0111] According to the noun metaphor recognition method based on large language models and graph neural networks in some embodiments of the present application, the second feature representation is classified through a linear layer W c and a bias term b c to output a scalar y pred , representing the probability that the sample belongs to a certain category.
[0112]
[0113] In the formula, y pred is the predicted output of the model, which represents the probability that the input sentence belongs to a noun metaphor, and sigmoid() represents the sigmoid function;
[0114] The loss function for training the metaphor recognition model is shown by the following formula:
[0115]
[0116] In the formula, LOSS all represents the total loss, and γ is a hyperparameter;
[0117] Among them:
[0118] Loss Ernie =-(y true log(y pred )+(1 - y true )log(1 - y pred ))
[0119] In the formula, Loss Ernie represents the loss of the preprocessing model, y true represents the true label, and y pred represents the predicted output of the model;
[0120]
[0121] In the formula, represents the loss of the graph convolutional network, C represents the number of categories, and the superscript (c) represents the number of the current category.
[0122] In a second aspect, embodiments of the present application further provide an electronic device, which includes: one or more processors, a memory, and one or more programs; wherein, the one or more programs are stored in the memory, and the one or more programs include instructions that, when executed by the electronic device, cause the electronic device to execute the first aspect and any possible technical solution of the first aspect.
[0123] In a third aspect, an embodiment of the present application further provides a computer-readable storage medium, where the computer-readable storage medium includes a computer program. When the computer program runs on an electronic device, the electronic device is caused to execute the first aspect and any possible technical solution of the first aspect.
[0124] Beneficial effects: The method for noun metaphor recognition based on a large language model and a graph neural network (GCN) of the present invention generates diverse corpora and verifies their effectiveness through data augmentation and distillation, reducing the dependence on large-scale labeled data. A graph is established using dependency syntax and pointwise mutual information (PMI), and the GCN learns the syntactic structure information in the sentence to improve the accuracy of metaphor recognition. At the same time, the ERNIE model captures the cross-domain mapping relationship between nouns through deep semantic modeling, and combines the graph structure information of the graph convolutional network to update the node features, and finally completes noun metaphor classification. Experimental results show that this method significantly improves the accuracy of metaphor recognition and has broad application potential. Specifically:
[0125] (1) Data augmentation and distillation: First, the large language model Qwen-max is used to augment the dataset to generate diverse corpora. Subsequently, the augmented data is distilled by the large language model ChatGLM to extract high-quality training data. This process effectively reduces the dependence on large-scale labeled data, enabling efficient fine-tuning of the pre-trained model even in the case of scarce data.
[0126] (2) Training based on Ernie and cross-domain mapping: The augmented and distilled dataset is finally input into the Ernie model for training. Ernie can efficiently process long texts and synthesize global context through the parallelization feature of its Transformer architecture, thereby accurately identifying metaphors in complex contexts. At the same time, the recognition of metaphors not only depends on the literal meaning but also involves the cross-domain mapping between the source domain and the target domain. Ernie can capture this cross-domain relationship and combine external knowledge (such as knowledge graphs) to assist in judging the semantic connection between the source domain and the target domain. In addition, metaphor understanding also involves multi-level semantic features, such as syntactic structure, emotional color, and cultural background. Through deep semantic modeling, Ernie effectively integrates this information, significantly improving the accuracy and robustness of metaphor recognition.
[0127] (3) Graph Convolutional Network Based on Dependency Syntax and PMI Modeling: In the present invention, the dependency syntax and PMI information of a sentence are extracted as the weights between nodes, and corresponding graphs between sentences and words are constructed respectively. The dependency syntax reflects the direct grammatical relationship between words, and PMI reflects the statistic of the association strength between two words. The nodes are represented by the output of Ernie, and the graph convolutional neural network (GCN) is used to learn these graph structures, so as to achieve effective information fusion. Through the graph convolutional network, the model can fully mine different levels of grammatical information in the sentence, comprehensively understand the structure and semantic relationship of the sentence, and improve the accuracy of metaphor recognition.
[0128] Additional aspects and advantages of the present invention will be given in part in the following description, will become apparent in part from the following description, or will be understood through the practice of the present invention. Brief Description of the Drawings
[0129] Figure 1 It is the technical roadmap in the embodiment.
[0130] Figure 2 It is the noun metaphor recognition model based on the large language model and graph convolutional network in the embodiment
[0131] Figure 3 It is the data augmentation - noun metaphor example in the embodiment.
[0132] Figure 4 It is the data augmentation - non - noun metaphor - non - metaphor example in the embodiment.
[0133] Figure 5 It is the data augmentation - non - noun metaphor - verb metaphor example in the embodiment. Detailed Description of the Specific Embodiment
[0134] The embodiments of the present application will be described in detail below with reference to the accompanying drawings. The examples of the embodiments are shown in the drawings. The present application provides a method and an electronic device. Among them, the method and the device are based on the same technical concept. Since the principles of solving problems by the method and the device are similar, the implementation of the device and the method can be referred to each other, and the repeated parts will not be described again.
[0135] Term Explanation:
[0136] Noun Metaphor: A noun metaphor is a rhetorical device in which a noun is used to represent another concept that is literally different from it, thereby creating an implicit and non - literal meaning. Through the hint or association of the context, such a metaphor enables a noun to go beyond its conventional, direct physical or concrete meaning and instead carry more abstract or symbolic content.
[0137] Data augmentation: Data augmentation refers to the process of generating new data samples by performing certain transformations on the original data. In the present invention, the original noun metaphor sentences or non-noun metaphor sentences are used to generate sentences with the same structure, aiming to expand the training data set, enhance the generalization ability of the model, and reduce overfitting.
[0138] Data distillation: Data distillation refers to filtering the existing data to obtain effective data to improve the data quality.
[0139] Word segmentation: Word segmentation is the process of splitting a continuous text sequence into the smallest language units (such as words, sub-words, characters, etc.) in natural language processing, laying the foundation for subsequent text analysis and modeling.
[0140] Word embedding: Word embedding is a technique that maps words into a high-dimensional dense vector space, making words with similar semantics closer in this space.
[0141] ERNIE: ERNIE (Enhanced Representation through Knowledge Integration) is a pre-trained language model based on knowledge enhancement proposed by Baidu. By masking Chinese word segmentation, it optimizes the representation learning of language.
[0142] Graph convolutional network: A graph convolutional network is a deep learning model for processing graph data. It aggregates information within the neighborhood of graph nodes to learn the representation of nodes or graphs.
[0143] Multi-head attention mechanism: The multi-head attention mechanism is an important part of the Transformer model. It splits the input attention calculation into multiple "heads", each head independently calculates and learns different attention representations, and finally merges the outputs of each head. This mechanism helps the model capture the dependency relationships of different aspects of the input.
[0144] Dependency syntax graph: A dependency syntax graph (Dependency Syntax Tree) is a graph representing the sentence structure, where nodes represent words and edges represent the dependency relationships between words. By analyzing the dependency relationships, the grammatical roles and mutual relationships of each word in the sentence can be revealed.
[0145] Feature fusion: Feature fusion (Feature Fusion) refers to merging features from different sources to construct a richer and more expressive input representation.
[0146] Softmax: Softmax is an activation function commonly used in classification tasks. It converts each element in a vector into a probability value between 0 and 1, and the sum of all elements is 1. Softmax is usually used in the output layer of multi-classification problems.
[0147] Binary classification: Binary classification refers to a classification task of dividing input data into two categories. In the present invention, it refers to noun metaphors and non-noun metaphors.
[0148] The method of the present invention aims to effectively identify metaphorical expressions with nouns as the core in text, provide support for improving the machine's understanding ability of complex semantics in natural language, and then be applied to various natural language processing tasks, including fields such as sentiment analysis, text generation, question answering systems, and intelligent dialogue systems.
[0149] In the first embodiment, the core of the noun metaphor recognition task lies in distinguishing whether there is a metaphorical relationship in a sentence and how to capture the non-literal semantic association between nouns through a model. As Figure 1 shown, the method for noun metaphor recognition based on a large language model and a graph neural network of the present invention is used for noun metaphor recognition, and this method mainly includes the following steps:
[0150] Construct and train a deep learning network model;
[0151] Use the trained model for recognition.
[0152] Among them, the input of the model is a sentence like "The lens of the camera is money", and the output of the model is a binary classification of noun metaphor or non-noun metaphor. The present invention mainly optimizes the model framework to better adapt to this specific task.
[0153] Among them, the model training method includes the following steps:
[0154] Step 1: Obtain input data
[0155] Among them, the input data is Chinese sentences containing noun metaphors or non-metaphors. The dataset used in the present invention comes from the CCL2018 Chinese metaphor detection task evaluation data, which contains a total of 4,394 Chinese sentences, covering various types of metaphorical sentences. Since the algorithm proposed in the present invention mainly focuses on identifying noun metaphors, the model will focus on distinguishing noun metaphor sentences and non-metaphor sentences during training and evaluation to improve the detection accuracy of noun metaphors.
[0156] Step 2: Data augmentation
[0157] Among them, in the noun metaphor recognition task, data augmentation is performed through the Qwen-max model to expand the data volume. Specifically, for the noun metaphor sentences in the dataset, the augmentation prompt is: "Generate a similar sentence containing a noun metaphor", as Figure 3 shown. For the non-noun metaphor sentences in the dataset, the prompt "Generate a similar sentence without a noun metaphor" is used, as Figure 4and Figure 5 As shown. The sentences generated during the augmentation process will be passed as training data to subsequent models.
[0158] Step 3: Data Distillation
[0159] Among them, it is used to verify the effectiveness of the sentences after data augmentation using ChatGLM. Specifically, the data from Step 2 is input into ChatGLM. Among them, for the noun metaphor sentences in the dataset and the sentences augmented according to the noun metaphor sentences, the prompt is used: "Are all of them metaphor sentences?" Among them, for the non-noun metaphor sentences in the dataset and the sentences augmented according to the non-noun metaphor sentences, the prompt is used: "Are none of them metaphor sentences?"
[0160] If the answer is "yes", it indicates that all the sentences augmented from the noun metaphor sentences in Step 2 are noun metaphor sentences, and all the sentences augmented from the non-noun metaphor sentences are non-noun metaphor sentences, and the data directly enters Step 4. If the answer is "no", it indicates that there are non-noun metaphor sentences among the sentences augmented from the noun metaphor sentences and / or there are noun metaphor sentences among the sentences augmented from the non-noun metaphor sentences in Step 2. For the non-noun metaphor sentences among the sentences augmented from the noun metaphor sentences, the prompt is used: "Delete the sentences that are not noun metaphors in the augmented sentences", and the non-noun metaphor sentences among the sentences augmented from the noun metaphor sentences are deleted. For the noun metaphor sentences among the sentences augmented from the non-noun metaphor sentences, the prompt is used: "Delete the sentences that are metaphors in the augmented sentences", and the noun metaphor sentences among the sentences augmented from the non-noun metaphor sentences are deleted, and then it enters Step 4.
[0161] Step 4: ERNIE Model Training
[0162] Among them, during the training process of the ERNIE model, first, the input data obtained through data distillation is tokenized and encoded, the sentence is disassembled into word or sub-word units, and a pre-trained word embedding is generated for each word. Then, by combining the word vectors with positional encodings, the model can capture the sequential relationship of words in the sentence. Then, through the multi-head attention mechanism, the representation of each word learns context information in multiple parallel attention heads, capturing long-range semantic dependencies. Through the self-attention mechanism, the model can adjust the word vector representation. Especially in metaphor sentences, ERNIE pays special attention to the semantic changes of nouns. After multi-layer attention mechanisms and context learning, the model finally obtains the vector representation of the entire sentence. To improve the semantic understanding of nouns, especially for the task of identifying noun metaphors, the model adjusts the embedding representation of nouns and outputs (the first feature representation). Through the cross-entropy loss function, the model optimizes the parameters for training by comparing the difference between the predicted probability values and the true labels.
[0163] Step 5: Modeling the Semantic Graph Based on Dependency Syntax and PMI
[0164] Among them, in the semantic graph modeling process based on dependency syntax and PMI, first, the distilled dataset is processed to construct a vocabulary and a word frequency dictionary. Specifically, each sentence in the sentence set is split into words, and the entire vocabulary is obtained through the union operation. The frequency of each word is calculated by traversing all documents and counting the number of times the word appears. Then, a word pair co-occurrence graph is constructed. By setting a context window of a fixed size, the co-occurrence times of word pairs are calculated, and a count dictionary is maintained to record the co-occurrence frequency of each pair of words. On this basis, the edge weights between word pairs are calculated: if there is a syntactic dependency relationship between two words, 10 is added to the PMI calculation result to increase its weight. Finally, the obtained word pair co-occurrence graph and dependency relationship information help to better capture the semantic associations between words and provide richer context information for the model.
[0165] Step Six: Training of the Graph Neural Network
[0166] The node feature representation (the first feature representation) output by ERNIE in Step Four and the semantic graph output in Step Five are further processed through a Graph Convolutional Network (GCN). GCN updates the features of each node through the relationships between nodes and edges in the graph structure, fusing information layer by layer, so as to capture the structural features of the graph. Finally, the model combines the output of the ERNIE classifier and the graph convolution output of GCN to provide a comprehensive feature representation (the second feature representation) for the final classification decision of the nodes.
[0167] Step Seven: Calculation of the Loss Value
[0168] Among them, the final loss value of the model is composed of the weighted combination of the losses of the ERNIE part and the graph convolutional neural network part. By weighted summing the losses of these two parts, the model can simultaneously optimize the learning of text features through the ERNIE model and the learning of graph structure information through the graph convolutional network during the training process.
[0169] In the second embodiment, specifically, the noun metaphor recognition of the present invention based on the large language model and the graph neural network includes the following steps:
[0170] Step One: Input Data
[0171] The input data is a Chinese sentence containing a noun metaphor or a non-noun metaphor. The formula is as follows:
[0172] Data Input ={Sentence}
[0173] Among them, Sentence represents the Chinese sentence containing a noun metaphor or the Chinese sentence including a non-noun metaphor (the Chinese sentence without a noun metaphor).
[0174] Step 2: Data Augmentation
[0175] In the noun metaphor recognition task, data augmentation not only needs to increase the amount of data, but also consider how to generate more sentences containing metaphors and non-metaphors. The present invention uses the Qwen-max model for data augmentation. For a sentence containing a noun metaphor, the augmentation prompt is: "[Input sentence] + Generate five similar sentences containing noun metaphors". For a sentence containing a non-noun metaphor, the prompt used is: "[Input sentence] + Generate five similar sentences without noun metaphors". For a sentence containing a noun metaphor:
[0176] Sentences metaphor = Qwen(Data Input + "Generate five similar sentences including noun metaphors?")
[0177] where Qwen represents Qwen-Max, and Sentences metaphor represents the augmented sentences of the sentence containing noun metaphors Data Input
[0178] For a sentence without a noun metaphor:
[0179] Sentences non-metaphor = Qwen(Data Input + "Generate five similar sentences without noun metaphors?")
[0180] where Sentences non-metaphor represents the augmented sentences of the sentence without noun metaphors Data Input
[0181] After augmentation, the quantity of the dataset is shown in Table 2:
[0182] Table 2 Augmented Data
[0183] Category Quantity Noun Metaphor 7626 Non-Noun Metaphor 8829 Verb Metaphor 7654 Non-Metaphor 1175 Total 16455
[0184] Step 3: Data Distillation
[0185] In this step, ChatGLM is used to verify the validity of the augmented sentences. Specifically, the original sentences in Step 2 and the data of the augmented sentences generated from the original sentences are input to ChatGLM. For the original sentences that include noun metaphors, the prompt is: "Are all the augmented sentences of this original sentence noun metaphor sentences?" For the original sentences that do not include noun metaphors, the prompt is: "Are all the augmented sentences of this original sentence not noun metaphor sentences?" If the answer is "yes", the data directly enters Step 4. If the answer is "no", then the prompts "Delete the sentences that are not noun metaphors in the augmented sentences of the original sentence" and "Delete the sentences that are noun metaphors in the augmented sentences of the original sentence" are used to remove the non-metaphor sentences from the augmented sentences of the noun metaphor original sentences and the metaphor sentences from the augmented sentences of the non-noun metaphor original sentences, and then enter Step 4.
[0186] Suppose the set of augmented sentences is input to ChatGLM and verified using the prompts, as shown below.
[0187] Verification = ChatGLM(Sentences metaphor ,"Are all of them metaphor sentences?")
[0188] Verification = ChatGLM(Sentences non-metaphor ,"Are all of them not metaphor sentences?")
[0189] If ChatGLM's answer is "yes", it means that all the augmented sentences of the original noun metaphor sentences are noun metaphor sentences, or all of the original non-noun metaphor sentences are not noun metaphor sentences. Then the formula is as follows:
[0190] Final data = Verification
[0191] That is, Final data is the union of the two Verifications.
[0192] If ChatGLM's answer is "no", it means that the augmented sentences of the original noun metaphor sentences contain non-noun metaphor sentences, and the augmented sentences of the original non-noun metaphor sentences contain noun metaphor sentences, and further screening is needed. At this time, the present invention removes the non-noun metaphor sentences contained in the augmented sentences of the original noun metaphor sentences and the noun metaphor sentences contained in the augmented sentences of the original non-noun metaphor sentences. The specific operation is as follows:
[0193] Filtered = ChatGLM(Verification, "Delete the sentences that are not metaphors")
[0194] Filtered = ChatGLM(Verification, "Delete metaphorical sentences")
[0195] Final effective data: The formula for the final effective data is:
[0196] Final data = Filtered
[0197] That is, Final data is the union of two Filtereds. Step 4: ERNIE model training
[0198] During the training process of the ERNIE model, the data obtained after data distillation in Step 3 is first subjected to tokenization and encoding. Tokenization breaks the sentence into word or sub-word units and performs word embedding using pre-trained word vectors. Next, by adding positional encoding to each word vector, the model can capture the sequential relationship of words in the sentence. Then, the processed word vectors enter the multi-head attention mechanism, and multiple parallel attention heads learn context information from different subspaces. Through the self-attention mechanism, the model can understand the relationship between each word and other words and adjust the word vector representation in multiple iterations to capture long-range semantic dependencies. In metaphorical sentences, ERNIE pays special attention to the semantic changes of nouns and continuously adjusts the word embeddings through context information to obtain a more accurate semantic representation. After passing through multiple layers of attention mechanisms and context learning, the model finally obtains the vector representation of the entire sentence.
[0199] For the input sentence Final data , first perform tokenization and encoding. Let the word or sub-word units after tokenization be {w 1 , w 2 , …, w n}, and generate pre-trained word embeddings E(w i ) for each word:
[0200] E(w i ) = Embedding(w i )
[0201] Then, combine the word embeddings with positional encoding to capture the sequential information of words in the sentence. The positional encoding can be represented as P(i), where i is the positional information of the word:
[0202]
[0203] where i is the positional information of the word, j is the dimension of the positional encoding, and d is the embedding dimension. Therefore, the final representation of each word:
[0204] H 0 (i) = E(wi ) + P(i)
[0205] Word vector H 0 (i) Enter the multi-head attention mechanism, and each attention head can learn context information in different subspaces. Next, the word vector H will be specifically described. 0 (i) Steps for the word vector H to enter the attention mechanism calculation:
[0206] First, from H 0 (i), three matrices, namely query (Query), key (Key), and value (Value), are generated through linear transformation:
[0207] Q = H 0 (i)W Q , K = H 0 (i)W K , V = H 0 (i)W V
[0208] is the projection matrix. represent the query matrix, key matrix, and value matrix respectively. d model represents the dimension of the model, h represents the number of heads, and d k represents the dimension of a single head.
[0209] In the self-attention mechanism, the goal is to calculate the attention distribution of each word to all other words in the sequence. First, calculate the similarity score between the query Q and the key K:
[0210]
[0211] Q i : The query vector of the i-th word. K j : The key vector of the j-th word. Score(i, j): Represents the relationship strength between position i and position j.
[0212] Then, the scores at all positions are integrated into a matrix form as:
[0213]
[0214] where: represents the correlation score of each position in the sentence to all other positions, and n represents the number of words.
[0215] To prevent the score value from being too large (which may lead to gradient disappearance or explosion), the score will be scaled to:
[0216]
[0217] Then, calculate the weighted sum of the values of each word:
[0218]
[0219] Among them, the numerator Converts the relevance score of the i-th position to the j-th position into a positive value for easy normalization. The denominator Calculates the sum of all relevance scores in the i-th row for normalization.
[0220] Finally, the attention weight a ij Is used to weight the value matrix V j To complete the weighted aggregation of context features:
[0221]
[0222] Among them, for the input H 0 (i) Generates multiple QKVs using different WQWKWV, then independently performs the above attention calculation for each head to obtain Attention1, Attention2,.....Attentionh. Finally, the outputs of each head are concatenated and mapped back to the dimension of the model through the linear transformation WO.
[0223] MultiHead(Q, K, V) = Concat(Attention 1 , Attention 2 ,..., Attention h )W O
[0224] Among them, Is the final linear transformation matrix. The shape of the output is (n, d model ).
[0225] Through the above self-attention mechanism, the model can understand the relationship between each word and other words, and adjust the representation of each word according to the context information. These adjusted word representations are the new output H L .
[0226] The output of each layer is normalized through layer normalization (LayerNormalization):
[0227] H L = LayerNorm(H 0 (i) + MultiHead(Q, K, V))
[0228] Among them, H 0 Is the input word representation. In the noun metaphor recognition task, ERNIE pays special attention to the semantic changes of nouns (such as the ontology words in the sentence) in the context, so it will assign a greater weight to nouns.
[0229] The first feature representation H of the final ERNIE model final = H L .
[0230] In the training of the ERNIE model, the cross-entropy loss function is adopted. This function is usually used to measure the gap between the probability distribution of the model output and the true label. Let y true be the true label (0 or 1, representing non-noun metaphor and noun metaphor respectively), and y pred be the predicted output of the model (the probability value after sigmoid activation). Then the cross-entropy loss function can be expressed as:
[0231] Loss Ernie = -(y true log(y pred ) + (1 - y true )log(1 - y pred ))
[0232] Step Five: Semantic graph modeling based on dependency syntax and PMI
[0233] Among them, the dataset after data distillation in Step Three is processed. It is divided into a sentence list and sentence content. The sentence list has each data as a sentence, and the construction of the sentence vocabulary. Assume S = {s 1 , s 2 , …, s N} is the sentence set, where each sentence s i contains a set of vocabulary.
[0234] Each sentence si is split into a series of vocabulary si = {w i1 , w i2 , …, w ij}, where w ij represents the j-th word in the sentence s i .
[0235]
[0236] Among them, ∪ represents the union operation on the vocabulary sets in all sentences to obtain the vocabulary list, vocab represents the vocabulary list, U represents the union operation, w ij represents the j-th word in the sentence s i , and N represents the number of sentences.
[0237] For each word w k ∈vocab, the word frequency word_freq(w k ) represents the number of occurrences of the word w k in all documents and can be expressed as:
[0238]
[0239] Among them, is an indicator function, which means that if the word wk appears in the sentence S i , then otherwise it is 0. Traverse all documents and count the frequency of each word. Finally, the frequencies of each word in all documents will be accumulated and recorded in the word frequency dictionary word_freq.
[0240] Next, construct a word pair co-occurrence graph. To calculate the co-occurrence frequency of word pairs, the present invention defines a context window of a fixed size (window_size), and calculates the co-occurrence times of word pairs by traversing the words in the sentence.
[0241] Suppose the vocabulary sequence in the sentence Si is S i = {w i1 , w i2 , …, w ij}, and the window size is window_size. The present invention will traverse the words in the sentence and construct word pairs based on the context window. For each word w i in the sentence S ij , the present invention defines its context window as:
[0242] window(w ij ) = {w ij-k , w ij-k+1 ,..., w ij+k}
[0243] where k is an integer from -window_size to window_size, that is, take window_size words on both the left and right of w ij as its context.
[0244] For the number of times each word pair (w p , w q ) appears in the context window, the present invention needs to maintain a count dictionary word_pair_count to record the number of times each pair of words w p and w q appears in the corpus.
[0245] word_pair_count(w p , w q ) = word_pair_count(w p , w q ) + 1
[0246] where w p and w qis any pair of different words in the window, and word_pair_count() represents the counting function.
[0247] Next, calculate the weight of the edge between two points. If there is no syntactic dependency relationship between the two points, the PMI value is used. If there is a syntactic dependency relationship, 10 is added to the PMI result to increase the weight. The PMI calculation is as follows.
[0248]
[0249] Among them, count(wi, wj) is the number of times the word pair (wi, wj) appears in the document.
[0250] Among them,
[0251] If there is a syntactic dependency relationship between two words, the weight value is increased by 10. The formula is as follows.
[0252] PMI(w i , w j ) = PMI(w i , w j ) + 10 if (w i , w j ) ∈ special_word_pairs
[0253] special_word_pairs is the set of syntactic dependency relationships between words obtained after calculating the syntactic dependency relationships by spacy.
[0254] Step Six: Training of the Graph Convolutional Neural Network
[0255] In this step, the node representation in Step Four and the graph structure in Step Five are combined, and the neighbor features are aggregated through the update formula of GCN to generate a new node representation. This update formula further updates and optimizes the features obtained in Step Four by using the graph information (neighbor set, degree normalization) constructed in Step Five. Simply put, the features in Step Four are input into the graph neural network to obtain the node representation, and at the same time, the graph constructed in Step Five is used for training. The formula for node update is as follows
[0256]
[0257] In the formula, H (L) is the node feature representation of the L-th layer. Initially, L = 0, and H (0) output by the preprocessing model is used. N(i) is the set of neighbor nodes of node i, and each node represents a word. is the adjacency matrix of the semantic graph, including self-connections (plus the identity matrix I), is the degree matrix of the nodes, which is calculated through the adjacency matrix. is the Laplacian normalization operation of the graph, which is used to smooth and standardize the graph structure data. W (I) is the trainable parameter of the I-th layer, which is used to perform a linear transformation on the features.
[0258] σ(·) is the activation function, and RELU is used in the present invention.
[0259] The process is as follows:
[0260] 1. Adjacency matrix normalization: Through perform the smoothing operation of the graph.
[0261] 2. Feature transformation: Utilize the parameter matrix W (I) to perform a linear transformation on the node feature H (L)
[0262] 3. Activation function: Introduce a non-linear transformation through σ(·), and finally obtain the node feature of the L+1-th layer
[0263] For the feature representation Use sigmod to classify through a linear layer Wc and a bias term bc, and finally output a scalar ypred, which represents the probability that the sample belongs to a certain category. Use cross-entropy to calculate the loss value, and the graph convolutional neural network also uses the cross-entropy loss function to calculate the cross-entropy loss LGCN between the model prediction value ypred and the true label ytrue:
[0264]
[0265] Step Seven: Loss calculation
[0266] The final loss value is the sum of the ERNIE part and the graph convolutional neural network part:
[0267]
[0268] where γ is a hyperparameter, a constant.
[0269] In the experimental examples of the present invention, the experimental parameters and environmental settings are as follows: This experiment was carried out for training and inference on two NVIDIA RTX 4090 servers, each equipped with 24GB of video memory. max_length(128) specifies the maximum length of the input text to ensure that the length of each input sequence is consistent. batch_size(32) determines the number of samples processed during each training session, affecting the training speed and memory usage. nb_epochs(60) sets the total number of training epochs, controlling the complete iteration times of model training. bert_lr(1e-4) sets the learning rate of the BERT model, affecting the step size of each parameter update. m(0.7) is used to weight the contribution ratio of the outputs of different modules (such as BERT and GCN) of the model. gcn_layers(2) represents the number of layers of the graph convolutional network, determining the capture depth of graph structure information. n_hidden(200) determines the number of hidden units in each layer, affecting the capacity of the model. dropout(0.5) is used for regularization training, randomly discarding 50% of the neurons to prevent overfitting. The parameter settings of the experiment are shown in Table 3.
[0270] Table 3 Experimental Parameter Settings
[0271] Parameter Name max_length 128 batch_size 32 nb_epochs 60 bert_lr 1e-4 m 0.7 gcn_layers 2 n_hidden 200 dropout 0.5
[0272] Comparison model: Long Short-Term Memory Network + Attention Mechanism (LSTM+Attention): The combination of the Long Short-Term Memory Network (LSTM) and the Attention Mechanism (Attention), i.e., LSTM+Attention, is a model enhanced on the basis of the classical LSTM network. It combines the temporal modeling ability of LSTM and the weighting characteristics of the Self-Attention mechanism, enabling the model to handle sequence data more flexibly and focus on the important parts in the sequence.
[0273] Capsule Network (CapsNet): A capsule is a group of neurons, and their input and output vectors are used to represent the instantiation parameters of a specific entity type (i.e., the occurrence probability of an object or concept entity and its related attributes). Essentially, the capsule network is a neural network designed to perform inverse graph parsing, consisting of multiple capsules. Each capsule can be regarded as a function that attempts to predict the existence of the target at a specific position and the instantiation parameters.
[0274] BERT: BERT (Bidirectional Encoder Representations from Transformers) is a pre-trained language model based on the Transformer encoder architecture, improved from GPT. Different from GPT's unidirectional language model, BERT uses bidirectional context information for word prediction and learns word-level and sentence-level representations through the Masked LM and Next Sentence Prediction tasks. The key features of BERT include: 1) Bidirectional encoding, predicting words through Masked LM and leveraging context information from both the front and back. 2) Using the Transformer encoder, which has deeper layers and better parallelism. 3) Sentence-level negative sampling, learning the relationship between sentence pairs to determine if a sentence is the next one.
[0275] Transformer: The Transformer is based on the Attention mechanism and is divided into an encoder and a decoder, each consisting of 6 stacked layers. Each layer includes a multi-head attention mechanism and a feed-forward network, and uses residual connections and normalization. Sequential information is injected through positional encoding, and the output predicts the next Token through Softmax. The Transformer uses three multi-head attention mechanisms: encoder-decoder Attention, encoder Self-Attention, and decoder Self-Attention (with Masking added to prevent illegal values). This structure is highly parallel, improving the training efficiency.
[0276] Llama-3.1-70B: Llama-3.1-70B is a large-scale pre-trained language model based on the Transformer architecture launched by Meta, containing 7 billion parameters. This model uses self-attention and multi-head attention mechanisms, and processes input data in parallel through multiple Transformer layers, capable of effectively modeling long-term dependencies and supporting multi-task processing. Llama-3.1-70B is trained on a large-scale, multi-lingual text corpus, with data sources including the Internet, books, Wikipedia, and other fields, ensuring its efficient performance and wide application capabilities in multiple natural language processing tasks such as text generation, question answering, and sentiment analysis.
[0277] GPT-4O: GPT-4O is an enhanced version based on GPT-4. It continues the basic design of the Transformer architecture and autoregressive language modeling, and has powerful multitasking capabilities. It is trained on a large-scale and diverse corpus, covering various types of text data, supports multilingual tasks, and performs well in fields such as text generation, question answering, translation, and summarization. GPT-4O may be optimized in terms of reasoning ability, task adaptability, and processing speed, and can provide customized services in specific application scenarios such as medical and legal fields, while retaining the powerful language understanding and generation capabilities of GPT-4.
[0278] Qwen2-1.5b: Qwen-2 is a large-scale pre-trained language model based on the Transformer architecture, optimized for Chinese and multilingual natural language processing tasks. It can efficiently perform various tasks such as text generation, text classification, sentiment analysis, and machine translation, and has powerful multitasking capabilities. Qwen-2 is trained on a vast amount of Chinese and multilingual datasets, has deep semantic understanding and reasoning capabilities, and can handle complex language expressions and context information.
[0279] Evaluation Metrics: The models of the present invention and the comparative experiments are analyzed through 4 metrics: accuracy, precision, recall, and F1 value.
[0280] Accuracy: Measures the overall correctness of a classification model, and the calculation formula is:
[0281]
[0282] Among them, Tp is the True Positive, Tn is the True Negative, Fp is the False Positive, and Fn is the False Negative.
[0283] Precision: Calculates the accuracy of the model's prediction of positive samples, and the formula is:
[0284]
[0285] Recall: Measures the ability of the model to correctly predict positive samples, and the formula is:
[0286]
[0287] F1 value: Considers the harmonic mean of precision and recall, and the formula is:
[0288]
[0289] For a given predicted label, the precision and recall can be quickly calculated by using a confusion matrix. The confusion matrix for a binary classification problem contains four basic results: true positives (Tp), false positives (Fp), true negatives (Tn), and false negatives (Fn). In the confusion matrix, the columns represent the true labels and the rows represent the predicted labels. The intersection of a row and a column corresponds to these four results, as shown in Table 4.
[0290] Table 4 Character Meanings of the Confusion Matrix
[0291]
[0292] Experimental Results
[0293] Table 5 Comparative Experimental Results
[0294]
[0295]
[0296] To verify the proposed noun metaphor recognition model MetaGCN-ERNIE based on large language models and graph neural networks, the present invention conducted experiments on a public dataset and comprehensively compared the proposed method with classical neural network models, deep learning models, and modern large language models. The experimental results show that the MetaGCN-ERNIE model outperforms all the above-mentioned comparison models in the metaphor recognition task. The main reasons are as follows: First, the model of the present invention generates more training data through data augmentation technology, enabling the model to more fully learn the complex features and diverse expressions of metaphors during the training process. Second, the model constructs a semantic graph using syntactic dependency relationships and pointwise mutual information (PMI). This process helps the model capture the deep grammatical and semantic connections between words, thereby more accurately identifying noun metaphors. Third, the model extracts features by using a powerful ERNIE pre-trained model, fully considering the context semantics of the text, especially in the understanding of long-distance dependency relationships, enhancing the metaphor recognition ability. In addition, the model uses a graph convolutional network to label sentences and words, and optimizes the representation of sentences through graph convolutional operations, enabling the final model to more effectively capture the potential information in metaphorical expressions. These innovative designs enable MetaGCN-ERNIE to achieve significant performance improvement when dealing with noun metaphors in complex contexts.
[0297] Based on the above embodiments, the embodiments of the present application further provide a computer program, which, when running on a computer, causes the computer to execute the method provided by the above embodiments.
[0298] Based on the above embodiments, an embodiment of the present application further provides a computer storage medium. A computer program is stored in the computer storage medium. When the computer program is executed by a computer, the computer is caused to execute the method provided in the above embodiments.
[0299] Among them, the storage medium can be any available medium that can be accessed by a computer. Taking this as an example but not limited to: the computer-readable medium may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, magnetic disk storage medium or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer.
[0300] Based on the above embodiments, an embodiment of the present application further provides a chip. The chip is used to read the computer program stored in the memory and implement the method provided in the above embodiments.
[0301] Based on the above embodiments, an embodiment of the present application provides a computer program product. When the computer program product runs on an electronic device, it implements the method provided in the above embodiments.
[0302] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.
[0303] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0304] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the process inFigure 1 one process or multiple processes and / or boxes Figure 1 the functions specified in one box or multiple boxes
[0305] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or boxes Figure 1 the one box or multiple boxes
[0306] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these changes and modifications.
Claims
1. A noun metaphor recognition method based on a large language model and graph neural network, characterized in that: A metaphor recognition model is used, wherein the metaphor recognition model includes a preprocessing model and a graph convolution model, and the recognition method includes the following steps: training the metaphor recognition model; Inputting the text to be processed into the trained preprocessing model, and the preprocessing model outputs a first feature representation; Constructing a semantic graph of the text to be processed according to syntactic dependencies and point mutual information; Inputting the first feature representation and the semantic graph into a trained graph convolutional network, and the graph convolutional network outputs a second feature representation; Classifying the second feature representation through a classifier to obtain a noun metaphor recognition result in the text to be processed; The step of training the metaphor recognition model includes: Performing data augmentation on the first data set by using the first large language model to construct a second data set; Performing data distillation on the second data set through a second language model to construct a third data set; Inputting the third data set into a preprocessing model, the preprocessing model outputting a first feature representation; Constructing a semantic graph of the third data set according to syntactic dependencies and point mutual information; Inputting the first feature representation and the semantic graph into a graph convolutional network, and the graph convolutional network outputs a second feature representation; Classifying the second feature representation through a classifier; The first data set includes Chinese sentences containing noun metaphors or Chinese sentences containing non-noun metaphors; The prompts for the first language model in the data augmentation include prompts for sentences containing noun metaphors and prompts for sentences not containing noun metaphors; The second data set includes augmented sentences of sentences containing noun metaphors and augmented sentences of sentences not containing noun metaphors; The prompts for the second largest language model in the data distillation include prompts for sentences containing noun metaphors and their augmented sentences, and prompts for sentences not containing noun metaphors and their augmented sentences.
2. The noun metaphor recognition method based on a large language model and a graph neural network according to claim 1 is characterized in that: The first data set includes Chinese sentences containing noun metaphors or Chinese sentences containing non-noun metaphors, which are represented by the following formula: Data Input ={Sentence} In the formula, Sentence represents a Chinese sentence containing a noun metaphor or a Chinese sentence containing a non-noun metaphor; The prompts for the first language model in the data augmentation, wherein the sentence Data containing noun metaphors Input The prompt is expressed as follows: Sentences metaphor =Qwen(Data Input + "Generate five similar sentences that include noun metaphors?") In the formula, Sentences metaphor Represents a sentence containing a noun metaphor Data Input The expanded sentence of The prompts for the first language model in the data augmentation, wherein the sentence Data that does not contain noun metaphors Input The prompt is expressed as follows: Sentences non-metaphor =Qwen(Data Input + "Generate five similar sentences that do not include noun metaphors?") In the formula, Sentences non-metaphor Represents a sentence that does not contain a noun metaphor Data Input The expanded sentence of The second data set includes Sentences metaphor and Sentences non-metaphor ; The prompts for the second largest language model in the data distillation, where the sentence Data containing noun metaphors Input And its augmented sentences metaphor The prompt is: Verification metaphor = ChatGLM(Sentences metaphor ,"Are all of them noun metaphors?" The prompts for the second largest language model in the data distillation, where Data Input And its augmented sentences non-metaphor The prompt is expressed as follows: Verification non-metaphor = ChatGLM(Sentences non-metaphor ,"Are all of them not noun metaphors?" If the answer to the prompt is yes, then the third data set is represented by the following formula: Final data =Verification metaphor +Verification non-metaphor If the answer to the prompt is no, then the sentence Data containing the noun metaphor Input And its augmented sentences metaphor The sentence that is not a noun metaphor is deleted from the sentence, which is represented by the following formula: Filtered metaphor = ChatGLM(Sentences metaphor ,"Delete sentences that are not noun metaphors") If the answer to the prompt is no, then the sentence Data that does not contain noun metaphors Input And its augmented sentences non-metaphor The sentence in which the noun metaphor is deleted is represented by the following formula: Filtered non-metaphor = ChatGLM(Sentences non-metaphor ,"Delete the sentence that is a noun metaphor") Then the third data set is expressed by the following formula: Final data =Filtered metaphor +Filtered non-metaphor In the formula, Final data The third data set.
3. The noun metaphor recognition method based on a large language model and a graph neural network according to claim 1 is characterized in that: The first large language model includes a Qwen model, the second large language model includes a ChatGLM model, and the preprocessing model includes an ERNIE model.
4. The noun metaphor recognition method based on a large language model and a graph neural network according to claim 1 is characterized in that: in, The step of inputting the third data set into a preprocessing model, and the preprocessing model outputting a first feature representation, comprises the following steps: S310. Finalizing the third data set data The sentences in the sentence are segmented, and the words after segmentation are {w1,w2,…,w n }, each word generates a pre-trained word embedding E(w i ), expressed by the following formula: E(w i )=Embedding(w i ) S320. Perform position encoding on the word, wherein the position encoding is represented by P(i), which is represented by the following formula: In the formula, i is the position information of the word, j is the dimension of the position encoding, and d is the embedding dimension; S330. Obtain the weighted sum of the value of each word through the attention mechanism of the word vector H0(i) and perform weighted aggregation of context features, concatenate the output of each attention head, and map it back to the dimension of the model through linear transformation to obtain the multi-head attention mechanism output MultiHead(Q, K, V); Among them, the word vector H0(i) of each word is expressed by the following formula: H0(i)=E(in i )+P(i); S340. The output of each layer is normalized by layer normalization, which is expressed as follows: H L =LayerNorm(H0(i)+MultiHead(Q,K,V)) S350. Output the first feature representation, which is represented by the following formula: H final =H L In the formula, H final Represents the first feature representation.
5. The noun metaphor recognition method based on a large language model and a graph neural network according to claim 4 is characterized in that: in, In step S330, the multi-head attention mechanism outputs MultiHead(Q, K, V), including the following steps: S331. Use different W for the input word vector H0(i) Q , W K , W V Generate multiple Q, K, V, and perform attention calculation on each head independently to obtain Attention1, Attention2, .....Attention h ; S332. Splice the output of each head and transform it by linear transformation W O Mapping back to the dimensions of the model is represented by the following formula: MultiHead(Q,K,V)=Concat(Attention1,Attention2,...,Attention h )W O In the formula, MultiHead(Q,K,V) means, represents the linear transformation matrix, d model represents the dimension of the model, h represents the number of heads, and d v Represents the dimensions of a single head; The calculation of Attention in step S331 includes the following steps: S3311. The word vector H0(i) is fed into the attention mechanism, and the query, key, and value matrices are generated from the word vector H0(i) through linear transformation, which is expressed as follows: Q=H0(i)W Q ,K=H0(i)W K ,V=H0(i)W V In the formula, Respectively represent the query projection matrix, key projection matrix, and value projection matrix; Represent the query matrix, key matrix, and value matrix respectively; S3312. Calculate the similarity score Score(i,j) between the query matrix Q and the key matrix K, which is expressed by the following formula: In the formula, Q i represents the query vector of the i-th word, K j represents the key vector of the j-th word, the similarity score Score(i,j) represents the strength of the relationship between position i and position j, and the superscript (t) represents the calculation of the t-th head in the multi-head attention mechanism; S3313. Integrate the scores of all positions into a matrix form, represented by the following formula: In the formula, represents the similarity score of each position in the sentence to all other positions, and n represents the number of words; S3314. Scale the similarity score as follows: S3315. Calculate the weighted sum of the value of each word, expressed by the following formula: In the formula, the molecule It means converting the correlation score of the i-th position to the j-th position into a positive value, and the denominator is represents the sum of all relevance scores in the i-th row for normalization; S3316. Add the weighted sum of the values of each word to a ij For the weighted value matrix V j , weighted aggregation of context features is performed, which is expressed as follows: In the formula, n represents the number of words.
6. The noun metaphor recognition method based on a large language model and a graph neural network according to claim 1, characterized in that: in, The step of constructing a semantic graph of the third data set according to the syntactic dependency relationship and the point mutual information includes the following steps: S410. Finalizing the third data set data The sentences in the vocabulary building include S411. Through S={s1,s2,…,s N } indicates the third data set Final data The set of sentences in , where each sentence s i Contains a set of vocabulary; S412. Each sentence s i Split into a series of words s i ={w i1 ,w i2 ,…,w ij }, where w ij Indicates sentence s i The jth word in ; S420. Perform a union operation on the vocabulary sets in all sentences to obtain a vocabulary table, which is represented by the following formula: In the formula, vocab represents the vocabulary, U represents the union operation, and w ij Indicates sentence s i The jth word in , N represents the number of sentences; S430. For each word w k ∈vocab, using word frequency word_freq(w k ) represents word w k The number of occurrences in all documents, traverse all documents, and count the frequency of each word. The frequency of each word in all documents is accumulated and recorded in the word frequency dictionary word_freq, which is expressed by the following formula: In the formula, is an indicator function, indicating the result of conditional judgment. In the formula, if word w k Appears in sentence S i In Otherwise, 0; S440. Define a fixed-size context window, and calculate the number of co-occurrences of word pairs by traversing the words in the sentence, including S441. Sentence S i The word sequence in is S i ={w i1 ,w i2 ,…,w ij }, where j represents the number of words, the window size is window_size, for sentence S i Each word w in ij , define its context window as: window(w ij )={w ij-k ,w ij-k+1 ,...,w ij+k } Where k is an integer from -window_size to window_size; S442. For each word pair (w p ,w q ) appears in the context window, by maintaining a counting dictionary word_pair_count, recording each pair of vocabulary w p and w q The number of occurrences in the corpus is expressed as follows: word_pair_count(w p ,w q )=word_pair_count(w p ,w q )+1 Where word_pair_count() represents the counting function, w p and w q is any pair of different words in the window; S450. Calculate the weight of the edge between two points, where any two points are represented by any two words. If there is no syntactic dependency between the two points, the PMI value is used as the weight. If there is a syntactic dependency, the PMI value plus 10 is used as the weight. Among them, the PMI calculation is as follows: In the formula, word_pair_count(w i ,w j ) is a word pair (w i ,w j ) appears in the document, N represents the number of sentences, where: Among them, if there is a syntactic dependency, the PMI value plus 10 is used as the weight, which is expressed by the following formula: PMI(w i ,w j )=PMI(w i ,w j )+10,if(w i ,w j )∈special_word_pairs Where special_word_pairs is the set of syntactic dependency relationships between words obtained by spacy after calculating the syntactic dependency relationships.
7. The noun metaphor recognition method based on a large language model and a graph neural network according to claim 1, characterized in that: in, In the step, the first feature representation and the semantic graph are input into a graph convolutional network, and the graph convolutional network outputs a second feature representation, including using an update formula of a graph convolutional network GCN to update and optimize the first feature representation through the semantic graph to aggregate neighbor features and output a second feature representation, wherein the update formula is expressed by the following formula: In the formula, H (L) is the node feature representation of the Lth layer, initially L = 0, and the output of the preprocessing model is H (0) ; is the adjacency matrix of the semantic graph, including self-connections, is the degree matrix of the node; is the Laplace normalization operation of the semantic graph; W (I) is the trainable parameter of layer I; σ() is the activation function; Represents the second feature representation.
8. The noun metaphor recognition method based on a large language model and a graph neural network according to claim 1, characterized in that: The second feature representation is passed through a linear layer W c and the bias term b c Classify and output a scalar y pred , indicating the probability that the sample belongs to a certain category; In the formula, y pred is the predicted output of the model, which indicates the probability that the input sentence belongs to a noun metaphor, and sigmoid() indicates the sigmoid function; The loss function of the metaphor recognition model training is shown in the following formula: Where, LOSS all Represents the total loss, γ is a hyperparameter; in: Loss Ernie =-(and true log(y pred )+(1-and true )log(1-y pred )) In the formula, Loss Ernie represents the loss of the preprocessing model, y true represents the true label, y pred represents the predicted output of the model; In the formula, represents the loss of the graph convolutional network, C represents the number of categories, and the superscript (c) represents the number of the current category.
9. An electronic device, comprising: One or more processors, a memory, and one or more programs; wherein the one or more programs are stored in the memory, and the one or more programs include instructions, which, when executed by the electronic device, enable the electronic device to perform any method of claims 1-8.
10. A computer-readable storage medium, comprising a computer program, which enables an electronic device to execute any one of the methods of claims 1 to 8 when the computer program is executed on the electronic device.
Citation Information
Patent Citations
Text-oriented metaphor recognition model and method
CN117669591A
Encoder, system and method for metaphor detection in natural language processing
US20210271822A1