Chinese text classification method based on metaphor association and label constrained contrastive learning
By constructing a Chinese text classification model based on metaphorical association and label constraints, the traditional model's shortcomings in ambiguity processing and semantic alignment are solved, and efficient classification and deep semantic understanding of Chinese text are achieved.
Patent Information
- Application Number
- CN202510738203.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-06-04
AI Technical Summary
Traditional Chinese text classification models are difficult to effectively capture the semantic ambiguity caused by the polyphonic and polysensory Chinese characters, and the large language model lacks a semantic alignment mechanism in the generation of Chinese metaphors, resulting in a degradation of classification performance.
A Chinese text classification model based on metaphorical association and label constraints is constructed, a Chinese character ambiguity is quantified through the ambiguity recognition module, and a large language model is used to generate metaphorical interpretation and label definition. A context and metaphorical attention mechanism are used to improve semantic alignment, and a triple comparison loss function optimization model is designed.
It improves the accuracy and robustness of Chinese text classification, can effectively capture the correlation between metaphor content and labels, and improves the deep semantic understanding ability of Chinese text classification.
Smart Images

Figure CN120256638B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of natural language processing and artificial intelligence, and specifically relates to a Chinese text classification method based on metaphor association and label-constrained contrastive learning. Background Art
[0002] With the rapid development of natural language processing technology, Chinese text classification plays an important role in fields such as public opinion analysis, literary research, and commercial applications. However, the language characteristics of Chinese as an ideographic writing system pose unique challenges to its in-depth semantic understanding. Compared with languages such as English that use alphabetic writing, Chinese characters carry rich cultural metaphors and context-dependent extended semantics through the combination of form, sound, and meaning. Especially in styles such as satirical comments and classical poems, metaphorical expressions are highly dense and present multi-level cognitive associations. For example, "A thousand sails pass by the side of the sunken ship, and ten thousand trees spring up ahead of the sick tree" metaphorically expresses the philosophical thought of adversity and new life through the contrast of natural scenes, and there is often a significant cognitive gap between its literal features and real semantics. Traditional text classification models (such as TextCNN, BiLSTM) mainly rely on surface features such as word frequency and syntax to construct sparse vector representations, making it difficult to capture implicit semantic associations and even more unable to effectively bridge the semantic gap between literal features and the label space, resulting in a significant decline in classification performance in metaphor-dense scenarios. Although pre-trained language models (such as BERT, RoBERTa) have improved semantic representation capabilities through large-scale corpus learning, their original design is for general contexts and does not quantify the semantic ambiguities caused by the unique polyphones and multiple word natures in Chinese. For example, the Chinese character "长" has multiple word natures such as adjective (cháng) and verb (zhǎng). Traditional methods lack a dynamic evaluation mechanism for such ambiguous features, making it difficult for the model to distinguish the contribution degrees of key semantic nodes in metaphorical texts.
[0003] In recent years, large language models (LLMs) have shown potential in the field of metaphor understanding. Through prompt engineering, metaphor explanations related to the context can be generated. However, existing research mostly focuses on English contexts, and there are two major bottlenecks in Chinese metaphor generation: First, the cultural specificity of Chinese character configurations makes metaphorical associations highly dependent on historical allusions and idiom knowledge, and it is difficult for general large language models LLMs (such as ChatGPT, Deepseek, etc.) to accurately capture such implicit associations; Second, there is a lack of a semantic alignment mechanism between label definitions and generated metaphors, resulting in the disconnection between the generated content and the classification target. Summary of the Invention
[0004] The present invention aims to provide a Chinese text classification method based on metaphor association and label-constrained contrastive learning. By quantifying text ambiguity and integrating metaphor interpretation with label definition, the method addresses the semantic gap between metaphorical expressions and literal features in Chinese texts. The method comprises the following steps: constructing a metaphor association model for classifying Chinese texts; the metaphor association model comprises an input layer, a representation layer, and a classification layer;
[0005] The input layer is used to classify the Chinese text Divide into character sequences , the Chinese text Mapping into metaphorical explanation sentences , and mapping classification labels to label definitions ;
[0006] The presentation layer is used to convert the character sequence , metaphorical explanation sentence and tag definitions Processed into text embedding vectors , metaphor interpretation embedding vector and labels define embedding vectors ; Use contextual attention mechanism to capture text embedding vector and labels define embedding vectors The contextual relationship between them, get the Chinese text The weighted context representation vector of the word is obtained by using the metaphor attention mechanism to adjust the metaphor interpretation under the label constraint, and obtain the metaphor attention representation vector. The context attention mechanism and the metaphor attention mechanism are used to identify the label directionality of literal features and metaphor features respectively, and improve the deep semantic alignment between text and labels.
[0007] The classification layer is used to classify the Chinese text The weighted context representation vector is concatenated with the metaphor attention representation vector and input into the linear transformation layer to obtain the classification prediction result; the concatenation refers to connecting the vectors in parallel to retain the independence of the features of the two.
[0008] Furthermore, the input layer includes an ambiguity recognition module and a prompt word management module;
[0009] The ambiguity recognition module is used to convert the Chinese text Divide into character sequences , a character sequence Expressed as:
[0010]
[0011] in, Indicates the first character in a sequence Chinese characters, It is Chinese text The total number of characters in ;
[0012] The prompt word management module is used to convert Chinese text Mapping into metaphorical explanation sentences , and mapping classification labels to label definitions .
[0013] Furthermore, the Chinese text Mapping into metaphorical explanation sentences , and mapping classification labels to label definitions , expressed as:
[0014] ,
[0015] ,
[0016] in, Indicates the generation of a large language model, Represents a prompt word template for generating context-based metaphor interpretation, Indicates the prompt word template for generating label definitions. represents a predefined set of classification labels, Indicates the number of tags, This is the task description for the dataset.
[0017] Furthermore, the ambiguity recognition module is also used to calculate Chinese text Ambiguity score , which is expressed as follows:
[0018] ;
[0019] in, Represents characters ambiguity score;
[0020] ;
[0021] in, For characters The number of pronunciation types, For characters The number of parts of speech.
[0022] By introducing the quantification of the polysemy of Chinese characters and combining the phonetic complexity and part-of-speech complexity of the characters, the ambiguity score of the text is calculated, further enhancing the model's ability to handle polysemous text.
[0023] Furthermore, a shared BERT series encoder is used at the representation layer to transform the character sequence , metaphorical explanation sentence and tag definitions Processed into text embedding vectors , metaphor interpretation embedding vector and labels define embedding vectors .
[0024] Furthermore, the contextual attention mechanism is used to capture the text embedding vector and labels define embedding vectors The contextual relationship between them is used to obtain the weighted context representation vector of the Chinese text. The specific steps are:
[0025] set up , is the embedding dimension, Indicates the The encoding embedding of characters;
[0026] set up ,in is the number of labels, Indicates the The encoding embedding defined by the labels;
[0027] The initial contextual attention distribution uses attention weights Expressed as:
[0028]
[0029] in, i∈[1, n] , j∈[1, m] , express The transpose of Indicates the Characters and The attention weights between the labels are defined, Indicates the The encoding embedding of characters;
[0030] Computing character embeddings The ambiguity adjustment factor , expressed as:
[0031]
[0032] Based on the ambiguity adjustment factor Attention weight Make adjustments to get the adjusted attention weight , expressed as:
[0033]
[0034] Adjust the attention weight Perform normalization to obtain the normalized adjusted attention weight , expressed as:
[0035] ;
[0036] in, Indicates the Characters and Adjust the attention weights between the label definitions;
[0037] Chinese text The weighted context representation vector , expressed as:
[0038]
[0039] in, Is to adjust the attention weight Matrix representation of ;
[0040] Next, a pooling operation is used to fix the representation vector The size of the final Chinese text is obtained The weighted context representation vector :
[0041]
[0042] Furthermore, the metaphor attention mechanism is used to adjust the metaphor interpretation under the label constraint to obtain the metaphor attention representation vector. The specific steps are as follows:
[0043] Calculating metaphor attention weight matrix , expressed as:
[0044] ,
[0045] in, represents the metaphor attention weight matrix, represents the metaphor interpretation embedding vector, Define embedding vectors for labels, Represents the label definition embedding vector The transpose of
[0046] The embedding vector obtained after the metaphorical attention mechanism is calculated , expressed as:
[0047]
[0048] Embedding vector The pooling operation is as follows:
[0049]
[0050] The final metaphorical attention representation vector defined by fusing attention information and labels.
[0051] Furthermore, the Chinese text The weighted context representation vector of is concatenated with the metaphor attention representation vector and input into the linear transformation layer to obtain the classification prediction result, which is:
[0052] Representing the context and metaphorical attention representation vector Splicing to form Chinese text The final representation vector , expressed as:
[0053] Z=[ R C ; R D ]
[0054] Will represent the vector Input linear transformation layer, converted into a numerical vector of classification labels :
[0055] ,
[0056] in, is the weight matrix, is the bias vector;
[0057] Then the Softmax function is used to obtain the probability distribution of the predefined categories, and finally the label with the highest probability is selected As Chinese text The classification prediction result is expressed as:
[0058] .
[0059] Furthermore, in the process of training the metaphor association model, the loss function used is:
[0060]
[0061] in, λ∈(0, 1] is a proportional adjustment hyperparameter used to balance the cross entropy loss and triplet contrast loss The weight of
[0062] Cross Entropy Loss , expressed as:
[0063] ;in, is a dataset for classification tasks, Represented as Chinese text Assigning Tags Calculate the probability of
[0064] ;
[0065] in, is the similarity distance function, i∈[1,n] For the data set Chinese text The index of Represents the first Chinese text The context representation vector of , as a negative sample, belongs to the context representation vector ; Represents the first Chinese text The metaphorical interpretation embedding vector of , also called the anchor point in the ternary contrast loss, belongs to the metaphorical interpretation embedding vector ; Represents the first Chinese text The label definition representation vector of is a positive sample, which belongs to the label definition representation vector ; is a marginal parameter.
[0066] Beneficial effects: The present invention models Chinese text based on the metaphor association model (L-MAM) of the large language model. Compared with traditional processing methods, it can more effectively use the advantages of the large language model to analyze the metaphorical content and label definitions of Chinese text, and realize the prediction of the category to which it belongs, thereby providing reference information for the analysis and mining of Chinese text data on the Internet. It has certain practical application value and can bring certain potential economic benefits to some related text information platforms.
[0067] In this paper, a metaphor association model (L-MAM) based on a large language model is used to model and classify Chinese texts. By combining a large language model with deep learning technology, it can provide certain technical support for the practice of Chinese natural language processing, text mining and other fields. It is a technical solution that integrates ambiguity quantification, dynamic attention and label-constrained contrastive learning.
[0068] (1) The ambiguity recognition module ARM in the present invention measures and quantifies the ambiguity of the input original text by deeply analyzing the number of pronunciations and different parts of speech of Chinese characters;
[0069] (2) A metaphor association mechanism is designed to enrich text representation by leveraging the capabilities of the Large Language Model (LLM) and introducing reasonable extended explanations and label definitions based on context generation.
[0070] (3) Two dual attention mechanisms are designed, including the contextual attention mechanism and the metaphorical attention mechanism; the contextual attention mechanism is used to align the contextual character semantics with the label definition, and the metaphorical attention mechanism is used to align the metaphorical interpretation with the label definition, thereby improving the accuracy and rationality of the text semantic representation.
[0071] (4) A “label-constrained context-metaphor contrast learning mechanism” was designed, which can effectively bring metaphor interpretation closer to label definition while keeping metaphor interpretation away from relatively irrelevant contextual character semantics. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments are briefly introduced below.
[0073] Figure 1 This is a flow chart of the Chinese text classification method based on metaphor association and label constraint contrast learning of the present invention.
[0074] Figure 2 It is a flow chart of the prompt word management module of the present invention mapping Chinese text into metaphorical explanation sentences.
[0075] Figure 3 This is a diagram of the Chinese text classification model based on metaphor association and label constraint contrast learning in the present invention. DETAILED DESCRIPTION
[0076] The following is a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0077] The present invention provides a Chinese text classification method based on metaphor association and label constraint contrast learning. The Chinese text classification model is used to effectively predict the category label to which a Chinese text belongs. By utilizing the capabilities of the Large Language Model (LLM), combined with potential metaphor interpretation and label definition, the effect of Chinese text classification tasks is improved. Given a Chinese text and a set of predefined classification labels , assign each text to the most appropriate category label. Let the classification function is the classification function that we want to learn. It can effectively capture the literal and metaphorical meanings hidden in the text, as well as the deep semantics defined by the tags. The problem can be defined as:
[0078] ,
[0079] in, According to Chinese text Extracted metaphorical interpretation, Represents a classification label set The tag definition, represents the task description of the dataset, represents the conditional probability.
[0080] The present invention proposes a Chinese text classification method based on metaphor association and label-constrained contrastive learning, including constructing a metaphor association model for Chinese text classification; in this specific embodiment, the metaphor association model is a metaphor association model based on a large language model (LLM-based Metaphorical Association Model, L-MAM), and the construction method is as follows Figure 1 As shown, the details are as follows;
[0081] Step 1: Input layer design
[0082] The input layer, which forms the foundation of the Metaphor Association Model (L-MAM) based on a large language model, is responsible for processing and encoding the input text. It converts the raw text input into a format suitable for processing by subsequent layers. It also includes the Ambiguity Recognition Module (ARM) and the Cue Word Management Module, driven by the large language model. The ARM extracts ambiguous features from the text, while the Cue Word Management Module generates metaphorical interpretations and label definitions.
[0083] Step 1.1, build the ambiguity recognition module ARM:
[0084] A character or word may have multiple different meanings, a phenomenon known as "ambiguity". In the context of Chinese language and culture, ambiguity is very common in certain literary forms, such as satire / criticism, classical literature, poetry, etc. In this paper, the ambiguity of Chinese characters is quantified by considering both the phonological and syntactic changes of Chinese characters. Since Chinese characters are the basic components of Chinese text, it is necessary to input Chinese text. Divide into character sequences ,in, Indicates the first character in a sequence Chinese characters, It's text The total number of characters in the language model is used to classify each Chinese character Perform Part-of-Speech (PoS) tagging. Then, calculate the number of different pronunciation types and the number of different词性 types represented by each Chinese character character, and calculate the ambiguity score for each character . The ambiguity score can be used to adjust the importance weights of each character in the attention distribution. The ambiguity score is expressed as:
[0085] ;
[0086] where, is the number of pronunciation types of the character , and is the number of词性 types of the character .
[0087] For example, the character "长" has 2 pronunciations (cháng / zhǎng) and 4词性 (adjective, verb, adverb, noun). The ambiguity score of the character "长" is
[0088] ;
[0089] In this embodiment, when the ambiguity score of a character is greater than 0.75, it is considered to have a high ambiguity potential. The ambiguity score of the character "长" is greater than 0.75, so the character "长" has a high ambiguity potential.
[0090] Next, define the ambiguity score of the Chinese text as the mean of the ambiguity scores of all characters, expressed as:
[0091]
[0092] where, , and the ambiguity score of the text is used for subsequent attention weight adjustment to guide the model to focus on high-ambiguity regions.
[0093] Step 1.2, construct an LLM prompt management module
[0094] Large language models (LLMs) can be used to solve the metaphor problem that traditional text classification models only rely on surface features. In this invention, a large language model is used as the prompt management module in this application, and the input Chinese text It maps the text into a metaphorical explanation sentence and maps the classification label to the corresponding label definition to activate the implicit semantic associations and cultural connotations in the text, perform metaphorical interpretation of the input text and semantic associations of the label definition, and improve the understanding of the deep semantics of Chinese text in text classification tasks, especially in texts with high information entropy and sparse literal features, so as to mine more implicit information.
[0095] Given input Chinese text , predefined classification label sets and dataset task description , the corresponding metaphorical explanation sentence and tag definitions It can be expressed as:
[0096] ,
[0097] ,
[0098] in, Indicates the generation of a large language model, Indicates the prompt word template for generating context-based metaphor interpretation, as shown in Table 1. Indicates the prompt word template for generating label definitions, as shown in Table 2. and prompt templates The preset metaphor-guided sentence structure and label explanation structure are in line with the Chinese text context.
[0099] Table 1. Large language model prompt word templates for generating metaphor interpretations
[0100]
[0101] Table 2 Large language model prompt word templates used to generate label definitions
[0102]
[0103] Step 2: Presentation layer design
[0104] The representation layer is responsible for encoding various inputs into vector representations, which include character sequences , metaphorical explanation sentence and tag definitions , and applies an attention mechanism to properly fuse the vector representations. The attention mechanism includes a contextual attention mechanism and a metaphorical attention mechanism, which enables metaphorical information, label definition information, and literal features to be properly integrated in the text representation process, thereby improving classification accuracy.
[0105] First, the shared BERT (Bidirectional Encoder Representations from Transformers) series encoder is used to process the character sequence after word segmentation. , metaphorical explanation sentence and tag definitions :
[0106] ,
[0107] in, Represents a character sequence The encoded embedding vector of , also written as text embedding vector, represents the embedding dimension, Metaphorical interpretation sentence The encoding embedding vector of , also written as metaphor interpretation embedding vector, Indicates tag definition The encoding embedding vector of , also written as the label definition embedding vector, Indicates the number of tags.
[0108] like Figure 2 Take the example of a line in a poem: "A thousand sails pass by the side of the sunken boat; a thousand trees bloom in front of the sick tree." On the surface, this line literally describes a natural scene. A non-native Chinese speaker or a traditional Chinese text representation model might be confused by this abstract, purely visual description. However, in reality, the contrasting imagery in Chinese text carries profound metaphorical meaning that goes far beyond literal meaning: the "sunken boat" and "sick tree" symbolize frustration and failure, representing negative states or experiences. In contrast, "a thousand sails pass by" and "a thousand trees bloom in spring" symbolize progress, hope, and rejuvenation, embodying positive change and new opportunities. This further leads to the metaphorical interpretation of this line: "The decline of old things does not hinder the flourishing of new things; all things in the world are constantly moving forward in constant renewal and hope."
[0109] Step 2.1: Build a contextual attention mechanism
[0110] Given the semantic difference between metaphorical information and literal features, a contextual attention mechanism is designed to capture the input text embedding vector and labels define embedding vectors The contextual relationship between them. The contextual attention mechanism enables the metaphor association model to focus on the most relevant part of the input text for each label definition.
[0111] set up , is the embedding dimension, Indicates the The encoding embedding of characters.
[0112] Similarly, let ,in is the number of labels, Indicates the The encoding embedding defined by the labels, then the initial context attention distribution is:
[0113] ,
[0114] in, i∈[1, n] , j∈[1, m] , represents the overall context attention weight matrix, express The transpose of Indicates the The encoding embedding of characters.
[0115] Next, considering that metaphorical semantics are often unrelated or contrary to the original semantics of literal features, this paper proposes to use the ambiguity score of the input text to adjust the attention contribution weight of each Chinese character. This process can help the model pay more attention to those parts of the text with higher ambiguity:
[0116] 1) Ambiguity adjustment: For each character embedding , calculate its ambiguity adjustment factor , the factor is based on the character Ambiguity score and Chinese text Average ambiguity score :
[0117]
[0118] if , it means that the ambiguity of this character is higher than the average level, and the metaphor association model of the present invention should pay more attention to this character. On the contrary, if , the model should pay less attention to this character.
[0119] 2) Adjusted attention distribution: Then, by dividing the initial attention weight and ambiguity adjustment factor Multiply to adjust the attention weight defined for each label to get the adjusted attention weight , expressed as:
[0120]
[0121] Then, adjust the attention weights for all characters Normalize to ensure attention weights are adjusted The sum of is 1, and the dynamically adjusted attention distribution matrix is obtained. :
[0122]
[0123] Indicates the Characters and Adjust the attention weights between the label definitions;
[0124] After the above weight adjustment process, Matrix and Multiply to get the input Chinese text The weighted context representation of:
[0125]
[0126] in, Is to adjust the attention weight Matrix representation of ;
[0127] Next, a pooling operation is used to fix the representation vector size:
[0128] ,
[0129] in, Enter Chinese text The context representation vector of .
[0130] Step 2.2: Build a metaphorical attention mechanism
[0131] Regarding the process of Chinese text comprehension, the deep semantics extracted from metaphorical associations may be consistent with the literal features or may be in sharp contrast with them.
[0132] Based on this, the metaphorical attention mechanism, a dual design of the contextual attention mechanism, corresponds to the positive reinforcement phenomenon in contrastive learning. Given the inherent alignment between the essence of metaphor and label definition, the metaphorical attention mechanism will be able to effectively implement metaphorical associations of literal semantics under label constraints, ensuring that the generated associations are both meaningful and targeted.
[0133] But the calculation process of the weight of the context attention mechanism is different. Represents the metaphor interpretation embedding vector and the metaphor attention weight matrix Calculated as follows:
[0134] ,
[0135] in, represents the metaphor attention weight matrix, represents the metaphor interpretation embedding vector, Define embedding vectors for labels.
[0136] The output of metaphor attention is:
[0137] ,
[0138] in, Represents the embedding vector of each label definition after the metaphorical attention mechanism is calculated.
[0139] In addition, in order to make the size of the embedding representation vector fixed, the pooling operation is adopted as follows:
[0140]
[0141] in, This is the final metaphorical attention representation vector that combines attention information and label definition.
[0142] Step 3: Classification layer design
[0143] The classification layer is designed to be responsible for the final text classification task based on the rich text vector representation obtained in the previous steps.
[0144] Finally, the context is represented as and metaphorical attention representation vector Splice together to form Chinese text The final representation vector :
[0145] Z=[ R C ; R D ]
[0146] Contrasting Context-Metaphor Learning:
[0147] To effectively integrate the literal and metaphorical semantics of Chinese text, this paper designs a "label-constrained contextual-metaphorical contrastive learning mechanism," or "label-constrained contrastive learning mechanism." This mechanism aims to effectively bridge the gap between metaphorical interpretation and label definition, maximizing similarity while simultaneously distancing metaphorical interpretation from relatively irrelevant contextual character semantics. This contrastive learning approach establishes stronger differentiation between metaphorical information, label definition information, and literal features, enabling the model to learn more valuable metaphorical associations and semantic connections during training. Accordingly, the triple contrastive loss function designed in this paper is as follows:
[0148]
[0149] in, is the similarity distance function, i∈[1,n] For the data set Chinese text The index of is the total number of texts in the dataset, Represents the first Chinese text The metaphor interpretation embedding vector of , also called the anchor point (metaphor interpretation embedding vector without attention calculation) in the ternary contrast loss, belongs to the metaphor interpretation embedding vector , Represents the first Chinese text The label definition representation vector of is a positive sample, which belongs to the label definition representation vector ; Represents the first Chinese text The context representation vector of , as a negative sample, belongs to the context representation vector ; is a margin parameter used to force the distance between the semantic anchor and the positive sample to be at least one interval smaller than the distance between the semantic anchor and the negative sample. ,If this condition is not met during the training process, the model parameters will be continuously adjusted and optimized through gradient descent.
[0150] The input of the classification layer is the final representation obtained through the metaphorical attention mechanism The present invention applies a linear layer to convert the final text vector representation into a numerical vector of classification labels :
[0151] ,
[0152] in, is the weight matrix, is the bias vector.
[0153] In the classification judgment process of the model, the numerical vector of the classification label The probability distribution of predefined categories will be obtained through the Softmax function, and the label with the highest probability will be selected. As input Chinese text Classification prediction results:
[0154]
[0155] The Softmax function maps the input vector to a probability distribution so that each output value is between [0, 1] and the sum of all values is 1. It is often used in the output layer of multi-classification problems.
[0156] The next step is to calculate the cross entropy loss between the predicted probability and the true label, which is more suitable for multi-class classification tasks because it encourages the model to output confident predictions for the correct category while penalizing incorrect predictions:
[0157] ;
[0158] in, is a dataset for classification tasks, Represented as Chinese text Assigning Tags The calculated probability of .
[0159] The metaphor association model constructed by the present invention is trained in the following steps:
[0160] To validate the effectiveness of our proposed method in Chinese text classification, we constructed and collated three Chinese metaphor datasets with varying degrees of metaphoricality and ambiguity. Furthermore, we provided character-level semantic support for the ambiguity recognition module (ARM). We used an online Chinese dictionary as the underlying database to capture linguistic attributes such as the number of pinyins, part of speech, and meaning distribution for each Chinese character.
[0161] The relevant statistical information of the three Chinese metaphor datasets involved in this invention is shown in Table 3:
[0162] Table 3 Three Chinese metaphor datasets
[0163]
[0164] 1. Weibo Sentiment Classification Dataset (WSC): This dataset contains 206,340 short Chinese Weibo messages covering 22 daily topics, categorized by sentiment: positive, neutral, and negative. The language in this dataset is close to everyday expressions, with a moderate level of metaphor, making it suitable for testing a model's general sentiment classification capabilities. The training set contains 165,072 data items, and the test set contains 41,268 data items.
[0165] 2. Chinese Classical Poetry Theme Classification Dataset (PTC): This dataset contains 17,103 classical Chinese poems, divided into 21 thematic categories. It is highly literary and rich in metaphorical expressions, making it suitable for topic classification tasks in metaphorical contexts. The training set consists of 15,393 poems and the test set consists of 1,710 poems. (Dataset reference: Y. Wei, L. Hu, Y. Zhu, J. Zhao, and B. Wu, “Knowledge-guided transformer for joint theme and emotion classification of Chinese classical poetry,” IEEE / ACM Transactions on Audio, Speech, and Language Processing, vol. 32, pp. 4783–4794, 2024.)
[0166] 3. Chinese Classical Poetry Semantic Matching Dataset (CPM): This dataset collects 24,498 modern paraphrased translations of classical Chinese poetry. The task is to select the sentence from four classical poetry candidates that best matches the text, given a modern text description. This task tests the model's ability to understand deep semantics and metaphorical content. The training set contains 21,778 sentences, and the test set contains 2,720. (Dataset reference: W. Li, F. Qi, M. Sun, X. Yi, and J. Zhang, “Ccpm: A Chinese Classical Poetry Matching Dataset,” arXiv preprint arXiv:2106.01979, 2021.)
[0167] The above three Chinese metaphor datasets cover texts of various styles and semantic understanding tasks, providing sufficient support for the systematic evaluation of the generalization ability, ambiguity recognition ability and metaphor association ability of the method of the present invention.
[0168] The Chinese text data in the dataset is preprocessed to remove null values or dirty data with a length of less than 5 characters in the Chinese text data; then the data in the dataset is shuffled, 80% of the data is used as the training sample set, and the remaining 20% is used as the test sample set;
[0169] Based on the Chinese text in the training sample set and the corresponding category label set, the metaphor association model is optimized and trained using the target loss function, and the performance of the metaphor association model is evaluated using the test sample set to obtain a trained metaphor association model.
[0170] Accordingly, the final objective loss function combines the standard cross entropy loss and triple contrast loss, and the metaphor association model parameters are updated through the optimization algorithm to minimize the total loss:
[0171] ,
[0172] in, λ∈(0, 1] is a scaling hyperparameter, The value range is (0,1], and the recommended range is [0.5,0.8], which is used to balance and The weight of , in order to achieve a trade-off between accurate classification and semantic comparison. It should be noted that when When , the main cross entropy loss responsible for the evaluation of the classification task will be completely removed, which is unreasonable in model design, so it is an open interval; when When , the total loss will be entirely cross-entropy loss, and the contrastive loss is completely removed. Through this design, combining cross-entropy loss and triplet contrastive loss, the model is able to distinguish different categories while effectively preserving and capturing the literal and metaphorical meanings in the text.
[0173] The method proposed in this paper is used for multi-category classification of Chinese texts and is suitable for specialized tasks involving metaphorical language expressions, such as sentiment classification, public opinion analysis, and automatic classification of poetry. A system constructed based on this method can be embedded in specific natural language understanding systems as a metaphor recognition module within an intelligent Chinese language processing platform, including but not limited to customer service question-and-answer systems, educational evaluation platforms, and automatic classification of cultural works.
[0174] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.
Claims
1. A Chinese text classification method, characterized in that: Construct a metaphor association model for classifying Chinese texts; The metaphor association model includes an input layer, a representation layer and a classification layer; The input layer is used to classify the Chinese text Divide into character sequences , the Chinese text Mapping into metaphorical explanation sentences , and mapping classification labels to label definitions ; The input layer includes an ambiguity recognition module and a prompt word management module; The ambiguity recognition module is used to convert the Chinese text Divide into character sequences , a character sequence Expressed as: ; in, Indicates the first character in a sequence Chinese characters, It is Chinese text The total number of characters in ; The prompt word management module is used to convert Chinese text Mapping into metaphorical explanation sentences , and mapping classification labels to label definitions ; The ambiguity recognition module is also used to calculate the Chinese text Ambiguity score , which is expressed as follows: ; in, Represents characters ambiguity score; ; in, For characters The number of pronunciation types, For characters The number of parts of speech; The presentation layer is used to convert the character sequence , metaphorical explanation sentence and tag definitions Processed into text embedding vectors , metaphor interpretation embedding vector and labels define embedding vectors ; Use contextual attention mechanism to capture text embedding vector and labels define embedding vectors The contextual relationship between them, get the Chinese text The metaphor attention mechanism is used to adjust the metaphor interpretation under the label constraint to obtain the metaphor attention representation vector; Use contextual attention mechanism to capture text embedding vectors and labels define embedding vectors The contextual relationship between them, get the Chinese text The weighted context representation vector of is constructed by the following steps: set up , is the embedding dimension, Indicates the The encoding embedding of characters; set up ,in is the number of labels, Indicates the The encoding embedding defined by the labels; The initial contextual attention distribution uses attention weights Expressed as: ; in, , , express The transpose of Indicates the Characters and The attention weights between the labels are defined, Indicates the The encoding embedding of characters; Computing character embeddings The ambiguity adjustment factor , expressed as: ; Based on the ambiguity adjustment factor Attention weight Make adjustments to get the adjusted attention weight , expressed as: ; Adjust the attention weight Perform normalization to obtain the normalized adjusted attention weight , expressed as: ; in, Indicates the Characters and Adjust the attention weights between the label definitions; Chinese text The weighted context representation vector , expressed as: ; in, Is to adjust the attention weight Matrix representation of ; Next, a pooling operation is used to fix the representation vector The size of the final Chinese text is obtained The weighted context representation vector : ; The metaphor attention mechanism is used to adjust the metaphor interpretation under label constraints to obtain the metaphor attention representation vector. The specific steps are as follows: Calculating metaphor attention weight matrix , expressed as: , in, represents the metaphor attention weight matrix, represents the metaphor interpretation embedding vector, Define embedding vectors for labels, Represents the label definition embedding vector The transpose of The embedding vector obtained after the metaphorical attention mechanism is calculated , expressed as: ; Embedding vector The pooling operation is as follows: ; The final metaphorical attention representation vector defined by fusing attention information and labels; The classification layer is used to classify the Chinese text The weighted context representation vector and metaphorical attention representation vector Perform splicing and input into the linear transformation layer to obtain the classification prediction result.
2. A Chinese text classification method according to claim 1, characterized in that: The Chinese text Mapping into metaphorical explanation sentences , and mapping classification labels to label definitions , expressed as: , , in, Indicates the generation of a large language model, Represents a prompt word template for generating context-based metaphor interpretation, Indicates the prompt word template for generating label definitions. represents a predefined set of classification labels, is the number of labels, This is the task description for the dataset.
3. A Chinese text classification method according to claim 1, characterized in that: In the representation layer, a shared BERT series encoder is used to transform the character sequence , metaphorical explanation sentence and tag definitions Processed into text embedding vectors , metaphor interpretation embedding vector and labels define embedding vectors .
4. A Chinese text classification method according to claim 1, characterized in that: The Chinese text The weighted context representation vector and metaphorical attention representation vector Perform splicing and input the linear transformation layer to obtain the classification prediction results, specifically: The context representation vector and metaphorical attention representation vector Splicing to form Chinese text The final representation vector , expressed as: ; Will represent the vector Input linear transformation layer, converted into a numerical vector of classification labels : , in, is the weight matrix, is the bias vector; Then the Softmax function is used to obtain the probability distribution of the predefined categories, and finally the label with the highest probability is selected As Chinese text The classification prediction result is expressed as: 。 5. A Chinese text classification method according to claim 1, characterized in that: In the process of training the metaphor association model, the loss function used is: ; in, is a proportional adjustment hyperparameter used to balance the cross entropy loss and triplet contrast loss The weight of Cross Entropy Loss , expressed as: ; in, is a dataset for classification tasks, Represented as Chinese text Assigning Tags Calculate the probability of ; in, is the similarity distance function, For the data set Chinese text The index of Represents the first Chinese text The context representation vector of , as a negative sample, belongs to the context representation vector ; Represents the first Chinese text The metaphorical interpretation embedding vector of , also called the anchor point in the ternary contrast loss, belongs to the metaphorical interpretation embedding vector ; Represents the first Chinese text The label definition representation vector of is a positive sample, which belongs to the label definition representation vector , is the marginal parameter.
Citation Information
Patent Citations
Emotion classification model construction method based on metaphor recognition
CN114942991A
Method for constructing sentiment classification model based on metaphor identification
US20230289528A1