Irony detection method based on multi-dimensional context modeling and dynamic self-checking
By employing multidimensional context modeling and dynamic self-verification, this method solves the problem of capturing implicit contradictions in complex contexts during satire detection, achieving high accuracy and reliability in identifying satirical texts. It can be applied to social media, public opinion analysis, and intelligent customer service systems.
Patent Information
- Application Number
- CN202511156095.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-18
- Publication Date
- 2025-11-07
AI Technical Summary
Existing methods for detecting irony struggle to effectively capture implicit contradictions and multi-dimensional contexts in complex situations, resulting in insufficient accuracy and reliability in detecting situational irony.
By employing multi-dimensional context modeling and dynamic self-verification methods, including multi-hop reasoning mechanisms, multi-modal cue decoupling, multi-dimensional context fusion, and dynamic self-verification, multi-dimensional contextual information is constructed to enhance the model's ability to understand satirical texts. Furthermore, the consistency of the reasoning path is verified through dynamic self-verification.
It significantly improves the accuracy and reliability of sarcasm detection, enabling the identification of sarcastic expressions in social media, public opinion analysis, and intelligent customer service systems, and providing more accurate user intent understanding and emotion monitoring capabilities.
Smart Images

Figure CN120911477A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to a sarcasm detection method based on multi-dimensional context modeling and dynamic self-checking, and belongs to the technical field of Internet and artificial intelligence. BACKGROUND
[0002] Sarcasm detection, as an important research direction in the field of natural language processing, aims to identify the implied semantic contrast and intent reversal in text through algorithmic models. Sarcasm detection relies on the analysis of deep language semantics, but the core challenge lies in capturing the contrast between literal meaning and true intent. For example, the sentence "Oh, great, another meeting! This is the perfect touch to my day!" superficially expresses positive emotions, but in the specific context (such as working for 12 hours), it can be inferred that the actual intent is to express negative emotions. The identification of such semantic reversal requires models to have multi-dimensional context understanding capabilities, including scenario inference, non-verbal cues (such as tone of voice, facial expressions) analysis, and deep mining of emotional states. Therefore, sarcasm detection is not only a language understanding task, but also a comprehensive test of complex context modeling capabilities, with important theoretical research value and broad practical application prospects.
[0003] Despite significant technological progress, the detection of contextual sarcasm still faces great challenges. Contextual sarcasm is often not directly manifested as a conflict between literal content and true meaning, making it difficult for models to identify its sarcastic expression (the same sentence may be interpreted as sarcastic or non-sarcastic in different contexts). Existing methods mainly rely on explicit emotional conflicts or obvious contrastive expressions, and these methods show poor detection results for more complex contextual sarcasm, as these sarcastic expressions often lack significant emotional conflicts. Traditional sarcasm detection models mainly rely on handcrafted feature extraction and rule systems, such as identifying the conflict between positive and negative emotions in text through contrastive sentiment analysis, or using social media tags (such as #sarcasm) for explicit labeling training, while combining context features in dialog scenes (such as the relationship between the speaker and the listener) to enhance detection capabilities. However, these methods are limited by surface feature engineering and simple classifiers, and are difficult to capture deep semantics and implicit contradictions in sarcastic contexts.
[0004] With the development of deep learning, pre-trained language models (PLMs) have significantly improved the sarcasm detection capability through task-specific fine-tuning and feature extraction techniques. For example, DC-Net introduces a dual-channel framework to model explicit and implicit semantics independently to better recognize emotional conflicts, while SarcPrompt optimizes the adaptability of pre-trained models to the sarcasm task through prompt tuning techniques. These models further enhance the modeling of sarcastic semantics in text by integrating diverse embedding techniques and semantic representation methods. However, the limitations of PLMs lie in their strong reliance on surface features, difficulty in dynamically generating complex reasoning paths, and tendency to miss key implicit contradictions when processing multi-hop contexts. Large language models (LLMs) leverage their large-scale pre-training and strong contextual understanding capabilities to significantly improve sarcasm detection accuracy through prompt engineering or thought chain techniques, especially when dealing with implicit contradictions and multi-modal contexts (such as combining visual instructions). Additionally, zero-shot and few-shot learning capabilities enable models to effectively work in low-resource scenarios, but the challenges of LLMs in terms of computational resource requirements and potential overfitting issues still need further optimization. For example, the high computational cost of LLMs limits their application in resource-constrained environments, and their overfitting to specific datasets may result in insufficient generalization capabilities in complex sarcasm scenarios.
[0005] Based on the limitations of existing research in sarcasm detection, especially for situational sarcasm detection, the invention introduces a sarcasm detection method based on multi-dimensional context modeling and dynamic self-checking. Through hierarchical context deduction and multi-modal cue decoupling techniques, it achieves deep mining of complex contexts. Unlike traditional methods, this invention does not rely on static external knowledge bases, but dynamically generates reasoning paths to gradually analyze the implicit meaning and subtle differences in emotions in the text, thereby more accurately capturing sarcastic intent. The invention integrates situational background, non-verbal cues, and emotional states into the reasoning process through structured multi-hop reasoning steps, significantly improving the detection capability of situational sarcasm with strong context dependence. At the same time, the invention combines a dynamic self-checking and cross-path consistency verification, which evaluates the consistency between paths and weights the aggregated results based on generating multiple reasoning paths, further reducing reasoning bias and improving model stability and prediction reliability.
[0006] More importantly, the method shows wide practical value in practical application. As a language phenomenon highly dependent on context, sarcasm widely exists in various texts such as social media, news comments, online forums, customer service dialogues, etc. Accurate identification of sarcasm not only helps to improve the understanding ability of natural language processing system to user's intention, but also plays an important role in the fields of public opinion analysis, brand reputation management, user emotion monitoring, intelligent customer service, etc. For example, in the social media platform, the application can be used to identify the sarcastic remarks of users in real time, and assist the platform to carry out more accurate public opinion guidance and content review; in enterprise public opinion monitoring, the method can help the brand party to identify potential negative emotions and avoid making wrong decisions due to misreading of user emotions; in the intelligent dialogue system, the improvement of sarcasm identification capability helps to enhance the semantic understanding ability of the dialogue system, so as to provide a more natural and more humanized interactive experience. In addition, the method can also be extended to multi-modal scenarios such as video comments, voice dialogues, etc., providing technical support for building a more comprehensive and more intelligent semantic understanding system.
[0007] Finally, the application builds a more accurate and reliable sarcasm language detection framework by integrating dynamic reasoning and self-consistency checking, effectively overcoming the problem of insufficient generalization ability of existing methods in complex context, not only promoting the development of sarcasm detection technology in the academic field, but also showing wide technical extensibility and landing possibility in industrial applications. SUMMARY
[0008] In order to solve the problems and deficiencies in the prior art, the application proposes a sarcasm detection method based on multi-dimensional context modeling and dynamic self-checking. In view of the problem that context dependence and implicit contradiction are difficult to capture in single-modal sarcasm detection, the application builds a detection framework from two dimensions of context reasoning and knowledge enhancement. The application introduces a multi-hop reasoning mechanism to gradually excavate semantic reversal, emotional contrast and sarcastic metaphor in the text, and combines an external knowledge base to build a multi-dimensional context representation, so as to make up for the deficiency of traditional methods in capturing implicit context. The application not only focuses on the explicit emotional conflict in the text, but also builds the emotional reversal path in the potential sarcastic expression through the multi-dimensional context reasoning chain, so as to comprehensively improve the modeling ability of complex sarcastic expression.
[0009] In order to achieve the above purpose, the technical scheme of the application is as follows: a sarcasm detection method based on multi-dimensional context modeling and dynamic self-checking, the method comprising the following steps:
[0010] Step 1: multi-modal corpus preprocessing and standardized modeling;
[0011] Step 2: hierarchical context deduction and multi-modal clue decoupling;
[0012]
[0012] Step 3: multi-dimensional context fusion and knowledge enhancement;
[0013] Step 4: Dynamic self-checking and cross-path consistency verification;
[0014] Step 5: Multi-task joint optimization and adaptive loss function tuning.
[0015] As an improvement of the present application, step 1: multi-modal corpus preprocessing and standardized modeling, first, through network crawler technology and data preprocessing method, the original corpus containing satirical semantics is obtained from the mainstream social network platform, and the collected data is cleaned, de-duplicated and standardized. Then, the preprocessed data set is divided into training set, validation set and test set according to the ratio of 8:1:1, among which the training set is used for model parameter learning, the validation set is used for hyperparameter tuning, and the test set is used for final performance evaluation. The core goal of this step is to identify and select text samples with typical satirical characteristics, providing high-quality data support for subsequent model training, thereby improving the model's ability to distinguish special satirical expression forms (such as implicit satire, situational satire, etc.).
[0016] As an improvement of the present application, step 2: hierarchical context deduction and multi-modal cue decoupling, the multi-hop reasoning of LLM is realized through the thinking chain prompting strategy. The core of this strategy is to decompose complex reasoning tasks into manageable sub-steps, thereby improving the prediction accuracy and detail of the model. In the satire detection scenario, this method guides the model to deduce the situational context (such as the dialogue scene), non-verbal cues (such as tone and body language), and the speaker's emotions, intentions and attitudes in stages, rather than directly outputting the prediction result. Through the staged reasoning process, the present application ensures that key elements such as context, non-verbal cues and emotional state are fully considered before generating conclusions, significantly enhancing the model's ability to capture explicit and implicit semantics, especially for satire semantic analysis tasks that rely on multi-dimensional context. The implementation of this step can be divided into the following sub-steps:
[0017] Sub-step 2-1, context reasoning. In this step, the model will be guided to analyze the possible logical context or environment in which the sentence is placed. By designing a specific prompt template {“[Given sentence X], in what context would this sentence be said?”}, the model is encouraged to consider a variety of contextual factors, including the physical environment, the relationship between the speaker and the listener, and the social norms and pragmatic background that govern the interaction. The first step can be represented as:
[0018] A=argmaxp(A|X)
[0019] Where A represents the inferred situational context generated according to the input sentence X and the structured prompt. The model uses its internal knowledge base and reasoning ability to infer the most likely context consistent with the text.
[0020] Sub-step 2-2, Non-verbal signal inference. In this step, the model will be guided to infer potential non-verbal signals in the text {“[Context], based on this context, what are the possible vocal tone, facial expression, body language?”}. These cues play an important role in sarcasm detection, as sarcasm often relies on the inconsistency between the literal expression and the actual conveyed meaning. For example, a negative statement with a cheerful tone is usually a clear signal of sarcasm. By analyzing such non-verbal features, the model can more accurately capture the emotional tone and implied meaning in the context, thus more effectively identifying the true intention of the speaker. The second step can be represented as:
[0021] B = argmax p(B | A, X)
[0022] wherein, wherein B represents the context text containing non-verbal cues, which is generated based on the second-step prompt template and the context inference A from the first step.
[0023] Sub-step 2-3, emotional state inference. In this step, the model will further analyze the speaker's psychological state based on the inference of non-verbal cues {“[Non-verbal signals], based on these non-verbal signals, what is the speaker's emotion, intention and attitude?”}. This process aims to enable the model to comprehensively capture potential emotional information and integrate it with the previously identified context background and non-verbal features. Through this step-by-step inference mechanism, the model can more deeply understand the true emotions behind the text, combining explicit expressions with implicit signals for comprehensive judgment, thus achieving more accurate and comprehensive analysis of sarcastic semantics. The third step can be represented as:
[0024] C = argmax p(C | X, H, B)
[0025] wherein, C represents the context text containing emotional states, generated by the third step of the prompt sequence. The dynamic dialogue history H is used in the inference process, which contains the current dialogue content as the context background; B represents the non-verbal cues extracted from the previous step. The dialogue history H is dynamically embedded throughout the inference process, affecting the understanding of the speaker's emotional state, enabling the model to make more accurate emotional inferences in combination with the dynamic changes in the dialogue.
[0026] As an improvement of the present application, step 3: multi-dimensional context fusion and knowledge enhancement, utilizes the multi-dimensional context obtained in step 2 to perform context expansion and semantic completion on the original text using the T5 model. By converting the inferred situational background, tone features, and speaker emotional state into structured text descriptions, and then concatenating and fusing them with the original input, an enhanced training sample with richer semantics and more complete context is constructed. This step aims to enrich the dataset with the output of the multi-hop thinking chain reasoning process in the knowledge enhancement phase. The original input sentence X is connected with the multi-dimensional context information (labeled as CoT), which contains the inferred context, non-verbal signals, and speaker emotions. The enhanced input is expressed as:
[0027] Y = Concat(X, CoT)
[0028] Where Y represents the input text after enrichment, X represents the original input sentence, and CoT contains multi-dimensional context information generated by a large language model. The function Concat() is used to integrate the original input with the context knowledge generated by the CoT process. To ensure the clarity and organization of the structured input, this concatenation process is further subdivided into specific components, and different context outputs are separated by special markers, as shown in the following equations:
[0029] Concat(X, CoT) = X || [EXT]CoT Ai || [EXT]CoT Bi || [EXT]CoT Ci
[0030] Where [EXT] represents a special marker used to separate different parts of the enhanced input, ensuring that the model can clearly process each component of the context information. The connection operator || links the original input sentence (X) with different CoT outputs:
[0031] Context information derived from situational reasoning provides insights into the possible scenarios or environments in which the sentence may occur.
[0032] Information related to non-verbal cues, such as inferred tone, facial expressions, and body language accompanying the sentence.
[0033] Information about the speaker's emotional state, inferred emotions, intentions, and attitudes through in-depth analysis of the sentence context.
[0034] As an improvement to this invention, step 4: Dynamic self-verification and cross-path consistency verification. Based on the design of the previous three steps, this invention proposes a dynamic self-verification and cross-path consistency verification mechanism. This mechanism can generate multiple inference paths from the same input, evaluate the consistency of conclusions from different perspectives, and finally synthesize the results of each path to generate a coherent and robust prediction. The implementation of this step can be divided into the following sub-steps:
[0035] Sub-step 4-1: Multi-level internal consistency verification and error analysis. After generating inference paths, the consistency between these paths is evaluated to ensure coordination and coherence across different perspectives. This process measures consistency between paths using probability distribution similarity. For each pair of inference paths P... i and P j The model calculates the KL divergence to quantify the differences between its predicted distributions:
[0036]
[0037] Specifically, P i (y|x) and P j (y|x) represent the predicted probability distributions of the i-th and j-th paths for the input x belonging to category y (e.g., "ironic" or "non-ironic"), respectively. D KL The KL divergence is always greater than or equal to 0, and equal to 0 only if the two distributions are completely identical. The lower the value, the more consistent the predictions of the two paths are. Generally, when the KL divergence is less than 0.5, the predicted distributions of the two paths can be considered highly similar; when the KL divergence is greater than 1.0, it indicates a significant difference between them. To evaluate the overall consistency among all inference paths, the model calculates the average KL divergence as the mutual consistency score C:
[0038]
[0039] Where N represents the total number of inference paths. The magnitude of the C value directly reflects the degree of coordination in the entire inference process: a lower C value (usually close to 0) indicates a high degree of consistency among the paths and a more robust model judgment; while a higher C value (e.g., greater than 1.0) suggests significant discrepancies between different inference paths, which may stem from ambiguity in the input context or uncertainty in the model's inference process. In this case, corresponding verification or fusion mechanisms will be activated to improve the reliability of the final prediction.
[0040] Sub-step 4-2: Multimodal inference path fusion and collaborative optimization. Once diverse inference sequences are generated, they are then evaluated by the T5 model. For each inference path i, its prediction result is determined by the prediction probability distribution P. i (y|x) represents this. To evaluate the reliability of this path, we base it on P. i(y|x) computes a confidence score, Confidence, reflecting the model's certainty about the accuracy of a particular explanation. The confidence score is defined as the maximum value in the predicted probability distribution:
[0041] Confidence = max(P(y|x))
[0042] The weight α assigned to the i-th reasoning path is determined based on its predicted probability distribution P i (y|x). This value reflects the model's confidence in the effectiveness of this path. More specifically, α i can be computed as follows: i
[0043]
[0044] To ensure that the weights are normalized so that their sum equals 1, i.e., ∑α i = 1 (i from 1 to N), and to assign greater weights to reasoning paths with higher confidence. In this way, the aggregated score S(y|x) can reflect the relative certainty of each path, allowing the model to synthesize a more reliable prediction based on the reasoning sequence it considers most persuasive. However, the model does not rely solely on the single path with the highest confidence, but rather aggregates the results by marginalizing multiple reasoning paths and selecting the final prediction with the highest consistency. Its synthesis process can be expressed as:
[0045]
[0046] where S(y|x) represents the synthesized score for class y, P i (y|x) represents the predicted probability distribution of the i-th reasoning path, and α i represents the weight assigned based on the path's certainty (satisfying ∑α i = 1).
[0047] As an improvement of the present application, step 5: multi-task joint optimization and adaptive loss function tuning, in this step, we use a batch size of 4, a learning rate of 1e-5, and use the AdamW optimizer to perform gradient backpropagation to update the model parameters; the initial learning rate is set to 2e-5, which is gradually reduced to 1e-5, and the T5-Base model is fine-tuned. This strategy not only effectively prevents overfitting, but also significantly improves the model's generalization ability on small data sets, allowing it to more accurately capture the nuances of sarcastic expressions in different contexts. In this way, the present application enhances the model's understanding and detection ability of complex sarcastic language, ensuring its high accuracy and reliability in various application scenarios.
[0048] An electronic device comprises a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the sarcasm detection method based on multi-dimensional context modeling and dynamic self-checking when executing the program.
[0049] A storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the sarcasm detection method based on multi-dimensional context modeling and dynamic self-checking.
[0050] Compared with the prior art, the advantages of the present application are as follows:
[0051] (1) The present application proposes a sarcasm detection method based on multi-dimensional context modeling and dynamic self-checking, which generates multi-dimensional context information including scene context, non-verbal cues and emotional state through hierarchical context deduction and multi-modal cue decoupling module, effectively enhances the understanding ability of the model to sarcastic text, provides key background knowledge for identifying sarcastic expressions, and thus more accurately captures its implied meaning.
[0052] (2) The present application further integrates multi-dimensional context fusion and knowledge enhancement module to improve the learning effect of pre-trained model on sarcastic semantics, and introduces dynamic self-checking and cross-path consistency verification to verify multiple reasoning paths. This comprehensive technical solution enables the model to analyze explicit and implicit emotional signals more carefully, significantly improving the accuracy and reliability of sarcasm detection.
[0053] (3) The present application significantly improves the accuracy and reliability of sarcasm detection through multi-dimensional context modeling and dynamic self-checking. Its application scenarios are wide, including social media monitoring, public opinion analysis and intelligent customer service system, etc. In social media, this method can identify sarcastic expressions in user-generated content in real time, helping brands and enterprises to understand public sentiment changes in a timely manner; in the field of public opinion analysis, it supports government agencies and research teams to make accurate policy and crisis management; in the intelligent customer service system, the sarcasm detection function of this invention can improve the understanding ability of the dialogue system and provide a more natural and smooth user experience. In addition, the present application also has scalability and is suitable for multi-modal data processing, and can efficiently run in resource-constrained environments, providing intelligent and humanized solutions for various intelligent systems. BRIEF DESCRIPTION OF DRAWINGS
[0054] Figure 1 The overall model diagram of the embodiment of the present application;
[0055] Figure 2 The method processing flowchart of the embodiment of the present application. DETAILED DESCRIPTION
[0056] In order to deepen the understanding and understanding of the present application, the following specific examples are further illustrated.
[0057] Example 1: A satire detection method based on multi-dimensional context modeling and dynamic self-checking. In order to solve the problem that the context dependence and implicit contradiction in single modal satire detection are difficult to capture, the present application constructs a detection framework from two dimensions of context reasoning and knowledge enhancement. The present application gradually excavates the semantic reversal, emotional contrast and satire metaphor in the text by introducing a multi-hop reasoning mechanism, and combines an external knowledge base to build a multi-dimensional context representation, so as to make up for the deficiency of traditional methods in capturing implicit context. The present application not only focuses on the explicit emotional conflict in the text, but also constructs the emotional reversal path in the potential sarcastic expression through the multi-dimensional context reasoning chain, thereby comprehensively improving the modeling ability of complex sarcastic expression.
[0058] The specific model is shown in Figure 1 , and the detailed implementation steps are as follows:
[0059] Step 1: Multimodal corpus preprocessing and standardized modeling. First, through network crawler technology and data preprocessing method, the original corpus containing sarcastic semantics is obtained from the mainstream social network platform, and the collected data is cleaned, de-duplicated and standardized. Then, the preprocessed data set is divided into training set, validation set and test set according to the ratio of 8:1:1, wherein the training set is used for model parameter learning, the validation set is used for hyperparameter optimization, and the test set is used for final performance evaluation. The core goal of this step is to identify and select text samples with typical sarcastic characteristics, providing high-quality data support for subsequent model training, thereby improving the model's ability to discriminate special sarcastic expression forms (such as implicit sarcasm, situational sarcasm, etc.).
[0060] Step 2: Hierarchical context deduction and multimodal cue decoupling. The multi-hop reasoning of LLM is realized through the thinking chain prompting strategy. The core of this strategy is to decompose complex reasoning tasks into manageable sub-steps, thereby improving the prediction accuracy and detail of the model. In the satire detection scene, this method guides the model to deduce the context (such as dialogue scene), non-verbal cues (such as tone and body language), and the speaker's emotion, intention and attitude in stages, rather than directly outputting the prediction result. Through the staged reasoning process, the present application ensures that key elements such as context, non-verbal cues and emotional state are fully considered before the conclusion is generated, significantly enhancing the model's ability to capture explicit and implicit semantics, especially for satire semantic analysis tasks that rely on multi-dimensional context. The implementation of this step can be divided into the following sub-steps:
[0061] Sub-step 2-1, Context Reasoning. In this step, the model will be guided to analyze the possible scenario or environment in which the sentence logically resides. By designing a specific prompt template {“[Given sentence X], in what context would this sentence be said?”}, the model is encouraged to consider a variety of contextual factors, including the physical environment, the relationship between the speaker and the listener, and the social norms and pragmatic background that govern the interaction. The first step can be represented as:
[0062] A = argmax p(A|X)
[0063] where A represents the inferred contextual scenario generated from the input sentence X and the structured prompt. The model utilizes its internal knowledge base and reasoning capabilities to infer the most probable context that aligns with the text.
[0064] Sub-step 2-2, Non-verbal Signal Reasoning. In this step, the model will be guided to infer potential non-verbal signals {“[Context], based on this context, what are the possible vocal tone, facial expressions, body language?”} in the text. These cues play a significant role in sarcasm detection, as sarcasm often relies on the discrepancy between the literal expression and the actual conveyed meaning. For example, a negative statement with a cheerful tone is often a clear signal of sarcasm. By analyzing such non-verbal features, the model can more accurately capture the emotional undertones and implied meanings in the context, thereby more effectively identifying the speaker's true intent. The second step can be represented as:
[0065] B = argmax p(B|A,X)
[0066] where B represents the contextual text containing non-verbal cues, generated based on the second-step prompt template and the context reasoning A from the first step.
[0067] Sub-step 2-3, Emotional State Reasoning. In this step, the model will further analyze the speaker's psychological state {“[Non-verbal signals], based on these non-verbal signals, what are the speaker's emotions, intentions, and attitudes?”} on the basis of the inferred non-verbal cues. This process aims to enable the model to comprehensively capture potential emotional information and integrate it with the previously identified contextual background and non-verbal features. Through this step-by-step reasoning mechanism, the model can gain a deeper understanding of the true emotions behind the text, combining explicit expressions with implicit signals for a comprehensive judgment, thereby achieving a more accurate and comprehensive analysis of sarcastic semantics. The third step can be represented as:
[0068] C = argmax p(C|X,H,B)
[0069] Where C represents the context text containing the emotional state, generated by the third step of the prompt sequence. A dynamic dialogue history H is used in the reasoning process, which contains the current dialogue content as a situational background; B represents the non-verbal cues extracted from the previous step. The dialogue history H is dynamically embedded throughout the reasoning process, affecting the understanding of the speaker's emotional state, enabling the model to make more accurate emotional inferences in combination with the dynamic changes in the dialogue.
[0070] Step 3: Multi-dimensional context fusion and knowledge enhancement, using the multi-dimensional context obtained in step 2, using the T5 model to perform context expansion and semantic completion on the original text. By converting the inferred situational background, tone features, and speaker emotional state into structured text descriptions, and then concatenating and fusing them with the original input, an enhanced training sample with richer semantics and more complete context is constructed. This step aims to use the output of the multi-hop thinking chain reasoning process to enrich the dataset in the knowledge enhancement phase. The original input sentence X is connected with the multi-dimensional context information (labeled as CoT), which contains the inferred context, non-verbal signals, and speaker emotions. The enhanced input is expressed as:
[0071] Y = Concat(X, CoT)
[0072] Where Y represents the input text after enrichment, X represents the original input sentence, and CoT contains the multi-dimensional context information generated by the large language model. The function Concat() is used to integrate the original input with the context knowledge generated by CoT. To ensure the clarity and organization of the structured input, this concatenation process is further subdivided into specific components, and different context outputs are separated by special markers, as shown in the following equation:
[0073] Concat(X, CoT) = X || [EXT] CoT Ai || [EXT] CoT Bi || [EXT] CoT Ci
[0074] Where [EXT] represents a special marker used to separate different parts of the enhanced input, ensuring that the model can clearly process each component of the context information. The connection operator || links the original input sentence (X) with different CoT outputs:
[0075] Context information derived from situational reasoning, providing insights into the possible scenarios or environments of the sentence.
[0076] Information related to non-verbal cues, such as inferred tone, facial expressions, and body language accompanying the sentence.
[0077] Information about the speaker's emotional state, inferred through in-depth analysis of the context of the sentence, emotions, intentions, and attitudes.
[0078] Step 4: Dynamic self-checking and cross-path consistency verification. Based on the design of the previous three steps, the present invention proposes a dynamic self-checking and cross-path consistency verification mechanism. This mechanism can generate multiple reasoning paths from the same input, evaluate the consistency of conclusions from different perspectives, and ultimately synthesize the results of each path to generate a coherent and robust prediction. The implementation of this step can be divided into the following sub-steps:
[0079] Sub-step 4-1, multi-level internal consistency verification and error analysis. After generating reasoning paths, the consistency between these paths is evaluated to ensure the coordination and coherence between different perspectives. This process measures the consistency between paths through probability distribution similarity. For each pair of reasoning paths P i and P j , the model calculates the KL divergence to quantify the difference between their prediction distribution:
[0080]
[0081] Specifically, P i (y|x) and P j (y|x) represent the prediction probability distribution of the i-th path and the j-th path for the input x belonging to the category y (e.g., "satire" or "non-satire"). D KL is always greater than or equal to 0, and is equal to 0 only when the two distributions are exactly the same. The lower the value, the more consistent the predictions of the two paths, and generally when the KL divergence is less than 0.5, the prediction distributions of the two paths are considered highly similar; when the KL divergence is greater than 1.0, it indicates that there are significant differences between them. To evaluate the overall consistency between all reasoning paths, the model calculates the average KL divergence as the mutual consistency score C:
[0082]
[0083] where N represents the total number of reasoning paths. The value of C directly reflects the coordination degree of the entire reasoning process: a lower C value (usually close to 0) indicates high consistency between paths, and the model's judgment is more robust; while a higher C value (e.g., greater than 1.0) indicates that there are significant differences between different reasoning paths, which may be due to the ambiguity of the input context or uncertainty in the model's reasoning process, at which point the corresponding verification or fusion mechanism will be activated to improve the reliability of the final prediction.
[0084] Sub-step 4-2: Multimodal inference path fusion and collaborative optimization. Once diverse inference sequences are generated, they are then evaluated by the T5 model. For each inference path i, its prediction result is determined by the prediction probability distribution P. i (y|x) represents this. To evaluate the reliability of this path, we base it on P. i (y|x) calculates a confidence score, which reflects the degree of certainty that the model is accurate for a particular explanation. This confidence score is defined as the maximum value in the predicted probability distribution:
[0085] Confidence = max(P(y|x))
[0086] The weight α assigned to the i-th inference path i Based on its predicted probability distribution P i (y|x) is determined, and this value reflects the model's confidence in the effectiveness of the path. More specifically, α i It can be calculated in the following ways:
[0087]
[0088] To ensure that the weights are normalized so that their sum equals 1, i.e., ∑α i =1 (i from 1 to N), and assign greater weight to inference paths with higher confidence. In this way, the aggregate score S(y|x) can reflect the relative certainty of each path, allowing the model to synthesize more reliable predictions based on the inference sequence it considers most convincing. However, the model does not rely solely on the single path with the highest confidence, but aggregates results by marginalizing multiple inference paths and selects the final prediction with the highest consistency. This synthesis process can be described as follows:
[0089]
[0090] Where S(y|x) represents the composite score of category y, P i (y|x) represents the predicted probability distribution of the i-th inference path, α i This represents the weights assigned based on the determinism of the path (satisfying ∑α). i =1).
[0091] Step 5: Multi-task joint optimization and adaptive loss function tuning, in this step, we use the batch size of 4, learning rate of 1e-5, and use AdamW optimizer for gradient back propagation to update the model parameters; the initial learning rate is set to 2e-5, gradually reduced to 1e-5, and the T5-Base model is fine-tuned. This strategy not only effectively prevents overfitting, but also significantly improves the model's generalization ability on small data sets, enabling it to more accurately capture the nuances of sarcastic expressions in different contexts. In this way, the present application enhances the model's understanding and detection ability of complex sarcastic language, ensuring its high accuracy and reliability in various application scenarios.
[0092] Based on the same inventive concept, the sarcasm detection method and device based on multi-dimensional context modeling and dynamic self-checking of the present application include a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is loaded into the processor, it implements the above-mentioned sarcasm detection method based on multi-dimensional context modeling and dynamic self-checking.
[0093] Those skilled in the art will appreciate that the embodiments described herein are intended to aid the reader in understanding the principles of the present application and should not be construed as limiting the scope of the present application, which is defined by the claims. After reading the present application, those skilled in the art will make various modifications to the present application, which fall within the scope of the claims.
Claims
1. A sarcasm detection method based on multi-dimensional context modeling and dynamic self-checking, characterized in that, The method comprises the following steps: Step 1: Multimodal corpus preprocessing and standardized modeling; Step 2: Hierarchical context deduction and multimodal clue decoupling; Step 3: Multidimensional context fusion and knowledge enhancement; Step 4: Dynamic self-checking and cross-path consistency verification; Step 5: Multi-task joint optimization and adaptive loss function optimization.
2. The sarcasm detection method based on multi-dimensional context modeling and dynamic self-checking according to claim 1, characterized in that, Step 1: Multimodal corpus preprocessing and standardized modeling, first, through network crawler technology and data preprocessing method, the original corpus containing satirical semantics is obtained from the mainstream social network platform, and the collected data is cleaned, de-duplicated and standardized, then the preprocessed data set is divided into training set, validation set and test set according to the ratio of 8:1:1, wherein the training set is used for model parameter learning, the validation set is used for hyperparameter optimization, and the test set is used for final performance evaluation.
3. The sarcasm detection method based on multi-dimensional context modeling and dynamic self-checking according to claim 1, characterized in that, Step 2: Hierarchical context deduction and multimodal clue decoupling, the multi-hop reasoning of LLM is realized through thinking chain prompting strategy, including the following sub-steps: Sub-step 2-1, context reasoning, in this step, the model will be guided to analyze the scene or environment that the sentence may be in logically, by encouraging the model to consider a variety of contextual factors, including the physical environment, the relationship between the speaker and the listener, and the social norms and pragmatic background that govern the interaction, the first step is represented as: A=argmaxp(A|X) Wherein, A represents the context of the scene generated according to the original input sentence X and the structured prompt, the model uses its internal knowledge base and reasoning ability to infer the most likely context consistent with the text; Sub-step 2-2, non-verbal signal reasoning, in this step, the model will be guided to infer the potential non-verbal signals in the text, the second step is represented as: B=argmaxp(B|A,X) Wherein, B represents the context text containing non-verbal clues, which is generated based on the second step of the prompt template and the context of the scene A from the first step; Sub-step 2-3, emotion state reasoning, the third step is represented as: C=argmaxp(C|X,H,B) Wherein, C represents the context text containing emotional state, which is generated by the third step of the prompt sequence, the dynamic dialogue history H is used in the reasoning process, which contains the current dialogue content as the context background; B represents the context text containing non-verbal clues, the dialogue history H is dynamically embedded in the whole reasoning process, which affects the understanding of the speaker's emotional state, so that the model can make more accurate emotional inference combined with the dynamic changes of the dialogue.
4. The sarcasm detection method based on multi-dimensional context modeling and dynamic self-checking according to claim 1, characterized in that, Step 3: Multidimensional context fusion and knowledge enhancement, the enhanced input is represented as: Y=Concat(X,CoT) Where Y represents the input text after enrichment, X represents the original input sentence, CoT contains multidimensional context information generated by large language model, and function Concat() is used to integrate the original input and the context knowledge generated by CoT process, in order to ensure the clarity and orderliness of the structured input, the splicing process is further subdivided into specific components, and different context outputs are separated by special markers, as shown in the following equation: Concat(X, CoT) = X || [EXT] CoT Ai ||[EXT] CoT Bi ||[EXT] CoT Ci where [EXT] represents a special token used to separate different parts of the enhanced input, ensuring that the model can clearly process each component of the context information, and the connection operator || links the original input sentence X with different CoT outputs: Contextual information derived from situational reasoning provides insights into the possible scenarios or environments in which a sentence may appear. information related to non-verbal cues such as inferred tone of voice, facial expression, and body language accompanying the sentence, information about the speaker's emotional state, inferred emotions, intentions, and attitudes through in-depth analysis of the sentence context.
5. The sarcasm detection method based on multi-dimensional context modeling and dynamic self-checking according to claim 1, characterized in that, Step 4: Dynamic self-checking and cross-path consistency verification, specifically including the following sub-steps: Sub-step 4-1, multi-level internal consistency check and error analysis, after generating the inference paths, the consistency between these paths is evaluated to ensure the coordination and coherence between different perspectives. The consistency between paths is measured by the similarity of probability distributions. For each pair of inference paths P i and P j , the model calculates the KL divergence to quantify the difference between their predicted distributions: In particular, P i (y | x) and P j (y | x) represent the prediction probability distribution of the i-th path and the j-th path for the input x belonging to the class y, respectively, D KL is always greater than or equal to 0, and is equal to 0 only when the two distributions are exactly the same, the lower the value, the more consistent the predictions of the two paths tend to be, and generally when the KL divergence is less than 0.5, the prediction distributions of the two paths can be considered highly similar; when the KL divergence is greater than 1.0, it indicates that there is a significant difference between the two. To evaluate the overall consistency between all reasoning paths, the model calculates the average KL divergence as the mutual consistency score C: where N represents the total number of inference paths, and the size of C directly reflects the coordination degree of the entire inference process, Sub-step 4-2, multimodal inference path fusion and collaborative optimization, once the diversified inference sequences are generated, they are then submitted to the T5 model for evaluation, and for each inference path i, its prediction result is represented by a prediction probability distribution P i (y|x) indicates that, in order to evaluate the reliability of the path, a confidence score Confidence is calculated based on P i (y|x) to reflect the model's determination of the accuracy of a specific explanation, and the confidence score is defined as the maximum value in the prediction probability distribution: Confidence=max(P(y|x)) The weight α assigned to the i-th inference path i is determined based on the predictive probability distribution P i (y | x) of the model, which reflects the degree of certainty of the model about the validity of the path, more specifically, α i is calculated by To ensure that the weights are normalized so that their sum equals 1, i.e.∑α i = 1 (i from 1 to N), and to assign greater weights to inference paths with higher confidence, in this way, the aggregated score S(y | x) can reflect the relative certainty of each path, enabling the model to synthesize a more reliable prediction result based on the inference sequence it considers most persuasive, the synthesis process is expressed as: where S(y | x) denotes the synthetic score for class y, P i (y | x) represents the predictive probability distribution of the i-th inference path, a i denotes the weight assigned based on the certainty of the path, satisfying∑a i = 1.
6. The sarcasm detection method based on multi-dimensional context modeling and dynamic self-checking according to claim 1, characterized in that, Step 5: Multi-task joint optimization and adaptive loss function tuning, using a mini-batch size of 4, a learning rate of 1e-5, and using the AdamW optimizer for gradient backpropagation to update the model parameters; the initial learning rate is set to 2e-5 and gradually reduced to 1e-5, and the T5-Base model is fine-tuned.
7. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that: The processor executes the program to implement the sarcasm detection method based on multi-dimensional context modeling and dynamic self-checking according to any one of claims 1 to 6.
8. A storage medium having stored thereon a computer program, characterized in that: The computer program is executed by the processor to implement the sarcasm detection method based on multi-dimensional context modeling and dynamic self-checking according to any one of claims 1 to 6.
Citation Information
Cited By
Emotion analysis method and system based on big language model reasoning chain generation
CN121524357A