A multi-modal collaborative evolution irony recognition method and device based on particle computing
By combining granular computing theory with multimodal data and employing multi-granularity partitioning and collaborative evolution mechanisms, the problem of unstable recognition results in multimodal irony recognition is solved, achieving comprehensive recognition and accurate understanding of ironic expressions.
Patent Information
- Application Number
- CN202511383455.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-26
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2045-09-26
AI Technical Summary
Existing technologies struggle to effectively process multimodal ironic information, lacking deep modeling of the complementarity and synergy of cross-modal information. This results in unstable recognition performance and insufficient accuracy, failing to capture subtle differences and deep semantics in ironic expressions. Furthermore, traditional methods lack a co-evolutionary mechanism when fusing multimodal features, affecting the robustness and accuracy of the recognition system.
A multimodal co-evolutionary irony recognition method based on granular computing is adopted. Through multi-granularity partitioning and co-evolution mechanism, multimodal multi-granularity knowledge representation is constructed, the optimal granularity level is selected, and a co-operative neural network is used for recognition to generate irony recognition results.
It improves the robustness and accuracy of irony recognition, enhances the ability to capture subtle features of ironic expressions, significantly improves the accuracy and generalization of recognition, and provides an innovative technical approach for multimodal fusion analysis.
Smart Images

Figure CN120892869B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of natural language processing technology, specifically to a multimodal collaborative evolutionary irony recognition method and apparatus based on granular computing. Background Technology
[0002] In today's booming natural language processing technology, irony recognition, as a key branch of sentiment computing and semantic understanding, is increasingly valued for its research and application. Irony, as a special form of linguistic expression, often manifests as a significant discrepancy between literal meaning and true intent. This "inconsistency" in semantics makes traditional surface-level text matching-based recognition strategies highly prone to failure. With the explosive growth of social media and short video platforms, the dissemination of ironic information has rapidly evolved from a single textual form to a complex interweaving of multiple modal elements, including text, images, voiceovers, emoticons, editing rhythm, and even visual atmosphere. This significantly increases the complexity of irony recognition tasks; a seemingly complimentary text paired with exaggerated expressions, a distinctive tone, or a contrasting image can create a strong ironic effect. However, most existing technologies remain at a superficial level of analyzing isolated semantics in text or single visual symbols in images, lacking deep modeling of the complementarity and synergy of cross-modal information. This makes it difficult to effectively process the rich information contained in multimodal data, leading to frequent misidentification and omissions in real-world open scenarios. For example, on social media platforms, users often express sarcasm through text with images or short videos, and relying solely on single-modal analysis makes it difficult to accurately understand their true meaning. On the other hand, the lack of multi-level and multi-angle analysis of data during feature extraction and representation learning prevents the model from fully capturing the subtle differences and deeper semantics in sarcasm expressions.
[0003] Meanwhile, irony recognition also suffers from incomplete knowledge representation and insufficient feature fusion. Traditional methods often employ fixed feature extraction strategies, making it difficult to adapt to the diversity of ironic expressions in different scenarios. Regarding multimodal information fusion, existing technologies lack effective co-evolutionary mechanisms and cannot fully utilize the complementary relationships between different modalities, which severely affects the stability and accuracy of recognition results. Specifically, ironic expressions naturally possess distinct context-dependent and dynamic evolutionary characteristics: the same sentence may exhibit diametrically opposed emotional polarities under different cultural backgrounds, narrative contexts, or emotional tones; and the rapid evolution of internet slang and meme culture further fragments and instantaneously changes ironic patterns. Traditional schemes based on fixed feature templates or single-granularity representation learning struggle to capture these subtle differences, often resulting in rigid knowledge representation and insufficient generalization ability. Furthermore, in the multimodal feature fusion stage, existing methods generally adopt simple strategies such as early splicing or late voting, ignoring the dynamic changes in the weights of each modality feature at different granularity levels, and lacking an effective collaborative evolution mechanism to correct the confidence of each channel in real time, thereby limiting the robustness and accuracy of the overall recognition system.
[0004] Even more challenging is the fact that ironic expressions are highly context-dependent and dynamically changing, requiring models with strong contextual understanding capabilities. Ironic signals often exhibit rhythmic changes in the temporal dimension, such as "rise followed by fall" or "feigned indifference," necessitating models with the ability to model long-distance dependencies across modalities and granularities. However, existing deep learning frameworks perform poorly in handling such high-order cognitive tasks (i.e., long-distance dependencies and dynamic contextual changes), still facing bottlenecks such as poor interpretability, sensitivity to noise, and massive training data requirements, making it difficult to accurately grasp the semantic transformations and emotional changes in ironic expressions. These problems severely restrict the practical application of irony recognition technology. Therefore, how to adaptively mine the most discriminative granular features from multimodal heterogeneous data and dynamically fuse evidence from various channels through co-evolutionary reasoning has become a core challenge urgently needing breakthroughs in the field of irony recognition.
[0005] In view of the above, this application is hereby submitted. Summary of the Invention
[0006] This invention provides a multimodal cooperative evolution irony recognition method and apparatus based on granular computing, which can at least partially improve the above-mentioned problems.
[0007] To achieve the above objectives, the present invention adopts the following technical solution:
[0008] A multimodal cooperative evolutionary irony recognition method based on granular computing, comprising:
[0009] Extract sample data of different modalities from a pre-defined sample source library, and construct a multimodal satirical dataset based on the sample data;
[0010] A hierarchical granularity analysis and partitioning process is performed on the multimodal irony dataset to construct a multimodal, multigranular knowledge representation;
[0011] The dependency of each mode is calculated based on the multimodal multigranularity knowledge representation, and the optimal granularity levels of each mode are selected to obtain the multimodal optimal granularity matrix.
[0012] Extract the prototype mode vector and test mode vector from the multimodal optimal granularity matrix, and construct the irony order parameter based on the prototype mode vector and test mode vector;
[0013] A collaborative neural network model for irony order parameters is used to perform collaborative recognition preprocessing on the order parameters to generate irony recognition results.
[0014] The present invention also provides a multimodal cooperative evolution irony recognition device based on granular computing, comprising:
[0015] The data processing unit is used to extract sample data of different modalities from a preset sample source library and construct a multimodal irony dataset based on the sample data;
[0016] Multi-granularity units are used to perform hierarchical granularity analysis and partitioning of multimodal irony datasets to construct multimodal multi-granularity knowledge representations;
[0017] The optimal granularity unit is used to calculate the dependency of each mode based on the multimodal multigranularity knowledge representation, select the optimal multiple granularity levels for each mode, and obtain the multimodal optimal granularity matrix.
[0018] The order parameter unit is used to extract the prototype mode vector and test mode vector from the multimodal optimal granularity matrix, and construct the ironic order parameter based on the prototype mode vector and test mode vector;
[0019] The collaborative recognition unit is used to perform collaborative recognition preprocessing on the order parameter using a collaborative neural network model of the irony order parameter, and generate irony recognition results.
[0020] In summary, the proposed multimodal co-evolutionary irony recognition method based on granular computing combines granular computing theory with multimodal data. Through multi-granularity partitioning and a co-evolutionary mechanism, it enhances the robustness and accuracy of irony recognition, providing a technical solution for intelligent irony understanding. Specifically, this method integrates multimodal data such as text, images, audio, and video, and combines hierarchical granularity analysis to achieve comprehensive recognition of irony samples. This method innovatively applies granular computing theory for multimodal knowledge representation, selects the optimal granularity combination through dependency functions, constructs irony order parameters, and uses a collaborative neural network for reasoning. Finally, it derives the recognition result based on weight fusion, effectively solving the limitations of traditional single-modal recognition methods.
[0021] This technical solution fully leverages the complementarity and synergy of multimodal data, capturing subtle features of ironic expressions through multi-granularity modeling and avoiding information loss that may result from single-granularity analysis. Simultaneously, the co-evolutionary mechanism effectively integrates the recognition results from different modalities, significantly enhancing the accuracy and generalization ability of irony recognition. This method not only achieves excellent performance in irony recognition tasks but also provides an innovative technical approach for multimodal fusion analysis in the field of natural language processing, possessing significant theoretical and practical value. Attached Figure Description
[0022] Figure 1 This is a flowchart illustrating the multimodal cooperative evolution irony recognition method based on granular computing provided in the first embodiment of the present invention;
[0023] Figure 2 This is a schematic diagram of multimodal multi-granularity modeling and optimal granularity selection provided in an embodiment of the present invention;
[0024] Figure 3 This is a schematic diagram of cooperative evolution provided in an embodiment of the present invention;
[0025] Figure 4 This is a schematic diagram of a multimodal collaborative evolution irony recognition device based on granular computing provided in the second embodiment of the present invention. Detailed Implementation
[0026] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0027] refer to Figures 1 to 3 As shown, the first embodiment of the present invention discloses a granular computing-based multimodal co-evolutionary irony recognition method, which can be executed by a granular computing-based multimodal co-evolutionary irony recognition device (hereinafter referred to as the recognition device), specifically, by one or more processors within the recognition device, to implement the following method:
[0028] S1, extract sample data of different modalities from a pre-defined sample source library, and construct a multimodal irony dataset based on the sample data;
[0029] Specifically, step S1 further includes: extracting sample data of different modalities from a preset sample source library to obtain multimodal sample data, wherein the modalities include text, images, audio, and video, and the preset sample source library can be well-known talk shows, such as stand-up comedy performances, variety shows, etc.
[0030] Multimodal sample data is preprocessed, and the preprocessed data is then annotated with irony labels. The annotation results are divided into two categories: irony and non-irony. The preprocessing includes word segmentation, stop word removal, and standardization for text data; size normalization, illumination adjustment, face detection and alignment for image data; noise reduction, volume standardization, and frame segmentation for audio data; and frame rate unification, resolution adjustment, and keyframe extraction for video data.
[0031] The labeled multimodal sample data are aligned according to the timestamps to construct a multimodal satirical dataset.
[0032] In this embodiment, the sample source library interface is first activated to read the original text, images, audio, and video streams of stand-up comedy specials and variety shows in batches. The text stream is processed by word segmentation, stop word removal, and standardization to obtain clean segments. The image stream is processed by size normalization, illumination adjustment, face detection, and alignment to obtain normalized frames. The audio stream is processed by noise reduction, volume standardization, and frame segmentation to obtain a unified energy spectrum. The video stream is processed by frame rate unification, resolution adjustment, and keyframe extraction to obtain a compact frame sequence. Subsequently, the four-modal data is manually annotated with irony / non-irony binary labels and precisely aligned according to timestamps to form a structured multimodal irony dataset for subsequent granular analysis and co-evolutionary recognition.
[0033] Specifically, for text data, Chinese word segmentation tools such as jieba were used to remove punctuation marks, stop words such as numbers, and to convert the text to lowercase. The text length was also standardized. For image data, all images were scaled to 224×224 pixels, histogram equalization was used to adjust image brightness and contrast, and OpenCV was used for face detection, with the detected face regions aligned. For audio data, bandpass filters were used to remove ambient noise, the volume was standardized to a uniform decibel level, and the audio signal was framed using a fixed window length of 20ms. For video data, a uniform frame rate of 30fps was used, the resolution was adjusted to 1280×720, and a frame was extracted every 0.5 seconds as a keyframe. Furthermore, the preprocessed raw samples were manually annotated, including whether the sample contained ironic expressions (-0 for yes / -1 for no), thus establishing a corpus with both sentiment and metaphor annotations.
[0034] Next, a unified data indexing system was established, assigning a unique identifier ID to each sample. Then, data from different modalities were aligned according to timestamps to ensure temporal synchronization of text, image, audio, and video data. For each sample, its text content, corresponding image frames, audio clips, and video clips were organized into a structured data format, storing the data paths and annotation information for each modality in JSON format. Finally, the processed dataset was randomly divided into training and validation sets in a 7:3 ratio to ensure a roughly consistent ratio of positive to negative samples in both datasets, and the partitioning results were saved as a configuration file.
[0035] S2, hierarchical granularity analysis and partitioning of the multimodal irony dataset are performed to construct a multimodal multigranular knowledge representation;
[0036] Specifically, step S2 further includes: extracting features from the multimodal irony dataset to obtain feature spaces for each modality, and obtaining different feature subsets from each feature space through feature clustering technology to construct knowledge representations with different granularities;
[0037] The multimodal satirical dataset is partitioned based on different feature subsets to obtain multi-granularity feature subspaces for each modality. The multi-granularity feature subspaces for each modality are then aligned to construct a multimodal multi-granularity knowledge representation.
[0038] In this embodiment, firstly, for the text modality, the system extracts text sequences from the dataset one by one after word segmentation and standardization, and maps them into a continuous vector space using a pre-trained language model. In this vector space, the system performs unsupervised feature clustering, automatically dividing the text features into several non-overlapping feature subsets. Each subset corresponds to a knowledge representation at a specific granularity, such as word-level, phrase-level, or sentence-level. For the image modality, the system extracts color, texture, and shape features from each keyframe to form a high-dimensional feature vector. The same clustering algorithm is then run in this feature space to obtain pixel-level, region-level, and whole-image-level granular descriptions. For the audio modality, the system extracts Mel-frequency cepstral coefficients and fundamental frequency features from audio segments, and then clusters them in its feature space to obtain frame-level, phoneme-level, and whole-sentence-level feature subsets. For the video modality, spatiotemporal features are extracted from keyframes and optical flow information, and then clustered into frame-level, shot-level, and segment-level granularities. After each modality completes its internal granularity partitioning, the system uses the original timestamp as an index to align features of the same granularity for text, images, audio, and video. If sentence-level text, region-level image, phoneme-level audio, and shot-level video features exist simultaneously at a given moment, the quadruple is marked as a corresponding sample at the same granularity level. If a modality lacks a corresponding granularity at that moment, a zero vector is used as a placeholder to ensure consistent dimensions in subsequent matrix operations. Through this alignment mechanism, the system ultimately forms a "multimodal, multi-granularity knowledge representation" covering four modalities—text, images, audio, and video—with multi-level granularity within each modality. This representation directly inherits the time alignment relationship of the original data without introducing additional parameters. It preserves the granularity structure of each modality and achieves a one-to-one mapping across modal granularities, laying a data foundation for subsequent dependency calculations and optimal granularity selection, and significantly reducing the risk of error propagation due to granularity mismatch.
[0039] Specifically, such as Figure 2 As shown, text features are extracted from text modal data, including part-of-speech features, and / or semantic role features, and / or phonological features, etc.; image features are extracted from image modal data, including color features, and / or posture features, and / or atmosphere features, etc.; audio features are extracted from audio modal data, including intonation features, and / or pitch features, and / or tone features, etc.; and video features are extracted from video modal data, including plot features, and / or motion features, and / or narrative features, etc.
[0040] Single-granularity knowledge representation refers to partitioning data using a specific subset of features. To comprehensively capture sample information, multi-granularity knowledge representation is introduced. Multi-granularity knowledge representation involves clustering the extracted feature set to obtain multiple feature subsets, and then partitioning the dataset differently based on these different feature subsets. For example... Figure 2As shown, the process of multimodal multi-granularity modeling and optimal granularity selection is illustrated.
[0041] S3. Calculate the dependency of each mode based on the multimodal multigranularity knowledge representation, select the optimal multiple granularity levels for each mode, and obtain the multimodal optimal granularity matrix.
[0042] Specifically, step S3 further includes: the granularity level list constructed from the modality M in the multimodal multigranularity knowledge representation is as follows: Where I is the number of granularities in the granularity level list of mode M. The granularity is for the i-th mode M;
[0043] Based on the decision set D and the granularity of the i-th mode M The dependency of mode M is calculated using the following formula: ;
[0044] The n granularities with the highest dependence in mode M are selected, and the optimal granularity matrix of the multimodal model is obtained based on the optimal n granularity levels of each mode, where n is the preset number of granularities.
[0045] In this embodiment, based on multimodal and multigranular knowledge representation, for each modality M (text, image, audio, video), a granularity level list has been formed internally according to the clustering results. Using the decision set (ironic / non-ironic tag set) as a benchmark, the contribution of each granularity to the classification decision is directly calculated using the rough set dependency formula; the positive domain in this formula consists of a set of samples where the granularity is unambiguous on the decision set. After calculation, all granularities of the current modality are sorted in descending order of dependency, and the top n granularities with the highest dependency are automatically selected as the optimal granularity combination for that modality; the value of n is preset by the user. After the above process is performed on the text, image, audio, and video modalities respectively, the system concatenates their respective optimal granularity subsets column-wise to form a multimodal optimal granularity matrix with consistent dimensions. This matrix retains the most discriminative granularity information for each modality and filters out low-contribution or noisy granularities through dependency quantification, demonstrating the dual advantages of adaptive granularity selection in terms of accuracy and efficiency.
[0046] Specifically, such as Figure 2 As shown, the dependency function is a mathematical tool in rough set theory used to measure the classification ability of conditional features (or granularity) on decision features. It quantifies the degree to which "equivalence classes partitioned by conditional features or granularity levels can accurately predict the decision class." Dependency is the specific value of the dependency function, representing the degree of dependence of the conditional attribute on the decision attribute, and its value ranges from [0,1]. Figure 2 The diagram illustrates the process of multimodal, multi-granularity modeling and optimal granularity selection. The granularity level lists for text, image, audio, and video modal data are as follows: , , and .
[0047] Based on the decision set D and the granularity level Calculate text dependency Select the n granularities with the highest text dependency. Based on the decision set D and the granularity level... Calculate image dependency Select the n granularities with the highest image dependency. Based on the decision set D and the granularity level... Calculate audio dependency Select the n granularities with the highest audio dependency. Based on the decision set D and the granularity level... Calculate video dependency Select the n granularities with the highest video dependency.
[0048] The formula for calculating text dependency is: The formula for calculating image dependency is: The formula for calculating audio dependency is: The formula for calculating video dependency is: . , , and These are the decision attribute D at different granularities. , , and The positive domain (i.e., the set of samples that determines a certain decision class). The cardinality of a set.
[0049] Furthermore, based on decision D, all samples in the sample set U are partitioned, denoted as partition U / D; based on , , and Each sample in the sample set U is divided into partitions, denoted as partition U / T. i U / P i U / A i U / V i Decision attribute D at granularity The formula for calculating the positive domain is: ; Decision attribute D at granularity The formula for calculating the positive domain is: ; Decision attribute D at granularity The formula for calculating the positive domain is: ; Decision attribute D at granularity The formula for calculating the positive domain is: The multimodal optimal granularity matrix includes the optimal granularity matrix for text, image, audio, and video.
[0050] S4. Extract the prototype mode vector and test mode vector from the multimodal optimal granularity matrix, and construct the irony order parameter based on the prototype mode vector and test mode vector.
[0051] Specifically, step S4 further includes: extracting the row vectors of the training set samples with known irony labels from the multimodal optimal granularity matrix as the prototype pattern vector of modality M. , where v is the prototype and m is the number of row vectors in the prototype pattern vector;
[0052] The row vectors of the test set samples are extracted from the multimodal optimal granularity matrix and used as the test mode vectors of mode M. Where u is the test and n is the number of row vectors in the test pattern vector;
[0053] Based on prototype pattern vector and test mode vector Construct an ironic order parameter that reflects the similarity between the prototype pattern vector and the test pattern vector.
[0054] The formula for calculating the irony order parameter is as follows: Where d is the number of samples in the test set. As an order parameter, Let T be the ironic order parameter of modality M, and T be the text.
[0055] In this embodiment, for each modality, firstly, all samples explicitly labeled as "irony" in the training set are traversed, and their row vectors in the corresponding modality's optimal granularity matrix are extracted sequentially, denoted as the prototype modality vector set. Subsequently, for each sample to be identified in the test set, its row vectors in the same modality's optimal granularity matrix are extracted, denoted as the test modality vector set. Since the matrix has already undergone dependency filtering, the dimensions of the prototype and test vectors remain highly consistent, avoiding the curse of dimensionality caused by redundant features. After obtaining the above two sets of vectors, the irony order parameter is directly constructed using the given similarity function: for the a-th sample in the test set, its cosine similarity with all prototype modality vectors is normalized to form the modality's irony order parameter.
[0056] Specifically, such as Figure 3 As shown, this method extracts the row vectors corresponding to the known ironic labels of the training set samples from the optimal granularity matrix of each modality as the prototype pattern vector for that modality. The prototype pattern vector represents the typical pattern of the known ironic labels. The prototype pattern vector represents the typical example feature set of a certain category and is a centralized representation of that category. Figure 3The diagram illustrates the process of co-evolution. The prototype pattern vector includes text prototype pattern vectors. Image prototype pattern vector Audio prototype mode vector and video prototype mode vector The test pattern vector includes text test pattern vectors. Image test mode vector Audio test mode vector and video test mode vector .
[0057] Furthermore, the computational model for the text irony order parameter is as follows: The computational model for the image irony order parameter is: The computational model for the audio irony order parameter is: The computational model for the video irony order parameter is: Where T represents text, P represents image, A represents audio, and V represents video.
[0058] S5 employs a collaborative neural network model for ironic order parameters to perform collaborative recognition preprocessing on the order parameters, generating ironic recognition results.
[0059] Specifically, step S5 further includes: constructing a collaborative neural network model based on the principle of synergetics, and reconstructing the order parameter using a collaborative neural network model with ironic order parameters;
[0060] The reconstructed order parameters are fused and sorted to select the prototype pattern vector corresponding to the highest order parameter in the sequence.
[0061] When it is determined that the sample to be identified belongs to the prototype pattern vector corresponding to the highest order parameter in the sequence, the irony labeling result of the sample to be identified is obtained, and the irony recognition result is generated.
[0062] The formula for the cooperative neural network model under mode M is: ,in, Note the parameters used to control the rate of change of the test mode; B is the lateral inhibition coefficient, and C is the self-inhibition coefficient. This is a pattern index.
[0063] Furthermore, in collaborative neural network models, parameter combinations Together, they determine the recognition performance of collaborative pattern recognition; This is a self-motivating term, representing the feedback and motivational effect of the pattern on itself; This is a self-inhibiting term, reflecting the pattern's inhibition of its own excessive growth; This is a lateral inhibition term, reflecting the mutual inhibition between patterns, and is essentially a penalty term.
[0064] In this embodiment, after calculating the irony order parameters, these parameters are immediately fed into a collaborative neural network model built based on the principle of synergetics. Based on the initial state, the model evolves within discrete time steps. During network iteration, each modality runs in parallel, requiring no additional training; the order parameters can be reconstructed solely through dynamic equations. After a few time steps, the internal competition stabilizes, and the reconstructed order parameters fully reflect cross-modal consensus. Subsequently, the four-modal reconstruction results are weighted and fused, with weights determined by the normalization of the obtained modal dependencies. The fused vectors are arranged in descending order of value, and the prototype mode vector corresponding to the highest value is considered the final decision criterion. The sample to be identified is determined to belong to the irony class only if and only if the highest order parameter exceeds a preset threshold; otherwise, it is assigned to the non-irony class, and the corresponding irony labeling result is output.
[0065] Specifically, the collaborative neural network model in the text modality is: The collaborative neural network model in the image modality is: The collaborative neural network model for the audio modality is: The collaborative neural network model in the video modality is: .
[0066] In the collaborative neural network model, , , and This is a self-motivating term, representing the feedback and motivational effect of the pattern on itself; , , and This is a self-inhibiting term, reflecting the pattern's inhibition of its own excessive growth; , , and This is a lateral inhibition term, reflecting the mutual inhibition between patterns, essentially a penalty term. In the formula, B=C=1.2, while... The values were 0.18, 0.36, 0.54, and 0.72, respectively.
[0067] Based on the order parameters of each modality, a weighted multimodal approach is introduced, and the weight coefficients of each modality are calculated using dependency. (These correspond to the weighting coefficients for text, image, audio, and video modalities, respectively). The formula for calculating the text modal weighting coefficient is: The formula for calculating the image modality weight coefficient is: The formula for calculating the weighting coefficients of audio modalities is: The formula for calculating the video modal weight coefficient is: .
[0068] Based on the weight coefficients of each modality and the fusion order parameter, calculate the multimodal fusion order parameter. The formula for calculating the multimodal fusion order parameter is: .in, The text irony order parameter indicates the text's irony. Indicates the image irony order parameter, Indicates the audio irony order parameter, Indicates the video irony order parameter. Based on the multimodal fusion order parameter... The process involves obtaining the irony annotation pattern of the sample to be identified, and then obtaining the irony recognition result and the final irony label.
[0069] Of course, the preset number of optimal granularities for the four modalities can also be different, depending on actual needs. If the preset number of optimal granularities for each modality is different, then when performing cross-modal fusion, the optimal granularities of each modality need to be vectorized and mapped to the same dimension. Specifically, this includes: first, vectorizing the features in the optimal granularities of each modality to obtain the feature vectors of each modality; second, to ensure subsequent operations, fully connected layers or other methods can be used to map the feature vectors of all modalities to the same dimension; then, multiplying each mapped feature vector by its corresponding weight coefficient; finally, concatenating all weighted modal feature vectors.
[0070] In summary, this method integrates multimodal data, combines hierarchical granular analysis, and utilizes co-evolution to achieve comprehensive identification of satirical samples. It leverages the complementarity and synergy of multimodal data, comprehensively analyzes multimodal data at different granularities, and provides a more accurate method for satirical identification. In short, the granular computing-based multimodal co-evolutionary satirical identification method extracts multimodal data and performs multi-granular hierarchical analysis to achieve accurate modeling and identification of satirical expressions, solving the technical challenges of insufficient accuracy and inadequate generalization ability in existing satirical identification techniques.
[0071] Please see Figure 4 The second embodiment of the present invention provides a multimodal cooperative evolution irony recognition device based on granular computing, which includes:
[0072] The data processing unit 101 is used to extract sample data of different modalities from a preset sample source library and construct a multimodal irony dataset based on the sample data;
[0073] Multi-granularity unit 102 is used to perform hierarchical granularity analysis and partitioning of the multimodal irony dataset to construct a multimodal multi-granularity knowledge representation;
[0074] The optimal granularity unit 103 is used to calculate the dependency of each mode based on the multimodal multigranularity knowledge representation, select the optimal multiple granularity levels for each mode, and obtain the multimodal optimal granularity matrix.
[0075] The order parameter unit 104 is used to extract the prototype mode vector and the test mode vector from the multimodal optimal granularity matrix, and construct the ironic order parameter based on the prototype mode vector and the test mode vector.
[0076] The collaborative recognition unit 105 is used to perform collaborative recognition preprocessing on the order parameter using a collaborative neural network model of the irony order parameter, and generate irony recognition results.
[0077] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A multimodal cooperative evolutionary irony recognition method based on granular computing, characterized in that, include: Extract sample data of different modalities from a pre-defined sample source library, and construct a multimodal satirical dataset based on the sample data; A hierarchical granularity analysis and partitioning process is performed on the multimodal irony dataset to construct a multimodal, multigranular knowledge representation; The dependency of each mode is calculated based on the multimodal multigranularity knowledge representation, and the optimal granularity levels of each mode are selected to obtain the multimodal optimal granularity matrix. Extract the prototype mode vector and test mode vector from the multimodal optimal granularity matrix, and construct the irony order parameter based on the prototype mode vector and test mode vector; A collaborative neural network model for ironic order parameters is used to perform collaborative recognition preprocessing on the order parameters to generate ironic recognition results. The formula for calculating the irony order parameter is as follows: Where d is the number of samples in the test set. As an order parameter, Let T be the ironic order parameter of modality M, and T be the text. The prototype pattern vector of mode M, where v is the prototype and m is the number of row vectors in the prototype pattern vector. Let u be the test pattern vector of modality M, where u is the test and n is the number of row vectors in the test pattern vector. A collaborative neural network model for ironic order parameters is used to perform collaborative recognition preprocessing on the order parameters to generate ironic recognition results, specifically: A collaborative neural network model is constructed based on the principle of synergetics, and the order parameter is reconstructed using a collaborative neural network model with ironic order parameters. The reconstructed order parameters are fused and sorted to select the prototype pattern vector corresponding to the highest order parameter in the sequence. When it is determined that the sample to be identified belongs to the prototype pattern vector corresponding to the highest order parameter in the sequence, the irony labeling result of the sample to be identified is obtained, and the irony recognition result is generated. The formula for the cooperative neural network model under mode M is: ,in, Note the parameters: B is the lateral inhibition coefficient, and C is the self-inhibition coefficient. This is a pattern index.
2. The multimodal cooperative evolution irony recognition method based on granular computing according to claim 1, characterized in that, Sample data of different modalities are extracted from a pre-defined sample source library, and a multimodal irony dataset is constructed based on the sample data, specifically as follows: Multimodal sample data is obtained by extracting sample data of different modalities from a pre-defined sample source library. The modalities include text, images, audio, and video. Multimodal sample data is preprocessed, and the preprocessed data is then annotated with irony labels. The annotation results are divided into two categories: irony and non-irony. The preprocessing includes word segmentation, stop word removal, and standardization for text data; size normalization, illumination adjustment, face detection and alignment for image data; noise reduction, volume standardization, and frame segmentation for audio data; and frame rate unification, resolution adjustment, and keyframe extraction for video data. The labeled multimodal sample data are aligned according to the timestamps to construct a multimodal satirical dataset.
3. The multimodal cooperative evolution irony recognition method based on granular computing according to claim 1, characterized in that, A hierarchical granularity analysis and partitioning process is performed on the multimodal irony dataset to construct a multimodal, multi-granularity knowledge representation, specifically: Features are extracted from the multimodal irony dataset to obtain the feature space of each modality, and different feature subsets are obtained from each modality feature space through feature clustering technology to construct knowledge representations of different granularities; The multimodal satirical dataset is partitioned based on different feature subsets to obtain multi-granularity feature subspaces for each modality. The multi-granularity feature subspaces for each modality are then aligned to construct a multimodal multi-granularity knowledge representation.
4. The multimodal cooperative evolution irony recognition method based on granular computing according to claim 1, characterized in that, The dependencies of each mode are calculated based on multimodal multigranularity knowledge representation. The optimal granularity levels for each mode are then selected, resulting in the multimodal optimal granularity matrix, as follows: The list of granularity levels constructed from the modality M in the multimodal multigranularity knowledge representation is as follows: Where I is the number of granularities in the granularity level list of mode M. The granularity is for the i-th mode M; Based on the decision set D and the granularity of the i-th mode M The dependency of mode M is calculated using the following formula: ; The n granularities with the highest dependence in mode M are selected, and the optimal granularity matrix of the multimodal model is obtained based on the optimal n granularity levels of each mode, where n is the preset number of granularities.
5. The multimodal cooperative evolution irony recognition method based on granular computing according to claim 1, characterized in that, Extract the prototype mode vector and test mode vector from the multimodal optimal granularity matrix, and construct the irony order parameter based on the prototype mode vector and test mode vector, specifically: Extract the row vectors of the training set samples with known ironic labels from the multimodal optimal granularity matrix as the prototype pattern vectors of modality M. ; The row vectors of the test set samples are extracted from the multimodal optimal granularity matrix and used as the test mode vectors of mode M. ; Based on prototype pattern vector and test mode vector Construct an ironic order parameter that reflects the similarity between the prototype pattern vector and the test pattern vector.
6. A multimodal cooperative evolution irony recognition device based on granular computing, characterized in that, include: The data processing unit is used to extract sample data of different modalities from a preset sample source library and construct a multimodal irony dataset based on the sample data; Multi-granularity units are used to perform hierarchical granularity analysis and partitioning of multimodal irony datasets to construct multimodal multi-granularity knowledge representations; The optimal granularity unit is used to calculate the dependency of each mode based on the multimodal multigranularity knowledge representation, select the optimal multiple granularity levels for each mode, and obtain the multimodal optimal granularity matrix. The order parameter unit is used to extract the prototype mode vector and test mode vector from the multimodal optimal granularity matrix, and construct the ironic order parameter based on the prototype mode vector and test mode vector; The collaborative recognition unit is used to perform collaborative recognition preprocessing on the order parameter using a collaborative neural network model of the irony order parameter, and generate irony recognition results. The formula for calculating the irony order parameter is as follows: Where d is the number of samples in the test set. As an order parameter, Let T be the ironic order parameter of modality M, and T be the text. The prototype pattern vector of mode M, where v is the prototype and m is the number of row vectors in the prototype pattern vector. Let u be the test pattern vector of modality M, where u is the test and n is the number of row vectors in the test pattern vector. A collaborative neural network model for ironic order parameters is used to perform collaborative recognition preprocessing on the order parameters to generate ironic recognition results, specifically: A collaborative neural network model is constructed based on the principle of synergetics, and the order parameter is reconstructed using a collaborative neural network model with ironic order parameters. The reconstructed order parameters are fused and sorted to select the prototype pattern vector corresponding to the highest order parameter in the sequence. When it is determined that the sample to be identified belongs to the prototype pattern vector corresponding to the highest order parameter in the sequence, the irony labeling result of the sample to be identified is obtained, and the irony recognition result is generated. The formula for the cooperative neural network model under mode M is: ,in, Note the parameters: B is the lateral inhibition coefficient, and C is the self-inhibition coefficient. This is a pattern index.
Citation Information
Patent Citations
Multi-modal Mongolian sentiment analysis method based on irony recognition and fine-grained feature fusion
CN113657115A
System to detect, assess and counter disinformation
US20220164643A1