Cross-modal multi-granularity humor recognition method and device based on composite discrimination measure
By preprocessing multimodal data and sampling at multiple granularities, and combining discriminative measures and variable coherence metrics for cross-modal fusion, the problems of insufficient information utilization and inaccurate feature combination in existing humor recognition technologies are solved, thereby improving the accuracy and robustness of humor recognition.
Patent Information
- Application Number
- CN202511156425.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2045-08-19
AI Technical Summary
Existing humor recognition technologies suffer from problems such as insufficient utilization of single-modal information, inadequate mining of multi-granularity knowledge, inability to accurately evaluate and select the optimal feature combination, and oversimplification of cross-modal fusion strategies, resulting in poor recognition performance.
A cross-modal multi-granularity humor recognition method based on composite discriminative measure is adopted. Through multimodal data preprocessing, multi-granularity sampling, optimal granularity combination and discriminative measure fusion, composite discriminative measure value is calculated to recognize humorous content.
It improves the accuracy and robustness of humor recognition, comprehensively portrays humorous content, overcomes the problem of insufficient utilization of single-modal information, and enhances the feature fusion effect.
Smart Images

Figure CN120822078B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of data mining and artificial intelligence technology, and more specifically, to a method and apparatus for cross-modal multi-granular humor recognition based on composite discriminative measures. Background Technology
[0002] With the rapid development of artificial intelligence technology, humor recognition has received widespread attention in fields such as human-computer interaction and content recommendation. Humor recognition aims to automatically determine whether content possesses humorous attributes by analyzing multimodal data such as text, images, audio, and video. However, due to the complex contextual understanding, cultural cognition, and deep logical relationships involved in humorous expression, its automated recognition faces numerous challenges. Humorous content typically relies on specific situations and contextual information, and exhibits diversity and innovation across different cultural backgrounds, further increasing the difficulty of technological implementation. Although advances in deep learning technology have enabled computers to demonstrate near-human comprehension capabilities in certain specific tasks, existing methods still have significant limitations when processing humorous content with high cognitive complexity.
[0003] The main problems with current humor recognition technology are as follows: First, existing research generally employs single-modal analysis methods, such as focusing only on text or audio features, making it difficult to comprehensively capture the rich information in multimodal data. This limitation prevents the system from fully utilizing non-textual humor elements such as visual and auditory elements, thus affecting recognition performance. Second, in the knowledge representation and feature extraction processes, existing methods lack effective coordination criteria, making it impossible to accurately select the optimal feature combination. Third, cross-modal feature fusion strategies are relatively simple, failing to establish effective measures to evaluate the importance and complementarity of different modal features, leading to information loss during fusion. Furthermore, when processing multi-granular information, the lack of quantitative indicators for system complexity and category discrimination makes it difficult to achieve optimal representation of data features. These problems severely limit the accuracy and robustness of humor recognition technology in practical applications.
[0004] In view of the above, this application is hereby submitted. Summary of the Invention
[0005] This invention aims to provide a cross-modal multi-granularity humor recognition method and device based on composite discriminative measures, in order to solve the problems of insufficient utilization of single-modal information, insufficient mining of multi-granularity knowledge, inability to accurately evaluate and select the optimal feature combination, and simplification of cross-modal fusion strategies in existing humor recognition technologies, so as to achieve a comprehensive characterization of humorous content and significantly improve the accuracy and robustness of humor recognition.
[0006] To solve the above-mentioned technical problems, the present invention is achieved through the following technical solution:
[0007] A cross-modal, multi-granular humor recognition method based on a composite discriminative measure includes:
[0008] S1. Acquire multimodal data and preprocess it to establish a multimodal dataset;
[0009] S2, Based on the multimodal dataset, multi-granularity sampling is performed in each modality to obtain a cross-modal multi-granularity sample space;
[0010] S3, based on the division of each modality data in the multimodal dataset relative to the preset decision set and granularity, select the optimal granularity from the multi-granularity sample space of each modality to satisfy the preset number of coordination requirements, and form the optimal granularity combination.
[0011] S4. Based on the optimal granularity combination, cross-modal fusion is performed using a discriminative measure and a variable coordination metric to calculate a composite discriminative measure value for identifying humorous content and obtaining humor recognition results.
[0012] Preferably, the multimodal data includes text data, and / or image data, and / or audio data, and / or video data.
[0013] Preferably, the preprocessing includes: segmenting text data into words, and / or removing stop words, and / or standardizing the text data; normalizing image data in size, and / or converting color space; denoising audio data, and / or segmenting the audio data into frames; and extracting frames and / or extracting motion from video data.
[0014] Preferably, the method also includes humorous annotation and alignment of the preprocessed multimodal data.
[0015] Preferably, the process of obtaining the cross-modal multi-granularity sample space specifically includes:
[0016] Features of each modality are extracted based on the multimodal dataset;
[0017] Multiple feature subsets are obtained based on the extracted features;
[0018] Different feature subsets constitute different granularities. The dataset is divided based on different feature subsets to obtain multi-granularity sample spaces for each modality.
[0019] By combining the multi-granularity sample spaces corresponding to each modality data, a cross-modal multi-granularity sample space is obtained.
[0020] Preferably, the selection process for the optimal granularity combination specifically includes:
[0021] Get the preset decision set ;
[0022] according to For the multimodal dataset All samples are divided into groups, denoted as partition . ;
[0023] Based on the cross-modal multi-granularity sample space, the multimodal dataset All samples are divided into groups, denoted as partition . , Indicates the granularity level; ; This indicates the total number of granularities; a larger number indicates a finer granularity. Indicates the first indivual Modal granularity;
[0024] Based on the granularity level from coarse to fine, identify the multi-granularity sample spaces corresponding to each mode that satisfy... of By calculating the granularity of each mode, the optimal granularity is obtained, thus forming the optimal granularity combination.
[0025] Among them, when the following conditions are met When, it indicates a mode. In particle size Above and in the decision set Coordination.
[0026] Preferably, the process of calculating the composite discriminant measure value through cross-modal fusion using a discriminant measure and a variable reconciliation metric is as follows:
[0027] Using a variable coordination metric to describe the relationship between granularity and decision-making, then modality In particle size Variable coordination measure The calculation formula is:
[0028] ;
[0029] The weighting coefficients of each mode are calculated using a variable compatibility metric. The formula for calculating the weighting coefficient is:
[0030] ;
[0031] in, Representing various modes; Indicates one of the modes; Indicates the preset number of units for the optimal granularity; Representing modes In particle size Variable coordination measure;
[0032] According to the division and division Calculate each mode at different granularities Conditional entropy ;
[0033] Then, modality In particle size Conditional entropy The calculation formula is:
[0034] ;
[0035] ;
[0036] ;
[0037] in, It is an equivalence class The probability of; It is an equivalence class The middle belongs to the decision-making category The conditional probability; The cardinality of a set.
[0038] Based on the conditional entropy of each mode Calculate each mode at different granularities Discrimination measure ;
[0039] Then, modality In particle size Discrimination measure The calculation formula is:
[0040] ;
[0041] ;
[0042] in, It is an equivalence class The probability of; It is modal In particle size Conditional entropy; The cardinality of a set.
[0043] Based on the weighting coefficients of each modality, the composite discrimination measure is calculated through cross-modal fusion. Then the composite distinguishing measure value The calculation formula is:
[0044] ;
[0045] in, Representing modes In particle size The distinguishing measure below.
[0046] Preferably, the process of identifying humorous content specifically involves:
[0047] Set humor recognition threshold If the composite distinguishing measure value If it is, it is considered to contain humorous content; otherwise, it is considered to contain non-humorous content.
[0048] The present invention also provides a cross-modal multi-granularity humor recognition device based on a composite discriminative metric, comprising:
[0049] The acquisition unit is used to acquire multimodal data and perform preprocessing to establish a multimodal dataset;
[0050] A multi-granularity sampling unit is used to perform multi-granularity sampling in each modality based on the multimodal dataset to obtain a cross-modal multi-granularity sample space;
[0051] The optimal granularity unit is used to select a preset number of optimal granularities from the multi-granularity sample space of each modality based on the division of each modality data in the multimodal dataset relative to the preset decision set and granularity, thereby forming an optimal granularity combination.
[0052] The recognition unit is used to perform cross-modal fusion based on the optimal granularity combination, using a discriminative measure and a variable coordination metric, to calculate a composite discriminative measure value, thereby recognizing humorous content and obtaining a humor recognition result.
[0053] The present invention also provides a cross-modal multi-granularity humor recognition device based on a composite discriminative metric, comprising a processor and a memory, wherein the memory stores a computer program that can be executed by the processor to implement the cross-modal multi-granularity humor recognition method based on a composite discriminative metric as described above.
[0054] The present invention also provides a computer-readable storage medium storing computer-readable instructions, which, when executed by a processor of the device in which the computer-readable storage medium resides, implement the cross-modal multi-granularity humor recognition method based on a composite discriminative metric as described above.
[0055] In summary, compared with the prior art, the present invention has the following beneficial effects:
[0056] First, this invention uses multimodal data collaborative representation to integrate multimodal information such as text, images, audio, and video, fully mining the multidimensional features of humorous content and overcoming the problem of insufficient utilization of single-modal information.
[0057] Second, this invention uses a multi-granularity knowledge mining strategy to extract features from different granularity levels, avoiding the information loss problem that may be caused by fixed-granularity feature extraction, and improving the comprehensiveness and flexibility of feature representation.
[0058] Third, this invention selects the optimal granularity through coordination, combines the discriminative measure and the variable coordination measure to perform the optimal granularity combination for cross-modal fusion, accurately selects the optimal feature combination, effectively integrates the complementarity and synergy of data from various modalities, and enhances the effect of feature fusion.
[0059] Fourth, this invention improves the accuracy and robustness of humor recognition by calculating composite discrimination measure values for humor identification and classification.
[0060] In summary, the method of this invention not only overcomes the limitations of existing technologies in cross-modal multi-granularity sampling and fusion, but also significantly improves the performance of humor recognition, providing new research ideas and technical paths for the field of humor recognition, and has important theoretical value and practical application prospects. Attached Figure Description
[0061] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained from these drawings without creative effort.
[0062] Figure 1 This is a schematic diagram of a cross-modal multi-granularity humor recognition method based on composite discriminative measures, provided in Example 1.
[0063] Figure 2 This is a schematic diagram of the multi-granularity sampling process provided in Example 1.
[0064] Figure 3 This is a schematic diagram of cross-modal multi-granularity sampling and optimal granularity selection provided in Example 1.
[0065] Figure 4 This is a schematic diagram of cross-modal fusion calculation of composite discriminative measure value provided in Example 1.
[0066] Figure 5 This is a schematic diagram of a cross-modal multi-granularity humor recognition device based on composite discrimination measure provided in Embodiment 2.
[0067] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Detailed Implementation
[0068] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0069] Example 1
[0070] Embodiment 1 of the present invention provides a cross-modal multi-granularity humor recognition method based on composite discriminative measure, which can be implemented by a cross-modal multi-granularity humor recognition device based on composite discriminative measure (hereinafter referred to as humor recognition device), specifically, executed by one or more processors within the humor recognition device.
[0071] In this embodiment, the humor recognition device may be an electronic device equipped with a processor, which carries a computer program for the cross-modal multi-granularity humor recognition method based on composite discrimination measure and the computer program can be executed, such as a computer, smartphone, smart tablet, workstation, etc., which are not limited here.
[0072] like Figure 1 As shown, a cross-modal multi-granularity humor recognition method based on composite discriminative measure includes steps S1 to S4.
[0073] S1. Acquire multimodal data and preprocess it to establish a multimodal dataset.
[0074] Specifically, step S1 includes steps S11 to S14.
[0075] S11, Collect TED Talks as multimodal data, which includes text data and / or image data, and / or audio data, and / or video data, etc. Extract multiple modalities of data, such as text, images, audio, and video, from the multimodal data.
[0076] S12 performs preprocessing on text data, including word segmentation, stop word removal, and standardization; preprocessing on image data, including size normalization and color space conversion; preprocessing on audio data, including noise reduction and frame segmentation; and preprocessing on video data, including frame extraction and motion extraction.
[0077] Specifically, for text data, firstly, natural language processing tools such as NLTK are used for word segmentation to divide continuous text content into word sequences; then, stop words such as "the", "a", and "an" that do not contribute much to semantics are removed; next, lemmatization is performed to unify words of different forms into their original forms; finally, standardization is performed to unify capitalization and remove special characters, etc.
[0078] For image data, all images are first scaled to a standard size such as a preset size (e.g., 299×299); then the RGB images are converted to color spaces such as HSV or Lab to obtain better feature representation; next, image enhancement processing is performed, including adjusting brightness and contrast; finally, normalization processing is performed to scale the pixel values to the [0,1] range.
[0079] For audio data, noise reduction algorithms are first used to remove background noise and improve signal quality; then the audio signal is processed in frames, typically with a frame length of 25ms and a frame shift of 10ms; next, acoustic features such as MFCC are extracted for subsequent analysis; finally, volume normalization is performed to ensure that the loudness of different audio segments remains consistent.
[0080] For video data, keyframes are first extracted at fixed intervals (e.g., 1 frame per second); then, tools such as OpenPose are used to extract human pose features from the video; next, optical flow features are extracted to capture motion information; finally, the frame sequence is temporally sampled to ensure that all video segments have a uniform length for easier subsequent processing.
[0081] S13, label the preprocessed multimodal data, including whether or not there are humor annotations.
[0082] Specifically, the collected raw samples are manually labeled, with the label indicating whether the sample contains humorous expressions (yes -0 / no -1).
[0083] S14. Assign data from different modalities to different time sequences or semantics, align multimodal data, and establish a multimodal dataset.
[0084] Preprocessing can remove noise from multimodal data (such as typos in text, blurred areas in images, and background noise in audio), ensuring data consistency and validity, and providing reliable input for subsequent granular analysis and model training.
[0085] S2, Based on the multimodal dataset, multi-granularity sampling is performed in each modality to obtain a cross-modal multi-granularity sample space.
[0086] Specifically, step S2 includes steps S21 to S26.
[0087] S21, extract text features based on the text data; wherein, text features include parts of speech, and / or grammatical structure, and / or semantic roles, and / or sentiment, and / or phonology, etc.
[0088] S22, based on the image data, extract image features; wherein, image features include color, and / or topological relationships, and / or posture, and / or action, and / or atmosphere, etc.
[0089] S23, based on the audio data, extract audio features; wherein, the audio features include pitch, and / or timbre, and / or stress, and / or intonation, and / or tone, etc.
[0090] S24, based on the video data, extract video features; wherein, video features include changes in motion, and / or scene transitions, and / or activities, and / or plot development, and / or transition effects, etc.
[0091] For different modalities (text, image, audio, video), features are extracted from "coarse-grained" to "fine-grained". For example:
[0092] Text modalities: coarse-grained (paragraph level, topic level), medium-grained (phrase level, sentence level), fine-grained (word level, character level);
[0093] Image modalities: coarse-grained (scene-level, semantic-level), medium-grained (region-level, object-level), and fine-grained (pixel-level, edge features);
[0094] Audio modalities: coarse-grained (sentence level, sentiment level), medium-grained (phoneme level, syllable level), fine-grained (spectral frame level);
[0095] Video modalities: coarse-grained (sequence level, storyline level), medium-grained (segment level, scene level), and fine-grained (frame level, motion detail level).
[0096] S25. Under each modal data, multiple feature subsets are obtained based on the extracted feature set. Different feature subsets constitute different granularities, i.e., multi-granular sampling, thereby obtaining text multi-granular sample space, image multi-granular sample space, audio multi-granular sample space and video multi-granular sample space respectively.
[0097] Preferably, single-granularity sampling refers to partitioning the data using a specific subset of features. To comprehensively capture sample information, multi-granularity sampling is introduced. Multi-granularity sampling involves obtaining multiple feature subsets based on the extracted features, and then partitioning the dataset differently based on these different feature subsets. For example... Figure 2 As shown, the process of multi-granularity sampling is illustrated.
[0098] For text data, different samples can be generated at the word level, phrase level, and sentence level; for image data, features at different levels can be extracted according to pixel blocks, local regions, and global structures; for audio data, layered sampling can be performed at the frame level, segment level, and overall waveform level.
[0099] For example, for the complete feature set consisting of N features extracted from a text modality, the feature subset can be generated by traversing all possible feature combinations, i.e., generating... A non-empty subset of features. Or, more specifically, for example, for the complete set of features extracted from a text modality {part-of-speech, grammatical structure, semantic role}, it can be combined into multiple feature subsets, such as single-feature granularity: {part-of-speech}, {grammatical structure}, {semantic role}; double-feature granularity: {part-of-speech, grammatical structure}, {part-of-speech, semantic role}, {grammatical structure, semantic role}; and triple-feature granularity: {part-of-speech, grammatical structure, semantic role}. Each feature subset is a granularity, which together constitute the multi-granularity sample space under this modality.
[0100] S26 combines the multi-granularity sample spaces corresponding to each modality data to obtain a cross-modal multi-granularity sample space.
[0101] In this step, these features at different levels are combined into a cross-modal multi-granularity sample space, which contains feature sets from different modalities and granularities for subsequent analysis and processing.
[0102] In other embodiments, those skilled in the art can extract other features beyond those described above, and the present invention does not specifically limit them.
[0103] S3. Based on the division of each modality data in the multimodal dataset relative to the preset decision set and granularity, select the optimal granularity from the multi-granularity sample space of each modality to achieve a preset number of granularities that satisfy coordination, and form the optimal granularity combination.
[0104] Specifically, such as Figure 2 and Figure 3 As shown, step S3 includes steps S31 to S33.
[0105] S31, Given a decision set ,according to sample set All samples in the (i.e., multimodal dataset) are divided into two groups, denoted as partitioning. The sample set This is a collection of TED talk samples.
[0106] Specifically, "decision set" "This refers to the set of classification criteria or target labels used to guide multi-granularity knowledge fusion. It contains the humor annotation information for each sample in the corpus, i.e., the humor label (humorous or not humorous). Decision set" N represents the set The total number of samples in the sample.
[0107] S32, Let the granularity of each modality of data be as follows: Text data: Image data: Audio data: Video data: , , This indicates the total number of granularities (features); the larger the number, the finer the granularity.
[0108] according to Sample sets All samples are divided into partitions, denoted as partitions _____. .
[0109] S33, based on the granularity level from coarse to fine, find the multi-granularity sample spaces corresponding to each mode that satisfy the coordination requirement. By calculating the granularity of each mode, the optimal granularity is obtained, thus forming the optimal granularity combination.
[0110] In this embodiment, coordination refers to the harmonious, consistent, or complementary characteristics exhibited by samples or features at different granularity levels under a certain evaluation criterion. In the process of finding the optimal granularity combination, coordination is used as an important criterion for evaluating the effectiveness of samples or features at different granularity levels. By comparing the coordination at different granularity levels, it is possible to determine which granularity levels contribute most significantly to the overall performance improvement.
[0111] In this embodiment, the evaluation criterion for coordination is: based on the granularity level from coarse to fine, identify the multi-granularity sample spaces corresponding to each mode that satisfy the following criteria. of By calculating the granularity, the optimal granularity for each mode can be obtained. That is:
[0112] According to the division and division choose of The coarsest granularity is used to obtain the optimal granularity for the text modality.
[0113] According to the division and division choose of The coarsest granularity is used to obtain the optimal granularity for image modality.
[0114] According to the division and division choose of The coarsest granularity is used to obtain the optimal granularity for the audio modality.
[0115] According to the division and division choose of The coarsest granularity is used to obtain the optimal granularity for the video modality.
[0116] Among them, when the following conditions are met When, it indicates a mode. In particle size Above and in the decision set Coordination.
[0117] The optimal granularity of each mode constitutes the optimal granularity combination.
[0118] S4. Based on the optimal granularity combination, cross-modal fusion is performed using a discriminative measure and a variable coordination metric to calculate a composite discriminative measure value for identifying humorous content and obtaining humor recognition results.
[0119] Specifically, such as Figure 4 As shown, step S4 includes steps S41 to S46.
[0120] S41 uses a variable coordination metric to describe the relationship between granularity and decision-making, providing a flexible description of the consistency of granularity within the dataset.
[0121] In this embodiment, variable coordination metric assesses the ability of a system or individual to dynamically adjust the degree of cooperation and harmony among its components under different conditions, speeds, rhythms, or environments. Its core lies in quantifying the level of coordination during this dynamic adaptation process. Variable coordination refers to the ability of a system or individual to flexibly adjust the degree of cooperation among its internal components when facing changes in conditions, speeds, rhythms, or environments to achieve optimal overall performance. This coordination is not only reflected in a static equilibrium state but also emphasizes adaptability and flexibility in dynamic changes.
[0122] Specifically, text modalities at the granularity Variable coordination measure The calculation formula is:
[0123] .
[0124] Image modalities at granularity Variable coordination measure The calculation formula is:
[0125] .
[0126] Audio modalities at the granularity Variable coordination measure The calculation formula is:
[0127] .
[0128] Video modalities at the granularity Variable coordination measure The calculation formula is:
[0129] .
[0130] in, Representing modes The next granularity; They represent particle size respectively The following divisions.
[0131] S42, calculate the weight coefficients of each mode using a variable compatibility metric. (These correspond to the weighting coefficients for text, image, audio, and video modalities, respectively).
[0132] Specifically, assuming the data has four modalities: text, image, audio, and video, and the preset number of optimal granularities is a uniform value (meaning the preset number of optimal granularities is the same for all four modalities), then the text modality weight coefficient... The calculation formula is:
[0133] .
[0134] Image modal weighting coefficients The calculation formula is:
[0135] .
[0136] Weighting coefficients of audio modalities The calculation formula is:
[0137] .
[0138] Video modal weighting coefficients The calculation formula is:
[0139] .
[0140] in, This indicates the preset number of optimal granularities. Of course, the preset number of optimal granularities for the four modalities can also be different, depending on actual needs.
[0141] If the preset number of optimal granularities for each modality is different, then when performing cross-modal fusion, it is necessary to first vectorize and map the optimal granularities of each modality to the same dimension. Specifically, this includes:
[0142] First, the features in the optimal granularity of each mode are vectorized to obtain the feature vector of each mode;
[0143] Secondly, to ensure subsequent operations, fully connected layers or other methods can be used to map the feature vectors of all modalities to the same dimension;
[0144] Then, each mapped feature vector is multiplied by its corresponding weight coefficient;
[0145] Finally, all weighted modal feature vectors are concatenated.
[0146] This process ensures that the complementarity and synergy between different modalities are fully utilized, while preserving important contextual information and intermodal correlation features.
[0147] S43, according to the division and division Calculate each mode at different granularities Conditional entropy ;
[0148] Specifically, text conditional entropy The calculation formula is:
[0149] ;
[0150] ;
[0151] .
[0152] Image conditional entropy The calculation formula is:
[0153] .
[0154] Audio conditional entropy The calculation formula is:
[0155] .
[0156] Video conditional entropy The calculation formula is:
[0157] .
[0158] in, It is an equivalence class The probability of; It is an equivalence class The middle belongs to the decision-making category The conditional probability; The cardinality of a set.
[0159] S44, based on the conditional entropy of each mode. Calculate each mode at different granularities Discrimination measure .
[0160] Specifically, discriminative measure is a concept in information theory that reflects the uncertainty and discriminative power of data. Therefore, text discriminative measure... The calculation formula is:
[0161] ;
[0162] .
[0163] Image discrimination measure The calculation formula is:
[0164] .
[0165] Audio Discrimination Measure The calculation formula is:
[0166] .
[0167] Video Discrimination Measure The calculation formula is:
[0168] .
[0169] in, It is an equivalence class The probability of; The cardinality of a set.
[0170] S45, Based on the weighting coefficients of each modality, calculate the composite discrimination measure value through cross-modal fusion. .
[0171] Composite Discrimination Measure The calculation formula is:
[0172] ;
[0173] in, This indicates the preset number of units for the optimal granularity.
[0174] S46 identifies humorous content and obtains humor recognition results.
[0175] Set humor recognition threshold If the composite distinguishing measure value If it is, it is considered to contain humorous content; otherwise, it is considered to contain non-humorous content.
[0176] This invention establishes a comprehensive dataset containing multimodal information, including text, images, audio, and video, through multimodal data acquisition and preprocessing. Text features such as part-of-speech, grammar, sentiment, and phonology are extracted from different modal data; image features such as color, posture, and action are extracted; audio features such as pitch and intonation are extracted; and video features such as scene and plot are extracted, constructing a cross-modal multi-granularity sample space. An optimal granularity selection method based on coordination is employed to select the optimal granularity for each modality. Cross-modal fusion is performed using a discriminative measure and a variable coordination metric, and a composite discriminative measure value is calculated to identify and classify humorous content.
[0177] Specific applications of this invention include, but are not limited to, automatic detection of humorous content on social media platforms, recommendation of humorous teaching resources in online education, and generation of humorous dialogues in intelligent customer service systems. On social media platforms, this invention can quickly identify and categorize humorous content by analyzing user-posted text, images, and voice messages, thereby improving user experience. In online education, this invention can assist teachers in selecting suitable humorous materials for classroom use, enhancing the fun of teaching. In intelligent customer service systems, this invention can help robots understand users' humorous intentions and generate more natural and human-like responses.
[0178] In summary, compared with the prior art, the present invention has the following beneficial effects:
[0179] First, by using multimodal data collaborative representation, it integrates multimodal information such as text, images, audio, and video, fully mining the multidimensional features of humorous content and overcoming the problem of insufficient utilization of single-modal information. For example, in the TED talk sample, the text modality provides humorous clues at the linguistic level, the image modality supplements facial expressions and scene information, the audio modality captures changes in tone and intonation, and the video modality records actions and plot development. Together, these pieces of information constitute a comprehensive description of the humorous content.
[0180] Second, by employing a multi-granularity knowledge mining strategy, features are extracted from different granular levels, avoiding the information loss that may result from fixed-granularity feature extraction and improving the comprehensiveness and flexibility of feature representation. For example, in text modalities, coarse-grained features focus on the overall semantic structure, while fine-grained features focus on local word collocations. The combination of the two can more accurately depict the semantic connotation of humorous expressions.
[0181] Third, this invention selects the optimal granularity through coordination, combining discriminative measures and variable coordination metrics to fuse the optimal granularity combination across modalities. This accurately selects the best feature combination, effectively integrating the complementarity and synergy of data from different modalities and enhancing the feature fusion effect. For example, in the optimal granularity selection process, coordination is used to select the feature subset that best represents humorous content, thereby improving the accuracy of feature representation. In the cross-modal fusion process, the importance of each modality is determined by weighting using variable coordination metrics, avoiding the problem of uneven weight distribution among modalities.
[0182] Fourth, this invention improves the accuracy and robustness of humor recognition by calculating composite discriminative measures for humor identification and classification. For example, by setting thresholds and performing composite fusion, it effectively captures various modal features in humorous content, thereby improving classification performance.
[0183] In summary, the method of this invention not only overcomes the limitations of existing technologies in cross-modal multi-granularity sampling and fusion, but also significantly improves the performance of humor recognition, providing new research ideas and technical paths for the field of humor recognition, and has important theoretical value and practical application prospects.
[0184] Example 2
[0185] like Figure 5 As shown, the second embodiment of the present invention also provides a cross-modal multi-granularity humor recognition device based on a composite discriminative metric, comprising:
[0186] The acquisition unit is used to acquire multimodal data and perform preprocessing to establish a multimodal dataset;
[0187] A multi-granularity sampling unit is used to perform multi-granularity sampling in each modality based on the multimodal dataset to obtain a cross-modal multi-granularity sample space;
[0188] The optimal granularity unit is used to select a preset number of optimal granularities from the multi-granularity sample space of each modality based on the division of each modality data in the multimodal dataset relative to the preset decision set and granularity, thereby forming an optimal granularity combination.
[0189] The recognition unit is used to perform cross-modal fusion based on the optimal granularity combination, using a discriminative measure and a variable coordination metric, to calculate a composite discriminative measure value, thereby recognizing humorous content and obtaining a humor recognition result.
[0190] Example 3
[0191] The third embodiment of the present invention also provides a cross-modal multi-granularity humor recognition device based on composite discriminative measure, which includes a memory and a processor. The memory stores a computer program, which can be executed by the processor to realize the cross-modal multi-granularity humor recognition method based on composite discriminative measure as described above.
[0192] Example 4
[0193] The fourth embodiment of the present invention also provides a computer-readable storage medium storing computer-readable instructions. When the computer-readable instructions are executed by the processor of the device where the computer-readable storage medium is located, they implement the cross-modal multi-granularity humor recognition method based on composite discriminative measure as described above.
[0194] In the several embodiments provided in this invention, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus and method embodiments described above are merely illustrative; for example, the flowcharts in the accompanying drawings illustrate apparatus and methods according to various embodiments of the invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0195] In addition, the functional modules in the various embodiments of the present invention can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0196] The terminology used in the embodiments of this invention is for the purpose of describing particular embodiments only and is not intended to limit the invention. The singular forms “a,” “the,” and “the” as used in the embodiments of this invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.
[0197] It should be understood that the term "and / or" used in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0198] Depending on the context, the word "if" as used here can be interpreted as "when," "when," "in response to determination," or "in response to detection." Similarly, depending on the context, the phrase "if determination" or "if detection (of the stated condition or event)" can be interpreted as "when determination," "in response to determination," "when detection (of the stated condition or event)," or "in response to detection (of the stated condition or event)."
[0199] The use of "first" and "second" in the embodiments is merely to distinguish similar objects and does not represent a specific ordering of objects. It is understood that "first" and "second" can be interchanged in a specific order or sequence where permitted. It should be understood that the objects distinguished by "first" and "second" can be interchanged where appropriate so that the embodiments described herein can be implemented in an order other than those illustrated or described herein.
[0200] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A cross-modal multi-granularity humor recognition method based on composite discrimination measure, characterized in that, The method comprises the following steps: S1, acquiring multi-modal data and preprocessing, and establishing a multi-modal data set; S2, based on the multi-modal data set, multi-granularity sampling is performed under each modality respectively to obtain a cross-modal multi-granularity sample space; S3, according to the division of each modality data in the multi-modal data set with respect to a preset decision set and granularity, a preset number of optimal granularities that satisfy coordination are selected from the multi-granularity sample space of each modality to form an optimal granularity combination; S4, based on the optimal granularity combination, a cross-modal fusion is performed by using a discrimination measure and a variable coordination measure to calculate a composite discrimination measure value, so as to identify humorous content and obtain a humorous recognition result. The process of calculating the composite discrimination measure value by using the discrimination measure and the variable coordination measure for cross-modal fusion is specifically as follows: The variable coordination metric is used to describe the relationship between the granularity of each modality and the decision; then, the modality The variable coordination metric under the granularity of the calculation formula is: ; wherein, is a predetermined decision set; is the multi-modal dataset; represents dividing all samples in the multi-modal dataset according to the cross-modal multi-granularity sample space; represents dividing all samples in the multi-modal dataset according to the cross-modal multi-granularity sample space; represents dividing all samples in the multi-modal dataset according to the cross-modal multi-granularity sample space; represents the granularity of the i-th modality; The weight coefficients of each modality are calculated by using a variable coordination metric The weight coefficients of each modality The calculation formula is: ; wherein, represents various modalities; represents one of the modalities; represents a preset number of optimal granularities; represents a modality variable coordination measure at a granularity of According to the partition and the partition The conditional entropy of each modality under different granularities is calculated ; Then, the modal The calculation formula of the conditional entropy under the granularity is: ; ; ; wherein, is the probability that is the probability that is the probability that is the conditional probability that is the conditional probability that denotes the cardinality of a set; Based on the conditional entropy of each mode Calculate each mode at different granularities Discrimination measure Then the mode In particle size Discrimination measure The calculation formula is: ; ; wherein, is the probability of the equivalence class . is the conditional entropy of the modal at the granularity . denotes the cardinality of a set. Based on the weight coefficients of each modality, a composite discrimination measure value is calculated through cross-modality fusion The calculation formula of the composite discrimination measure value is as follows: ; wherein represents a modality a discrimination measure at a granularity under.
2. The cross-modal multi-granularity humor recognition method based on composite discrimination measure according to claim 1, characterized in that The multi-modal data comprises text data, and / or image data, and / or audio data, and / or video data.
3. The cross-modal multi-granularity humor recognition method based on composite discrimination measure according to claim 2, characterized in that The preprocessing comprises: word segmentation, and / or stop word removal, and / or standardization processing for text data; size normalization, and / or color space conversion processing for image data; noise reduction, and / or frame division processing for audio data; frame extraction, and / or action extraction processing for video data.
4. The cross-modal multi-granularity humor recognition method based on composite discrimination measure according to claim 1, characterized in that The method further comprises humor annotation and alignment on the preprocessed multi-modal data.
5. The cross-modal multi-granularity humor recognition method based on composite discrimination measure according to claim 1, characterized in that The process of obtaining the cross-modal multi-granularity sample space is specifically as follows: Features of each modality data are extracted based on the multi-modal data set; A plurality of feature subsets are obtained based on the extracted features; Different feature subsets constitute different granularities, and the data set is divided based on different feature subsets to obtain a multi-granularity sample space of each modality; The multi-granularity sample spaces corresponding to each modality data are combined to obtain a cross-modal multi-granularity sample space.
6. The cross-modal multi-granularity humor recognition method based on composite discrimination measure according to claim 1, characterized in that The process of selecting the optimal granularity combination is specifically as follows: acquiring a preset decision set ; According to partitioning all samples in the multi-modal dataset into a set of clusters, denoted as partition ; According to the cross-modal multi-granularity sample space, all samples in the multi-modal data set are divided, denoted as division , represents the granularity level; ; represents the total number of granularities, and the larger the number, the finer the granularity division; represents the granularity of the modal According to the granularity level from coarse to fine, find out the granularity that satisfies the formula in the multi-granularity sample space corresponding to each modality of the formula, obtain the optimal granularity of each modality, and constitute the optimal granularity combination. wherein, when represents the modality on the granularity on the set of decisions is coordinated.
7. The cross-modal multi-granularity humor recognition method based on composite discrimination measure according to claim 1, characterized in that The process of identifying humorous content is specifically as follows: Setting a humor recognition threshold If the calculated composite discrimination measure value is determined to be humorous content, otherwise non-humorous content.
8. A cross-modal multi-granularity humor recognition device based on composite discrimination measure, configured to implement a cross-modal multi-granularity humor recognition method based on composite discrimination measure according to any one of claims 1-7, characterized in that, The method comprises the following steps: An acquisition unit is configured to acquire multi-modal data and perform preprocessing to establish a multi-modal data set; A multi-granularity sampling unit is configured to perform multi-granularity sampling under each modality based on the multi-modal data set to obtain a cross-modal multi-granularity sample space; An optimal granularity unit is configured to select a preset number of optimal granularities that satisfy coordination from the multi-granularity sample space of each modality according to the division of each modality data in the multi-modal data set with respect to a preset decision set and granularity to form an optimal granularity combination; An identification unit is configured to perform cross-modal fusion by using a discrimination measure and a variable coordination measure based on the optimal granularity combination to calculate a composite discrimination measure value, so as to identify humorous content and obtain a humorous recognition result.
Citation Information
Patent Citations
Knowledge graph and cross-modal attention-based multi-modal siphonage detection method
CN114330334A
Health management data mining method and system based on deep learning
CN120015357A