A multi-modal multi-view dispute detection method and system

By combining video and text features, a multimodal, multi-view dispute detection method is used to capture disputes in social media videos. This solves the problems of neglecting visual functions and poor text detection performance in existing technologies, and achieves efficient dispute detection and risk control.

CN118887028BActive Publication Date: 2026-01-16SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411345180.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-26
Publication Date
2026-01-16
Estimated Expiration
2044-09-26

AI Technical Summary

Technical Problem

Existing social media controversy detection methods mainly focus on the text modality, neglecting the potential of visual functions, and perform poorly when the text is missing or incomplete.

Method used

A multimodal, multi-view dispute detection method is adopted, which captures semantic and sentiment inconsistencies in comments through video feature extraction, text feature extraction, integration of multimodal features, context graph learning, and sentiment matrix, and generates dispute detection results using a classifier.

Benefits of technology

It improves the accuracy of controversy detection, enabling timely detection of video controversies when early comment information is limited, reducing public opinion risks, providing video recommendations, and curbing the spread of controversies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118887028B_ABST
    Figure CN118887028B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of social media controversy detection, and provides a multi-modal multi-view controversy detection method and system, video feature extraction and text feature extraction are respectively performed to obtain multi-modal features; the multi-modal features are integrated together to learn overall controversial features of video content; the relationship between the video and the comment is modeled through context graph learning to obtain semantic and structural relationship features between the video and the comment; an emotion matrix and a semantic matrix are used to capture inconsistent features between comment semantics and emotions; the overall controversial features, the semantic and structural relationship features and the inconsistent features are spliced to obtain a controversy detection result through a classifier; the application can effectively model multi-modal video content and capture the interaction between social contexts, and improves the accuracy of controversy detection.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of social media controversy detection, and specifically relates to a multi-modal multi-view controversy detection method and system. BACKGROUND

[0002] The statements in this section merely provide background information related to the present application and do not necessarily constitute the prior art.

[0003] The openness of social platforms has led to the exchange of heated discussions and different opinions, and the spread of some videos can even bring about negative public opinion influence. Therefore, a large number of video diffusion needs to perform risk management and control measures. Previous research on controversy detection mainly focuses on the text modality, ignoring the necessity of incorporating multi-modal in the case of limited discourse information.

[0004] The existing methods for detecting social media controversies mainly focus on utilizing the semantic and structural features of target posts and their comments, which have two problems: first, they ignore the potential of utilizing the visual features available on social media platforms; second, current controversy detection models often perform poorly when faced with incomplete or missing text. SUMMARY

[0005] In order to solve the problems of the prior art, the present application provides a multi-modal multi-view controversy detection method and system, which can effectively model multi-modal video content and capture the interaction between social contexts, improving the accuracy of controversy detection.

[0006] In order to achieve the above purpose, the present application adopts the following technical scheme:

[0007] In a first aspect, the present application provides a multi-modal multi-view controversy detection method.

[0008] A multi-modal multi-view controversy detection method includes the following processes:

[0009] Video feature extraction and text feature extraction are performed respectively to obtain multi-modal features;

[0010] Integrate the multi-modal features together to learn the overall controversial features of the video content;

[0011] Model the relationship between the video and the comments through context graph learning to obtain the semantic and structural relationship features between the video and the comments;

[0012] Utilize the sentiment matrix and the semantic matrix to capture the inconsistency features between the comment semantics and the sentiment;

[0013] The overall controversial feature, the semantic and structural relationship feature and the inconsistency feature are spliced to obtain a controversial detection result through a classifier.

[0014] In a second aspect, the present application provides a multi-modal multi-view controversial detection system.

[0015] A multi-modal multi-view controversial detection system comprises:

[0016] A multi-modal feature extraction unit is configured to perform video feature extraction and text feature extraction respectively to obtain multi-modal features.

[0017] An overall controversial feature acquisition unit is configured to integrate the multi-modal features together to learn an overall controversial feature of video content.

[0018] A semantic and structural relationship feature acquisition unit is configured to model a relationship between a video and a comment through context graph learning to obtain a semantic and structural relationship feature between the video and the comment.

[0019] An inconsistency feature acquisition unit is configured to capture an inconsistency feature between comment semantics and emotions by using an emotion matrix and a semantic matrix.

[0020] A controversial detection result generation unit is configured to splice the overall controversial feature, the semantic and structural relationship feature and the inconsistency feature to obtain a controversial detection result through a classifier.

[0021] Compared with the prior art, the present application has the following beneficial effects:

[0022] 1. The present application splices the overall controversial feature, the semantic and structural relationship feature and the inconsistency feature to obtain a controversial detection result through a classifier, which can help to detect controversies on a social video platform, can provide a reference for video recommendation, can timely suppress controversy spread and reduce public opinion risk.

[0023] 2. The present application can effectively capture controversies in video content itself and controversies generated by interactions between a video and its related comments or between comments themselves, and a large number of experiments prove the effectiveness of the proposed scheme.

[0024] 3. The present application can perform controversy detection in the initial stage of video publishing and timely control the risk of spread under the condition that early comment information is limited.

[0025] The advantages of the additional aspects of the present application will be partially given in the following description, partially become obvious from the following description, or be learned through the practice of the present application. BRIEF DESCRIPTION OF DRAWINGS

[0026] The accompanying drawings, which form a part of this specification, are included to provide a further understanding of the application, and are incorporated into and constitute a part of this specification. The embodiments of the application, and their

[0027] Figure 1 A schematic diagram of the overall process of the multi-modal multi-view controversy detection method provided for Embodiment 1 of the application;

[0028] Figure 2 A schematic diagram of step S1 provided for Embodiment 1 of the application;

[0029] Figure 3 A schematic diagram of step S2 provided for Embodiment 1 of the application;

[0030] Figure 4 A schematic diagram of step S3 provided for Embodiment 1 of the application;

[0031] Figure 5 A schematic diagram of step S4 provided for Embodiment 1 of the application;

[0032] Figure 6 A schematic diagram of step S5 provided for Embodiment 1 of the application;

[0033] Figure 7 A schematic diagram of a multi-modal multi-view controversy detection system provided for Embodiment 2 of the application. DETAILED DESCRIPTION

[0034] The application will be further described below with reference to the drawings and embodiments.

[0035] It should be noted that the following detailed description is exemplary in nature and is intended to provide further description of the application. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs.

[0036] The embodiments in the application and the features in the embodiments can be combined with each other without conflict.

[0037] Embodiment 1:

[0038] The present implementation proposes a multi-modal multi-view controversy detection method, which aims to detect whether a given video and its related content contain controversy. Based on multi-modal feature extraction, the controversy between the video itself, the video and the comments, and the comments and the comments is modeled, and applied to the controversy detection of social media videos. First, the technical terms and related concepts involved in this processing scheme are briefly introduced, including:

[0039] CN-CLIP: Chinese version of CLIP model, trained on large-scale Chinese data (database contains 200 million image-text pairs), visual side skeleton uses ViT-H / 14, and text side skeleton uses RoBERTa-wwm-Large.

[0040] ViT (Vision Transformer): A deep learning model based on Transformer architecture for image recognition and computer vision tasks, unlike traditional convolutional neural networks (CNN), ViT directly treats images as a sequential input and uses self-attention mechanisms to handle pixel relationships in images.

[0041] Transformer: A deep learning model architecture for natural language processing and sequence-to-sequence tasks that introduces self-attention mechanisms and multi-head attention mechanisms, allowing it to consider all positions in the input sequence simultaneously.

[0042] RoBERTa (Robustly Optimized BERT Pretraining Approach): A pre-training language model based on the BERT model, proposed by Facebook AI in 2019.

[0043] MoE (Mixture of Expert): A deep learning model architecture based on sparse MoE layers, which splits large models into multiple small models (experts), and each iteration decides to activate a part of the experts for calculation based on the sample, achieving the effect of saving computing resources; and introduces a trainable gate mechanism to ensure sparsity, to ensure the optimization of computing power.

[0044] GCN (Graph Convolution Network): Graph convolution network, a convolutional neural network that can directly act on graphs and utilize their structural information.

[0045] ffmpeg: A free software project initiated by programmer Fabrice Bellard, aiming to provide tools and libraries for handling multimedia data.

[0046] MaxPooling: Max pooling, the entire image is divided into several non-overlapping small blocks of the same size, and only the maximum number in each small block is taken, and after discarding other nodes, the original planar structure is maintained to obtain the output.

[0047] SenticNet: Concept-level sentiment analysis, which uses semantics and linguistics to complete tasks such as polarity detection and sentiment recognition, rather than simply relying on word co-occurrence frequency.

[0048] ReLU (Rectified Linear Unit): A linear rectifier function, also known as a rectified linear unit, is a commonly used activation function in artificial neural networks, usually referring to a non-linear function represented by a ramp function and its variants.

[0049] Dropout: A commonly used regularization method that reduces overfitting by randomly setting some neuron outputs to zero.

[0050] Frobenius norm: Also known as F-norm, it is a matrix norm, denoted as ||·||F. F The Frobenius norm of matrix A is defined as the sum of the absolute value of each element of matrix A, that is, it can be used to approximate a single data matrix using a low-rank matrix.

[0051] Bi-LSTM: Bidirectional Long Short-Term Memory Network, a variant of Recurrent Neural Network (RNN), unlike ordinary RNN network, LSTM adds three gate units: forget gate, input gate and output gate, which can selectively store and clean data, effectively solving the problems of gradient disappearance and gradient explosion. Bi-LSTM is composed of two sets of LSTM networks with opposite directions, which is used to model the context information with time sequence relationship.

[0052] Specifically, the multi-modal multi-view controversy detection method includes the following processes:

[0053] S1: Multi-modal feature extraction, using multi-modal pre-training model CN-CLIP to extract multi-modal features in video and text, as shown in Figure 1 ;

[0054] S2: Modality perception learning, integrating multiple modalities and aligning features, learning the overall controversial features of video content, considering that different modes have different effects on different audiences, and must address the challenges caused by inconsistent and insufficient attention to these modes; In order to overcome these challenges, the Mixture of Expert (MoE) architecture is used to enhance the overall modeling performance;

[0055] S3: Context graph learning, modeling the relationship between video and comments, in order to effectively capture the controversy between video and comments, after aligning the features, Graph Convolution Network (GCN) is used to capture the semantic and structural relationship between the two modalities;

[0056] S4: inconsistency reinforcement learning, which focuses on capturing the inconsistency between review semantics and sentiment, given the inherent inconsistency observed in controversial reviews, which usually includes the difference between content and emotion, an inconsistency matrix (i.e., a sentiment matrix and a semantic matrix) is calculated, and fusion calculation is performed using the sentiment matrix and the semantic matrix to capture and simulate these inconsistencies;

[0057] S5: integration and prediction, splicing the features output by S2-S4, and judging whether it is controversial through a classifier.

[0058] In step S1 of the present implementation, as shown in Figure 2 , specifically, it includes:

[0059] S11: In order to extract frame-level features, key frames are extracted from each video using the ffmpeg tool, and the extracted key frames are input into the image encoder Vision Transformer (ViT) in CN-CLIP to generate video features, represented as , where represents the feature vector extracted from the i-th key frame, represents the total number of key frames in the video, represents the encoded image dimension.

[0060] S12: The structured text content includes video description, publisher profile, ASR text, and review content, and the text features are extracted using RoBERTa in CN-CLIP, and the extracted features are represented as , where represents the dimension of the encoded text, represents the feature vector of the i-th review in , and represents the total number of reviews.

[0061] In step S2 of the present implementation, as shown in Figure 3 , specifically, it includes:

[0062] S21: The features obtained in step S12 are aligned through a separate fully connected layer, and then these aligned features are connected and input into a single Transformer layer (i.e., TL in equation (2)) to capture time information, and the calculations involved in this process are as follows:

[0063] (1);

[0064] (2);

[0065] where, , ​​, and These represent aligned video features, aligned video description features, aligned publisher profile features, and aligned ASR text features, respectively. , , and These represent video features, video description features, publisher profile features, and ASR text features, respectively. , , and The weight parameters represent the video features, video description features, publisher profile features, and ASR text features, while , , and These represent the bias parameters corresponding to the weight parameters of video features, video description features, publisher profile features, and ASR text features, respectively. This is a concatenation function.

[0066] S22: The MoE architecture consists of multiple expert layers (represented as...) (Taking three expert layers as an example) and a gated layer (called The output features obtained in step S21 are composed of... Input to the gating layer and expert level To obtain the output:

[0067] (3);

[0068] (4);

[0069] in, Indicates learnable parameters, It is a function used to determine input features Before choosing The highest threshold value, This indicates the ordinal position of the expert layer, and m represents the number of expert layers. Indicates the first One gate layer, Indicates the first A group of experts, It is a normalized exponential function.

[0070] S23: Apply a balanced loss to the MoE architecture to ensure fair load and importance among experts, calculated as follows:

[0071] (5);

[0072] in, denotes the coefficient of variation, is a smoothing function, and and are used to balance the importance and load of experts.

[0073] In step S3 of the present implementation, as shown in Figure 4 , specifically, it includes:

[0074] S31: comment features are aligned through a fully connected layer to establish unified features:

[0075] (6);

[0076] wherein, denotes a weight parameter, denotes a bias parameter.

[0077] S32: a video context graph is constructed for each video, denoted as , wherein the node set is composed of features of videos or comments, and there is an edge between a video and its corresponding comment when the comment is associated with the specific video. The initial representation of the node can be defined as:

[0078] (7);

[0079] In the present implementation, the center node of the video features processed through the fully connected layer is input into the GCN, and the adjacent node of the comment features processed through the fully connected layer is input into the GCN.

[0080] S33: in the message passing process, each node updates its representation according to the aggregated information obtained from its adjacent nodes and their features, so that the learned representation can contain information from the content and structure of the graph. Specifically, for a given node , the update rule can be represented as:

[0081] (8);

[0082] wherein denotes the hidden state of node in the GCN layer, denotes the hidden state of node in the GCN layer, denotes a rectified linear unit (ReLU) activation function, denotes the neighbors of node (including the node itself), is an aggregation function, is an aggregation function, is an aggregation function, denotes the bias term.

[0083] At the layer level, the embedding vectors are used as input to a two-layer GCN, resulting in a condensed representation denoted as , which is computed as follows: The incoming messages from the neighbor set are aggregated by a function

[0084] , which is implemented as a linear function. Thus, for the l-th layer, the propagation rule is as follows:

[0085] where contains all node vectors of the l-th layer, is the normalized adjacency matrix, is the weight matrix, denotes the bias parameter of the l-th layer.

[0086] S34: Finally, the most important features are extracted using MaxPooling in order to allow for subsequent computations:

[0087] (10).

[0088] In step S4 of the present implementation, as shown in Figure 5 , specifically, includes:

[0089] S41: In order to capture and analyze sentiment inconsistencies in the reviews, a sentiment matrix is constructed based on a set of reviews , where each element is computed as follows:

[0090] (11).

[0091] where denotes the sentiment score of review computed using the external sentiment lexicon SenticNet, denotes the sentiment score of review computed using the external sentiment lexicon SenticNet, denotes the absolute value computation.

[0092] By employing this approach, the weight of the corresponding edge increases with the degree of sentiment reversal between each two reviews, thus allowing to focus the attention on reviews that exhibit opposite sentiments.

[0093] S42: Next, the semantic matrix ​to measure the semantic inconsistency between reviews. Specifically, for each pair of review features The attention score is calculated as:

[0094] (12).

[0095] where is a trainable parameter matrix, denotes matrix transpose.

[0096] S43: Combining the learned sentiment matrix and semantic matrix to obtain a more expressive representation (i.e., the fusion matrix ):

[0097] (13).

[0098] where is a hyperparameter.

[0099] S44: In addition, we added a regularization loss to improve the quality of the learned semantic information:

[0100] (14).

[0101] where is the Frobenius norm of a matrix, is a sparsity hyperparameter.

[0102] S47: To comprehensively analyze the review features , a bidirectional long short-term memory (Bi-LSTM) model is used to capture the context information :

[0103] (15).

[0104] S46: To consider both inconsistencies simultaneously, the context information is processed through a fully connected layer and then combined with using an inner product operation to obtain the output:

[0105] (16).

[0106] where is a weight matrix.

[0107] S47: Subsequently, the MaxPooling function is applied to derive the final output :

[0108] (17).

[0109] In step S5 of the present implementation, as shown in Figure 6 , specifically, includes:

[0110] S51: In order to obtain the final integrated output from the three modules, first , , connect (i.e. Figure 6 output feature 1, output feature 2 and output feature 3) in

[0111] (18);

[0112] S52: The output of step S51 is passed to the constructed classifier to obtain the final output. The classifier consists of a stack of layers, including two fully connected layers, and after the fully connected layers, a normalization layer, a ReLU activation function layer and a Dropout layer are added. The final probability distribution is calculated as follows:

[0113] (19);

[0114] where , , and are model parameters, denotes the layer normalization function. The probability matrix includes and , which represent the predicted probabilities of the label being 0 (no dispute) and 1 (dispute), respectively.

[0115] S53: The predicted label is defined as:

[0116] (20);

[0117] S54: We use the cross-entropy loss function to measure the difference between the predicted probability and the true label, assuming denotes the true label, then the loss function is calculated as follows:

[0118] (21);

[0119] S55: Finally, the three loss functions in steps S23, S44 and S54 are added to obtain the final loss function:

[0120] (22).

[0121] Example 2:

[0122] As shown in Figure 7As shown, the present implementation provides a multi-modal multi-view controversy detection system, comprising:

[0123] A multi-modal feature extraction unit is configured to perform video feature extraction and text feature extraction respectively to obtain multi-modal features; the specific working process of this unit is described in step S1 of embodiment 1, which will not be repeated here.

[0124] An overall controversial feature acquisition unit is configured to integrate the multi-modal features together to learn the overall controversial features of the video content; the specific working process of this unit is described in step S2 of embodiment 1, which will not be repeated here.

[0125] A semantic and structural relationship feature acquisition unit is configured to model the relationship between the video and the comments through context graph learning to obtain the semantic and structural relationship features between the video and the comments; the specific working process of this unit is described in step S3 of embodiment 1, which will not be repeated here.

[0126] An inconsistency feature acquisition unit is configured to capture the inconsistency features between the comment semantics and the emotions using the emotion matrix and the semantic matrix; the specific working process of this unit is described in step S4 of embodiment 1, which will not be repeated here.

[0127] A controversy detection result generation unit is configured to concatenate the overall controversial features, the semantic and structural relationship features, and the inconsistency features to obtain a controversy detection result through a classifier; the specific working process of this unit is described in step S5 of embodiment 1, which will not be repeated here.

[0128] It can be understood that the above-mentioned units can be combined into one or several other units respectively or entirely, or some of the units can be further split into a plurality of units with smaller functions to constitute, which can achieve the same operation without affecting the implementation of the technical effects of the embodiments of the present application. The above-mentioned units are divided based on logical functions, and the functions of one unit can also be realized by multiple units, or the functions of multiple units can be realized by one unit. In other embodiments of the present application, the multi-modal multi-view controversy detection system can also include other units, and these functions can also be realized by other units in actual application, and can be realized by multiple units in cooperation.

[0129] According to another embodiment of the present application, the system described in the embodiment and the multi-modal multi-view dispute detection method of the embodiments of the present application can be constructed by running a computer program (including program codes) capable of performing each step involved in the corresponding method described in Embodiment 1 on a general computing device such as a computer including processing elements and storage elements such as a Central Processing Unit (CPU), a Random Access Memory (RAM), a Read Only Memory (ROM), etc., and the computer program can be recorded on a computer readable recording medium, loaded into the above computing device through the computer readable recording medium, and run therein.

[0130] The above merely provides preferred embodiments of the present application but should not be used to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A multi-modal multi-view dispute detection method, characterized in that, The method comprises the following processes: respectively performing video feature extraction and text feature extraction to obtain multi-modal features; aligning the multi-modal features through separate fully connected layers, then connecting the aligned features and inputting them into a single Transformer layer to capture time information, inputting the captured time information into a gating layer and an expert layer to obtain the overall controversial features; capturing the temporal information input to the gating layer and the expert layer obtaining the overall contentiousness feature comprising: ; ; in, Represents the learnable parameters. It is a function used to determine input features Before choosing The highest threshold value, This indicates the ordinal position of the expert layer, and m represents the number of expert layers. Indicates the first One gate layer, Indicates the first A group of experts, It is a normalized exponential function; modeling the relationship between the video and the comment through context graph learning to obtain semantic and structural relationship features between the video and the comment; capturing the inconsistency features between the comment semantics and the sentiment by using a sentiment matrix and a semantic matrix; splicing the overall controversial features, the semantic and structural relationship features, and the inconsistency features to obtain a controversy detection result through a classifier.

2. The multi-modal multi-view controversy detection method according to claim 1, wherein the multi-modal features comprise video features, video description features, publisher profile features, speech recognition text features, and comment content features.

3. The multi-modal multi-view controversy detection method according to claim 1, wherein a balanced loss is used for the expert layer, comprising:

4. The multi-modal multi-view controversy detection method according to claim 1, wherein a video context graph is constructed for each video, the node set in the video context graph is composed of features of the video or the comment, there is an edge between the video and the corresponding comment when the comment is associated with the specific video, and the initial representation of the node in the node set is defined as:

5. The multi-modal multi-view controversy detection method according to claim 1, wherein the expert layer comprises: ; wherein, denotes the coefficient of variation, is a smoothing function, hyperparameters and for balancing expert importance and load.

6. The multi-modal multi-view controversy detection method according to claim 1, wherein the expert layer comprises: Review features First align through a fully connected layer to establish uniform features: ; wherein, denotes a weight parameter, denotes a bias parameter, denotes a rectified linear unit activation function; 7. The multi-modal multi-view controversy detection method according to any one of claims 1-6, comprising: ; wherein, representing aligned video features, using embedding vectors as input to a two-layer GCN, resulting in a condensed representation ; extracting condensed representations using max pooling the semantic and structural relationship features . a multi-modal feature extraction unit configured to respectively perform video feature extraction and text feature extraction to obtain multi-modal features; Based on a set of reviews Constructing a sentiment matrix , For the 1st review, For the 2nd review, For the th review, For the total number of reviews, each element in the sentiment matrix is calculated as: , represents the sentiment score of the review calculated using an external sentiment lexicon, represents the sentiment score of the review calculated using an external sentiment lexicon, represents the absolute value calculation; computing a semantic matrix to measure semantic inconsistency between reviews, for each pair of review features , the attention score , is a trainable parameter matrix, denotes matrix transpose; combining the learned sentiment matrix and semantic matrix to obtain a more expressive representation: where is a hyperparameter; For the review features , a bidirectional long short-term memory model is used to capture context information , is the first context information, is the context information; The output is obtained using an inner product operation: wherein is a weight matrix; max-pooling the output of the inner product operation to obtain the inconsistent feature ; A regularization loss is added to improve the quality of the learned semantic information: wherein, is the Frobenius norm of a matrix, is a sparsity hyper-parameter. an overall controversial feature acquisition unit configured to integrate the multi-modal features together to learn overall controversial features of the video content, align the multi-modal features through separate fully connected layers, then connect the aligned features and input them into a single Transformer layer to capture time information, input the captured time information into a gating layer and an expert layer to obtain the overall controversial features; The classifier includes two fully connected layers, a normalization layer, a ReLU activation function layer and a Dropout layer are added after the fully connected layers, and the final probability distribution includes: wherein, , , and are model parameters, represents a layer normalization function, and a probability matrix includes and , respectively, represent that the label is 0 and 1, 0 represents a non-controversial prediction probability, and 1 represents a controversial prediction probability, is a splicing result of the overall controversial feature, the semantic and structural relationship feature and the inconsistency feature. The cross-entropy loss function is used to measure the difference between the predicted probabilities and the true labels, assuming that y represents the true labels, then the cross-entropy loss function is computed as follows: ​​ a semantic and structural relationship feature acquisition unit configured to model the relationship between the video and the comment through context graph learning to obtain semantic and structural relationship features between the video and the comment; The overall loss function is wherein, is the balancing loss for the expert layer, is the regularization loss, is the cross-entropy loss function.

8. A multi-modal multi-view dispute detection system, characterized in that, an inconsistency feature acquisition unit configured to capture the inconsistency features between the comment semantics and the sentiment by using a sentiment matrix and a semantic matrix; a controversy detection result generation unit configured to splice the overall controversial features, the semantic and structural relationship features, and the inconsistency features to obtain a controversy detection result through a classifier. ​ capturing the time information input to the gating layer and the expert layer obtaining the overall contentiousness feature comprising: ; ; in, Represents the learnable parameters. It is a function used to determine input features Before choosing The highest threshold value, This indicates the ordinal position of the expert layer, and m represents the number of expert layers. Indicates the first One gate layer, Indicates the first A group of experts, It is a normalized exponential function; ​ ​ ​

Citation Information

Patent Citations

  • Video comment generation method, system and device and storage medium

    CN114339450A

  • Exploiting multi-modal affect and semantics to assess the persuasiveness of a video

    US20160328384A1