Short video platform-oriented multi-modal event information propagation strategy identification method and system
Through multimodal data analysis and hierarchical strategy identification architecture, combined with video image, audio and text features, the accuracy and comprehensiveness of publicity strategy identification on short video platforms are solved, and transparent and interpretable efficient recognition effect is achieved.
Patent Information
- Application Number
- CN202510339967.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-25
AI Technical Summary
It is difficult for the existing technology to fully identify publicity strategies on short video platforms, especially when it faces complex contexts and cross-document context associations, and deep learning methods have problems such as difficult to explain and lack of generalization capabilities.
A multimodal data analysis method is adopted, combining video images, audio and text features, and the improved MTCNN+ResNet50 model detects facial position and expression changes, introduces a feature fusion network of attention mechanisms, and builds a hierarchical strategy recognition architecture, including recognition systems in multiple dimensions such as emotional manipulation and false information.
It improves the accuracy and comprehensiveness of publicity strategy identification, provides transparent and interpretable analysis results, is highly adaptable, can quickly adapt to new application scenarios and strategy types, and overcomes the limitations of single-dimensional analysis and specific field construction.
Smart Images

Figure CN120378656A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of information security and data analysis, and particularly relates to a method and system for identifying multi-modal event information dissemination strategies for short video platforms. Background Art
[0002] Short video platforms have become one of the core channels for Internet content dissemination today. With their unique characteristics of being short, concise, and fast in dissemination, they have attracted a large number of users and advertisers. Globally, short videos not only play an important role in fields such as entertainment, education, and news dissemination, becoming an important tool for social event dissemination, but also become potential carriers of information manipulation and malicious behaviors.
[0003] By identifying the publicity strategies of event information and then identifying malicious behaviors in the process of social event dissemination, it can help short video platform operators establish effective early warning and intervention mechanisms, avoid the spread of bad information, and improve the security and fairness of platform content. For the identification of event information publicity strategies, the existing technologies mainly include publicity strategy identification methods based on sentiment analysis and tendency judgment, and publicity strategy identification methods based on deep learning. The publicity strategy identification methods based on sentiment analysis and tendency judgment are difficult to accurately capture complex context information and implicit publicity intentions because they mainly rely on preset sentiment dictionaries and rules, and their dictionaries and rules are often constructed for specific fields and have insufficient generalization ability when facing new fields or expressions. The publicity strategy identification methods based on deep learning, although having strong feature extraction capabilities, have problems such as dependence on large-scale labeled data, insufficient model interpretability, and limited context understanding. These problems mainly stem from the high cost of obtaining training data for deep learning models, the difficulty of explaining the decision basis as the model runs as a black box, and the limitations in dealing with cross-document context associations. Summary of the Invention
[0004] The main purpose of the present invention is to overcome the problem that it is difficult to comprehensively identify publicity strategies in the process of short video dissemination in the prior art, and propose a method and system for identifying multi-modal event information dissemination strategies for short video platforms. Through multi-modal data analysis, by integrating the characteristics of video images, audio, and text, publicity strategies can be accurately and comprehensively identified.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] A method for identifying multi-modal event information dissemination strategies for short video platforms includes the following steps:
[0007] S1. Data collection, obtaining data from short video platforms, including video image data, audio data, subtitle text data, and user interaction data;
[0008] S2. Data processing, which processes the collected video image data, audio data, and subtitle text data, and establishes a feature library including video image data, audio data, subtitle text data, and user interaction data;
[0009] S3. Define publicity strategies and classify them;
[0010] S4. Conduct targeted analysis according to the feature requirements of different publicity strategies to identify event information dissemination strategies.
[0011] The present invention also includes a multi-modal event information dissemination strategy identification system for short video platforms. The system adopts the multi-modal event information dissemination strategy identification method provided by the present invention. The system includes a data acquisition module, a data processing module, a publicity strategy module, and a dissemination strategy identification module;
[0012] The data acquisition module is used to obtain data from short video platforms, including video image data, audio data, subtitle text data, and user interaction data;
[0013] The data processing module is used to process the collected video image data, audio data, and subtitle text data, and establish a feature library including video image data, audio data, subtitle text data, and user interaction data;
[0014] The publicity strategy module is used to define publicity strategies and classify them, including emotional manipulation and emotional polarization strategies, false information and misleading information dissemination strategies, image and sound manipulation strategies, information overload and repeated reinforcement strategies, social proof and group effect strategies, visual hint and symbol transmission strategies, and clickbait strategies;
[0015] The dissemination strategy identification module conducts targeted analysis according to the feature requirements of different publicity strategies to achieve the identification of event information dissemination strategies.
[0016] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0017] 1. The present invention breaks through the limitations of the prior art that only relies on single - dimension analysis by introducing a multi - modal data analysis method and combining the comprehensive features of video images, audio, and text, significantly improving the accuracy and comprehensiveness of propaganda strategy recognition; by establishing a complete strategy recognition system, including multiple dimensions such as emotional manipulation, false information, and image - sound manipulation, it solves the problem of lack of systematicness in the prior art; at the same time, the present invention adopts a feature fusion network with an attention mechanism, overcomes the shortcoming that the black box in traditional deep - learning methods is difficult to explain, and provides more transparent and interpretable analysis results. The modular design enables the system to have good adaptability and scalability, can quickly adapt to new application scenarios and types of propaganda strategies, and overcomes the limitations of being constructed for specific fields in the prior art.
[0018] 2. In terms of facial expression analysis, the present invention adopts an improved MTCNN + ResNet50 model architecture, which can not only detect the facial position but also recognize subtle expression changes, while the prior art mainly relies on preset emotion dictionaries and rules and cannot capture complex visual context information.
[0019] 3. In terms of data fusion, the present invention innovatively introduces a feature fusion network with an attention mechanism, which can adaptively adjust the weights of visual, audio, and text features, while existing deep - learning methods often use simple feature concatenation or average fusion and are difficult to highlight the importance of key information.
[0020] 4. Hierarchical strategy recognition architecture; the present invention constructs a two - level classification system from basic strategy types to specific techniques. Through the combined analysis of multi - dimensional features, it can accurately locate different types of propaganda strategies. In contrast, the prior art mostly uses a single classification model, lacking a systematic understanding and classification of propaganda strategies, resulting in one - sided or inaccurate recognition results; in the detection of social iconic elements, the present invention also particularly introduces a special visual symbol analysis module, which makes important supplements and innovations to the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] The present invention will be further described in detail below in conjunction with embodiments and the accompanying drawings, but the embodiments of the present invention are not limited thereto.
[0023] Embodiment
[0024] As Figure 1 shown, the multi - modal event information dissemination strategy recognition method of the present invention for short - video platforms includes the following steps:
[0025] S1. Data collection: Obtain data from short video platforms, including video image data, audio data, subtitle text data, and user interaction data (such as comments, likes, shares, etc.); specifically including:
[0026] Obtain video image data, specifically:
[0027] Extract key-frame images of the video, retaining important information such as scenes, objects, and colors; assume the short video stream is represented as S = {V1, V2, …, V n}, where V n represents the video frame arriving at time n.
[0028] Initialize the parameters of the video data processing method, including the video frame set F, the video frame feature library F feat , the image data I wt of each video frame, and the emotion label E wt of each frame; the video frame set F is used to store all video frame data extracted from the short video.
[0029] Regard each video frame V t in the short video stream as the image data I t , input it into the video frame set F, and update the image feature library F t using the video frame V feat , where F feat = {I t , C t , E t}, I t is the video frame image data, containing RGB color values, representing the visual information in the video, C t is the hue information, representing the color distribution or hue of the video frame, extracted through a color histogram or average hue, and E t is the emotion label, representing facial expressions, emotional tendencies, etc. in the video frame;
[0030] Obtain audio data, specifically including extracting the audio part in the video, performing emotion recognition and speech-to-text processing through audio analysis, and extracting information such as the emotional color, speech rate, and pitch change of the speech.
[0031] Obtain subtitle text data, specifically including information such as the title, description, subtitles, and comments of the video, and perform text analysis through natural language processing techniques (such as word segmentation, emotion analysis, keyword extraction, etc.) to identify the emotional tendency and topic focus of the video.
[0032] Obtain user interaction data, specifically by collecting the behavior data of the audience, including the number of likes, comments, and forwards, etc.
[0033] S2. Data processing: Process the collected video image data, audio data, and subtitle text data to establish a feature library that includes video image data, audio data, subtitle text data, and user interaction data. The processing of video image data includes facial expression detection, hue and color analysis, and social iconic element detection.
[0034] Among them, facial expression detection includes face detection and expression classification.
[0035] Face detection is used to identify the position of the face from video frames. Specifically:
[0036] Input the video frame image I t , I t represents a video frame arriving at a time point t, represented in the form of an image matrix, I t ∈R H×W×3 , where H is the height of the image, W is the width of the image, and 3 represents the three RGB color channels.
[0037] Each pixel point I t (i, j) of the image consists of the intensity values of red, green, and blue colors: I t (i, j) = {I t (r) (i, j), I t (g) (i, j), I t (b) (i, j)}, where I t (r) , I t (g) , I t (b) represent the pixel values of the red, green, and blue channels in the image at position (i, j), respectively.
[0038] Convert the color image to a grayscale image. The specific calculation formula is as follows:
[0039] I t (gray) (i, j) = 0.2989 × I t (r) (i, j) + 0.587 × I t (g) (i, j) + 0.114 × I t (b) (i, j)
[0040] Among them, I t (gray) (i, j) represents the pixel value at the (i, j) position in the grayscale image.
[0041] Normalize the grayscale values to scale them between 0 and 1, I t (norm) (i,j) = I t (gray) (i,j) / 255;
[0042] Use MTCNN for face region detection. MTCNN is a cascaded multi-task convolutional neural network for face detection and key point localization, consisting of three cascaded sub-networks, namely P-Net, R-Net, and O-Net. Among them, P-Net is the first-stage network, adopting a shallow CNN structure of 12×12×3. It generates candidate face bounding boxes by sliding window scanning on the image pyramid, and the output includes face classification scores, bounding box regression vectors, and facial key point coordinates, and uses non-maximum suppression to remove overlapping candidate boxes. R-Net is the second-stage network, adopting a CNN structure of 24×24×3. It performs more detailed feature extraction and classification on the candidate bounding boxes output by P-Net, outputs more accurate face classification scores and bounding box positions, and further removes mis-detected boxes through more strict threshold screening. O-Net is the last-stage network, adopting a deep CNN structure of 48×48×3. It not only performs face detection but also locates five facial key points (eyes, nose, and the two corners of the mouth), outputs the final face detection bounding box, confidence score, and facial key point coordinates, and finally obtains the final detection result using a more strict non-maximum suppression threshold;
[0043] When performing face detection, first construct an image pyramid S = {I1, I2,..., I n} where I i represents the i-th layer image, and the scaling ratio between adjacent layers is scale_factor. Set a sliding window of size W win ×H win for each layer of the image. This window starts from the top-left corner (0, 0) of the image and slides within the image range [0, W i -W win ×[0, H i -H win with step sizes stride_x and stride_y, where W i and H i are the width and height of the i-th layer image respectively;
[0044] The sub-image region I sub(x,y) ∈R Wwin×Hwin×3The input is processed by the P-Net, and the P-Net extracts features from the input through a three-layer cascaded convolutional structure. In the three-layer cascaded convolutional structure, the first layer of convolution uses a convolutional kernel W1 with a size of 3×3×3×10 to process the input to obtain a feature map F1∈R10×10×10. After PReLU activation and max pooling, it is input into the second layer. The second layer uses a convolutional kernel W2 with a size of 3×3×10×16 to obtain a feature map F2∈R4×4×16. The third layer uses a convolutional kernel W3 of 3×3×16×32 to finally obtain a feature map F3∈R2×2×32, that is, F sub(x,y) ;
[0045] The feature maps obtained by feature extraction are then input into two parallel fully connected layers, which are respectively used for face classification to obtain a probability p(x,y)∈[0,1] and bounding box regression to obtain an offset Δbox(x,y)=(Δx1,Δy1,Δx2,Δy2);
[0046] For positions where the probability exceeds the threshold θ p , according to the predicted offset, calculate the candidate box box(x,y)=T(x,y,W win )+α·Δbox(x,y), where T is the position mapping transformation function and α is the regression coefficient. Thus, a set of candidate boxes B={b1,b2,...,b m} and the corresponding confidence set P score ={p1,p2,...,p m} are obtained;
[0047] To remove overlapping detection boxes, the non-maximum suppression method is used. First, calculate the intersection-over-union ratio of any two boxes B1 and B2:
[0048]
[0049] When its value exceeds the preset threshold Threshold, keep the box with a higher confidence;
[0050] The candidate boxes after the preliminary screening are scaled to 24×24 and then input into the R-Net. Through a deeper convolutional network, features F r ∈R 128 are extracted, and a more stringent threshold θr is used for secondary screening and position refinement;
[0051] Finally, the candidate boxes are scaled to 48×48 and then input into the O-Net. The O-Net outputs the final detection box bboxO-Net, the positions of 5 key points landmarks, and the confidence p through five layers of convolution and three parallel fully connected layers. Each final face detection result is represented as Bi=(x1,y1,x2,y2,p,landmarks), which includes the coordinates of the upper left and lower right corners of the detection box, the confidence, and the set of facial key point coordinates;
[0052] Facial expression classification is used to analyze the detected facial regions and identify the corresponding expressions according to facial features, including the following steps:
[0053] Input the image I of the facial region face , and use an improved cascaded regression facial landmark detector based on OpenCV to extract facial key points. This detector adopts a multi-level cascaded regression structure, and each level contains two sub-modules: local feature extraction and shape regression. First, normalize the input image to the standard size N×N, where N = 96, and then extract the coordinate points of the key features in the facial key regions including eyes, eyebrows, lips, nose, etc. For each key point k i =(x i ,y i ), obtain the key point through the following iterative optimization formula:
[0054]
[0055] where T i (x i ,y i ) is a predefined facial region template, using a shape-constrained local response map as the template feature, and k i is the coordinate of the facial key point; in each iteration, extract the local feature based on the currently estimated key point position and update the position estimate until convergence or the maximum number of iterations is reached;
[0056] Finally, obtain the set of extracted facial key point coordinates {k1, k2,..., k n}, where the value of n is 68. These facial key points precisely describe the geometric structure and morphological features of the face;
[0057] Generate the feature vector of the expression by calculating the local features within each facial region; the calculated feature vector is f face =[f1, f2,…, f m ; where f1, f2,…, f m are the respective features extracted from the facial region;
[0058] Adopt a convolutional neural network CNN to automatically learn features from the facial image for facial expression classification; in the convolutional neural network CNN, the network structure includes a convolutional layer, a pooling layer, and a fully connected layer, and finally outputs a probability distribution representing the probability of each expression:
[0059]
[0060] where W i is the weight corresponding to each expression category, and P(emotion i ∣Iface ) represents the image I face The probability belonging to the expression i;
[0061] Output the predicted expression category y through the classification model emotion and the predicted confidence P(emotion i ) for each expression category, y emotion ∈{smile,angry,sad,…}, P(emotion i ) represents the recognition probability of this expression.
[0062] The hue and color analysis is specifically as follows:
[0063] For the convenience of analysis, first convert the RGB image to the HSV color space, specifically:
[0064] Obtain the maximum and minimum values in RGB: C max = max(R,G,B), C min = min(R,G,B), where C max and C min are the brightest and darkest color values in the three RGB channels respectively;
[0065] Calculate the hue H through the maximum value in RGB, and the formula is as follows:
[0066]
[0067] If the calculated H value is less than 0, add 360 to it to ensure that the hue is in the range of [0,360);
[0068] Calculate the saturation S, and the formula is as follows:
[0069]
[0070] If C max = 0 then S = 0;
[0071] Calculate the brightness V, V = C max ;
[0072] Finally, realize the conversion of the RGB image to the HSV color model.
[0073] According to the converted HSV color model, conduct hue analysis, and judge whether the image has a certain publicity strategy according to the hue distribution; for example, red is usually used to arouse emotions such as alertness and passion; blue is often used for the transmission of calmness and trust; green is mostly used for the publicity of concepts such as health, nature, and environmental protection.
[0074] Saturation analysis, by analyzing the saturation changes in an image, to identify whether there are certain emotions or strategies conveyed through color harmonization; for example, high saturation may be used to evoke strong emotional responses from the audience (e.g., bright red or yellow); low saturation may represent a certain low-key, deep or calm strategy.
[0075] Brightness analysis, brightness affects the visual perception of an image, and by brightness analysis to judge whether the image wants to convey a specific atmosphere; for example, bright colors may give people a positive and energetic impression. Dark tones may be used to express a heavy, serious or mysterious atmosphere.
[0076] In short videos, the gradual change, change or strong contrast of colors can be a deliberately used promotional strategy. For example, to attract the audience's attention or emotional fluctuations through the transition or contrast of colors. By analyzing the color changes in the video frame sequence, color change features can be extracted and combined with the results of emotional analysis to identify possible promotional strategies. For example, the change of hue: for example, the transition from cool tones (such as blue) to warm tones (such as red) may mean a turning point in mood or plot, conveying a certain emotional change. The fluctuation of saturation: rapid saturation changes may be used to arouse the audience's emotional fluctuations, for example, the transition from low saturation to high saturation may cause the concentration of attention. The jump of brightness: to attract the audience's attention through bright images, or to express a certain heavy or dangerous atmosphere through dim scenes.
[0077] The detection of social iconic elements is specifically as follows:
[0078] By detecting specific social iconic elements in video images, to identify emotions, values or cultural symbols related to promotional strategies; social iconic elements are usually images, logos, symbols, words, figures, etc. with wide recognition or symbolic meaning in society.
[0079] Use a CNN-based method to detect social iconic elements, collect image data with iconic elements, and annotate each image to mark the areas of social iconic elements in it. These data will be used to train the convolutional neural network CNN model; the CNN model automatically learns the features in the images and extracts high-level features related to the iconic elements;
[0080] Through the trained CNN model, for the input image I t make predictions, output the category and location of the iconic elements. If a specific logo or symbol is recognized, the model will output the coordinate box and category label of the element. The output includes:
[0081] The bounding box (x min , y min , x max,y max ), indicating the position of the landmark element in the image;
[0082] The class label C corresponding to each detected landmark element label ; for example, national emblems, political symbols, etc.;
[0083] For each recognized element, the confidence P(C label ∣I t ) of the output model, indicating the probability that the element belongs to a certain class.
[0084] In step S2, the processing of the audio data includes:
[0085] Let the audio signal in the video be A t , indicating the audio signal arriving at time point t, A t is a time-domain signal, A t ={a1,a2,…,a n}, a i ∈R, a i represents the sample point of the audio signal, and the sampling frequency is f s Hz, that is, f s samples are collected per second;
[0086] For the audio signal, preprocessing is first performed, including noise suppression and spectral enhancement; the improved Wiener filtering method is used for noise suppression. Let the original signal be x (t) , and its power spectrum is P x(f) , and the noise power spectrum is P n(f) , then the frequency response of the filter is:
[0087] H(f) = P x(f) / (P x(f) +α·P n(f) )
[0088] where α is a smoothing factor, and its value range is [0,1];
[0089] The filtered signal is denoted as x'(t), and then speech separation processing based on a deep neural network is used to extract the target speech from the mixed signal; let the mixed signal x'(t) be expressed as the superposition of multiple signals:
[0090] x'(t) = s 1(t) +s 2(t) +...+s n(t)
[0091] where s 1(t) , s 2(t) ,..., s n(t) respectively represent the signal components from different sources;
[0092] Speech separation adopts the time-frequency masking method. First, calculate the short-time spectrum X(f, t) of the signal:
[0093] X(f, t) = STFT(x'(t))
[0094] where STFT represents the short-time Fourier transform;
[0095] Then estimate the time-frequency masking matrix M of each sound source i(f,t) , and the target speech signal s 1(t) is obtained in the following way:
[0096] S 1(f,t) = M 1(f,t) · X(f, t)
[0097] s 1(t) = ISTFT(S 1(f,t) )
[0098] The important features of the extracted speech signal in the time domain include the zero-crossing rate, which is used to distinguish between voiced and unvoiced sounds. For the sampling point sequence {a i}}, the zero-crossing rate calculation formula is:
[0099]
[0100] where I is the indicator function, and sgn(a i ) represents the sign function of the signal sample a i :
[0101]
[0102] To obtain the time-frequency characteristics of the signal, perform a short-time Fourier transform on the processed speech signal to obtain the time-varying spectrum:
[0103]
[0104] where x(n) is the audio signal, and w(t - n) is the Hamming window function, expressed as:
[0105] w(n) = 0.54 - 0.46cos(2πn / N), 0 ≤ n ≤ N - 1
[0106] where f is the frequency, t is the time, and N is the window length; the frequency distribution characteristics of the signal at different time points are obtained through STFT for subsequent feature extraction and analysis.
[0107] In step S2, the processing of the subtitle text data includes:
[0108] The text is segmented and normalized, and specific word segmentation strategies are adopted for different languages. For Chinese text, a word segmentation model based on conditional random fields is adopted. This model uses the feature function f i (y t-1 ,y t ,x,t) describes the contextual dependency of word boundaries, where y represents the tag sequence, x represents the input text sequence, and t represents the position index; the probability distribution of the model is expressed as:
[0109]
[0110] Among them, λ i is the feature weight, Z(x) is the normalization factor;
[0111] For English text, we use a combination of rule-based and statistical methods to segment words, including space segmentation, punctuation recognition, and special phrase processing, followed by stop word filtering to remove function words, conjunctions, and auxiliary words that have no actual semantic contribution, while retaining key words with emotional color and topic indication. For English text, we also perform case unification processing, converting all letters to lowercase, and then use the neural network-based Word2Vec model to map words into vector representations. The Word2Vec model uses the Skip-gram architecture to maximize the target word w t Predict its context word w t+j The logarithmic probability of learning word vectors:
[0112]
[0113] Where c is the context window size, θ is the model parameter, and p(w t+j |w t ) is calculated by the softmax function:
[0114]
[0115] Among them, v w and v' w are the input and output vector representations of word w, respectively. Finally, each word w is mapped to a d-dimensional vector space to obtain a dense word vector v w ∈R d , d ranges from 100 to 300;
[0116] Let the text sequence T t =(w1,w2,...,w L ), L is the length of the text, w i is the i-th word; text T t The feature representation of adopts weighted average based on self-attention mechanism:
[0117]
[0118] where α i is the attention weight of the i-th word, calculated as follows:
[0119]
[0120] where W a ∈ R d×d , b a ∈ R d and v a ∈ R d are learnable parameters of the attention mechanism;
[0121] The sentiment analysis of the text is based on a deep model that combines a bidirectional long short-term memory network and an attention mechanism. The input to the model is a sequence of word vectors (v w1 , v w2 ,..., v wL ). Context-related feature representations h i are extracted through Bi-LSTM:
[0122]
[0123] The sentiment classifier calculates the probability distribution of the sentiment label y_emotion based on the feature vector h:
[0124] P(y emotion |T t ) = softmax(W e h + b e )
[0125] where W e and b e are parameters of the sentiment classifier, and y emotion ∈ {positive, negative, neutral};
[0126] Meanwhile, a hierarchical attention network is used for topic classification, considering semantic information at both the word and sentence levels, and outputting the topic category y topic and its probability distribution, where y topic ∈ {politics, entertainment, sports, technology,...}.
[0127] S3. Define the publicity strategy and classify it; the defined publicity strategies include:
[0128] Emotional manipulation and emotional polarization strategies, false information and misleading information dissemination strategies, image and sound manipulation strategies, information overload and repeated reinforcement strategies, social proof and group effect strategies, visual hint and symbol transmission strategies, and clickbait strategies;
[0129] Emotional manipulation and emotional polarization strategies, whose manifestations include using fear to trigger behavioral responses and using compassion to guide behavior;
[0130] False information and misleading information dissemination strategies, whose manifestations include taking out of context and fabricating facts;
[0131] Image and sound manipulation strategies, whose manifestations include video frame manipulation and audio manipulation;
[0132] Information overload and repeated reinforcement strategies, whose manifestations include quickly switching scenes and information and repeatedly reinforcing key information at a high frequency;
[0133] Social proof and group effect strategies, whose manifestations include group effects and social proof, and using celebrity effects;
[0134] Visual hint and symbol transmission strategies, whose manifestations include using specific symbols or icons and color psychology;
[0135] Clickbait strategies, whose manifestations include extreme titles and emotional narratives.
[0136] S4. Conduct targeted analysis according to the characteristic requirements of different publicity strategies to identify event information dissemination strategies; specifically:
[0137] Extract features and preprocess the input short video to establish a feature library containing video images, audio, subtitle text, and user interaction data; the video feature library includes information such as facial expression sequences, hue changes, and the positions of social iconic elements, the audio feature library contains speech signal features and audio emotion features, the text feature library stores word segmentation results, emotional tendencies, and keyword information, and the interaction feature library includes user comment emotions and dissemination path data;
[0138] After completing feature extraction, conduct targeted analysis according to the characteristic requirements of different publicity strategies, as follows:
[0139] For emotional manipulation and emotional polarization strategies, use a multi-modal emotion analysis framework for identification. First, extract facial expression features in the key frames of the video through a deep neural network model to establish an expression-emotion mapping matrix E = [e ij , where e ij represents the association strength between the i-th expression category and the j-th emotional state; at the same time, apply spectral analysis and prosody feature extraction to the audio signal to construct a speech emotion feature vector S = (s pitch , s energy,s tempo ,s timbre ), where each component represents pitch, energy, rhythm, and timbre features respectively; calculate the emotional polarization index:
[0140]
[0141] where, e i is the emotional value of the i-th frame, is the average emotional value, σ e is the emotional standard deviation, w i is the weight coefficient; when the PI value exceeds the threshold θ PI , it is determined that there is an emotional polarization strategy; calculate the emotional change rate ΔE / Δt, and when its value increases abnormally, it is determined that there may be emotional incitement behavior;
[0142] For the spread strategies of false information and misleading information, a cross-source information consistency verification method is adopted. First, extract the key fact point set F = {f1, f2,..., f m} from the video subtitles and the results of speech-to-text conversion through named entity recognition technology. Each fact point f i contains a triple of subject, predicate, and object; then access the trusted knowledge base K and the fact-checking API to calculate the information consistency score:
[0143]
[0144] where the verify function returns the consistency score of the fact point f i and the knowledge base K; construct the event representation vector v event through the deep contrast learning method, and calculate the semantic distance matrix D = [d event 1 , v event 2 ,..., v event k} with the set of representation vectors of videos of the same type of event {v ij}, where d ij = cos(v event i , v event j ); when there are abnormally low values in D and CS is lower than the threshold θ CS , it is determined that there is false or misleading information;
[0145] For the image and sound manipulation strategies, a temporal multi-modal consistency detection algorithm is adopted. This algorithm constructs the visual emotion sequence V = (v1, v2,..., v n ) and the audio emotion sequence A = (a1, a2,..., a n), calculate the sequence matching degree DTW(V,A) through the dynamic time warping algorithm, and evaluate the intrinsic correlation degree of the two modalities based on the mutual information MI(V;A); when DTW(V,A)>θ DTW and MI(V;A)<θ MI it is identified as a manipulation mode with inconsistent images and sounds; at the same time, analyze the time lag of modal correlation, and identify the deliberate emotional mismatch phenomenon by calculating the lag cross-correlation coefficient ρ lag(V,A)
[0146] For information overload and repeated reinforcement strategies, analyze through the information entropy change rate and scene jump detection algorithm. First, calculate the term frequency-inverse document frequency value of the key information in the video, and construct the term frequency distribution P=(p1,p2,...,p k ), and calculate the information entropy
[0147]
[0148] When the repetition index R = 1 - H(P) / log(k) exceeds the threshold θ R it is determined that there is a repeated reinforcement strategy; for scene change detection, calculate the structural similarity index SSIM(I t ,I t+1 ), construct the scene change sequence SC=(sc1,sc2,...,sc n-1 ), and count the frequency of its change rate exceeding the threshold θ SC When the high-frequency scene switching and key information repetition show significant correlation in time, it is determined that the information overload strategy is adopted;
[0149] For social proof and group effect strategies, based on the social network propagation analysis model, first construct the user interaction graph G=(V,E), where V represents the set of interacting users and E represents the interaction relationship; calculate the user group cohesion index
[0150]
[0151] and the influence centrality vector BC to identify the key opinion leader nodes; analyze the emotional consistency of the review text and calculate σ sentiment 2 / μ sentiment as the emotional polarization index; evaluate the plasticity of group opinions through the social resilience analysis method. When the communication network shows high cohesion, low heterogeneity and the emotional polarization index exceeds the threshold θ SP it is determined that there is a social proof strategy;
[0152] For visual hint and symbol transmission strategies, a symbol semantic analysis framework is used for identification. First, a set of visual symbols S = {s1, s2,..., s p} in the video is extracted through a pre-trained deep convolutional neural network, and the detected symbols are mapped to the corresponding semantic space using the symbol-semantic mapping matrix M = [m ij ; the temporal distribution T = {t1, t2,..., t p} and the screen position distribution P = {p1, p2,..., p p} of the symbols are analyzed to evaluate the significance and concealment of symbol presentation; combined with HSV color space analysis, the distribution characteristics of colors in the emotional dimension are calculated. When a significant correlation is detected between a specific symbol, a specific color, and a specific emotion, and this correlation exceeds the random expectation θ SC , it is determined that there is a visual hint strategy;
[0153] For the clickbait strategy, a title-content consistency evaluation model is adopted. First, the semantic vector v title of the title and the semantic vector v content of the video content are extracted through natural language processing techniques, and the cosine similarity cos(v title , v content ) between them is calculated; the density ρ emotion of emotional words and the usage frequency f extreme of extreme words in the title are analyzed to construct the clickbait index
[0154] TBI = α(1 - cos(v title , v content )) + β·ρ emotion + γ·f extreme
[0155] where α, β, and γ are weight coefficients; when TBI exceeds the threshold θ TBI , it is determined that the clickbait strategy is adopted; at the same time, the consistency between the emotional excitation intensity of the video content and the title promise is evaluated to identify the content presentation patterns of deliberate exaggeration and violation of expectations.
[0156] In another embodiment, a multi-modal event information dissemination strategy recognition system for short video platforms is also provided. The system adopts the multi-modal event information dissemination strategy recognition method described in the above embodiment. The system includes a data collection module, a data processing module, a publicity strategy module, and a dissemination strategy recognition module;
[0157] The data collection module is used to obtain data from short video platforms, including video image data, audio data, subtitle text data, and user interaction data;
[0158] A data processing module for processing the collected video image data, audio data, and subtitle text data, and establishing a feature library including video image data, audio data, subtitle text data, and user interaction data;
[0159] A publicity strategy module for defining and classifying publicity strategies, including emotional manipulation and polarization strategies, false information and misleading information dissemination strategies, image and sound manipulation strategies, information overload and repeated reinforcement strategies, social proof and group effect strategies, visual hint and symbol transmission strategies, and clickbait strategies;
[0160] A communication strategy recognition module for conducting targeted analysis according to the feature requirements of different publicity strategies to achieve the recognition of event information communication strategies.
[0161] It should also be noted that in this specification, terms such as "including", "comprising", or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or elements inherent to such a process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including the said element.
[0162] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for identifying multi-modal event information dissemination strategies for short video platforms, characterized in that, It includes the following steps: S1. Data collection, obtaining data from short video platforms, including video image data, audio data, subtitle text data, and user interaction data; S2. Data processing, processing the collected video image data, audio data, and subtitle text data, and establishing a feature library containing video image data, audio data, subtitle text data, and user interaction data; S3. Defining and classifying publicity strategies; S4. Conducting directional analysis according to the feature requirements of different publicity strategies to identify event information dissemination strategies.
2. The multimodal event information dissemination strategy recognition method for short video platforms according to claim 1, wherein Step S1 specifically includes: Obtaining video image data, specifically: Extract key frame images of the video, retaining important scenes, objects, and colors; assume the video stream is represented as S = {V1, V2, …, V n}, V n represents the video frame arriving at time n; Initialize the parameters of the video data processing method, including the video frame set F, the video frame feature library F feat , the image data I of each video frame wt and the emotion label E of each frame wt ; The video frame set F is used to store all video frame data extracted from the short video; Treat each video frame V in the video stream t as image data I t , input the set F of video frames and use the video frame V t to update the image feature library F feat , where F feat = {I t , C t , E t}, I t is the video frame image data, containing RGB color values, representing the visual information in the video, C t is the hue information, representing the color distribution or hue of the video frame, extracted through a color histogram or average hue, E t is the emotion label, representing the facial expressions and emotional tendencies in the video frame; Obtaining audio data, specifically including extracting the audio part from the video, performing sentiment recognition and speech-to-text processing through audio analysis, and extracting the emotional color, speech rate, and pitch change of the speech; Obtaining subtitle text data, specifically including the title, description, subtitles, and comments of the video, and performing text analysis through natural language processing technology to identify the emotional tendency and topic focus of the video; Obtaining user interaction data, specifically collecting the behavior data of the audience, including but not limited to the number of likes, comments, and forwards.
3. The method for identifying a multi-modal event information dissemination strategy for a short video platform according to claim 1, wherein In step S2, the processing of video image data includes facial expression detection, hue and color analysis, and social iconic element detection; Among them, facial expression detection includes face detection and expression classification; Face detection is used to identify the position of the face from the video frame, specifically: Input video frame image I t , I t is represented in the form of an image matrix, I t ∈R H×W×3 , where H is the height of the image, W is the width of the image, and 3 represents the three RGB color channels; Each pixel I of the image t (i, j) consists of intensity values of three colors: red, green, and blue, that is, I t (i, j) = {I t (r) (i, j), I t (g) (i, j), I t (b) (i, j)}, where I t (r) 、I t (g) 、I t (b) represent the pixel values of the red, green, and blue channels in the image at position (i, j), respectively; Converting the color image to a grayscale image, and the specific calculation formula is as follows: I t (gray) (i,j) = 0.2989×I t (r) (i,j) + 0.587×I t (g) (i,j) + 0.114×I t (b) (i,j) Among them, I t (gray) (i, j) represents the pixel value at the (i, j) position in the grayscale image; Normalize to scale the grayscale value between 0 and 1, I t (norm) (i,j) = I t (gray) (i,j) / 255; Using MTCNN for face region detection. MTCNN is a cascaded multi-task convolutional neural network for face detection and key point localization, consisting of three cascaded sub-networks, namely P-Net, R-Net, and O-Net; among them, P-Net is the first-stage network, adopting a shallow CNN structure of 12×12×3, generating candidate face frames by sliding window scanning on the image pyramid, and the output includes face classification scores, bounding box regression vectors, and facial key point coordinates, and using non-maximum suppression to remove overlapping candidate boxes; R-Net is the second-stage network, adopting a CNN structure of 24×24×3, performing more detailed feature extraction and classification on the candidate boxes output by P-Net, outputting more accurate face classification scores and bounding box positions, and further removing false detection boxes through more stringent threshold screening; O-Net is the last-stage network, adopting a deep CNN structure of 48×48×3, not only performing face detection, but also locating five facial key points, namely the eyes, nose, and the two corners of the mouth, outputting the final face detection box, confidence score, and facial key point coordinates, and finally obtaining the final detection result using a more stringent non-maximum suppression threshold; When performing face detection, first construct an image pyramid \(S = \{I_1, I_2, \ldots, I\) n \}, where \(I\) i represents the \(i\)-th layer image, and the scaling ratio between adjacent layers is scale_factor; set a sliding window of size \(W\) win \times H win for each layer image. The window starts from the top-left corner \((0, 0)\) of the image and slides within the image range \([0, W\) i - W win \times[0, H\) i - H win with step sizes stride_x and stride_y, where \(W\) i and \(H\) i are the width and height of the \(i\)-th layer image respectively; Sub-image region I extracted during each sliding sub(x,y) ∈R Wwin×Hwin×3 is input to the P-Net for processing. The P-Net extracts features from the input through a three-layer cascaded convolutional structure. In the three-layer cascaded convolutional structure, the first convolutional layer uses a convolutional kernel W1 of size 3×3×3×10 to process the input to obtain a feature map F1 ∈ R10×10×10. After passing through PReLU activation and max pooling, it is input to the second layer. The second layer uses a convolutional kernel W2 of size 3×3×10×16 to obtain a feature map F2 ∈ R4×4×16. The third layer uses a convolutional kernel W3 of 3×3×16×32 to finally obtain a feature map F3 ∈ R2×2×32, that is, F sub(x,y) ; The feature maps obtained by feature extraction are then input into two parallel fully connected layers, which are respectively used for face classification to obtain a probability p(x,y)∈[0,1] and bounding box regression to obtain an offset Δbox(x,y)=(Δx1,Δy1,Δx2,Δy2); For positions where the probability exceeds the threshold θ p , calculate the candidate box box(x, y) = T(x, y, W win ) + α·Δbox(x, y) according to the predicted offset, where T is the position mapping transformation function and α is the regression coefficient. Thus, obtain the candidate box set B = {b1, b2,..., b m} and the corresponding confidence set P score = {p1, p2,..., p m}; To remove overlapping detection boxes, the non-maximum suppression method is adopted. First, calculate the intersection over union of any two boxes B1 and B2: When its value exceeds the preset threshold Threshold, retain the bounding boxes with higher confidence; The candidate boxes that have completed the preliminary screening are scaled to a size of 24×24 and then input into the R-Net to extract the feature F through a deeper convolutional network r ∈R 128 and a more stringent threshold θr is used for secondary screening and position refinement; Finally, the candidate bounding boxes are scaled to 48×48 and then input into the O-Net. The O-Net outputs the final detection bounding box bboxO-Net, the positions of 5 key points landmarks, and the confidence p through five convolutional layers and three parallel fully connected layers. Each final face detection result is represented as Bi=(x1,y1,x2,y2,p,landmarks), which includes the coordinates of the upper left and lower right corners of the detection bounding box, the confidence, and the set of facial key point coordinates; Facial expression classification is used to analyze the detected facial region and identify the corresponding expression based on facial features, including the following steps: Input image I of the facial region face , use an improved cascade regression face calibrator based on OpenCV to extract facial key points. This calibrator adopts a multi-level cascade regression structure, and each level contains two sub-modules: local feature extraction and shape regression. First, normalize the input image to the standard size N×N, where N = 96, and then extract the coordinate of the feature points of the key facial regions including eyes, eyebrows, lips, and nose. For each key point k i =(x i ,y i ), obtain the feature points through the following iterative optimization formula: Among them, T i (x i , y i ) is a predefined facial region template, using a shape-constrained local response map as the template feature, and k i are the coordinates of the facial key points; in each iteration, local features are extracted based on the currently estimated key point positions and the position estimates are updated until convergence or the maximum number of iterations is reached; Finally, the extracted set of facial key point coordinates {k1, k2,..., k n} is obtained; Generate a feature vector of the expression by calculating local features within each facial region; the calculated feature vector is f face =[f1,f2,…,f m , where f1,f2,…,f m are the respective features extracted from the facial regions; Train a convolutional neural network CNN to automatically learn features from facial images for facial expression classification, and finally output a probability distribution representing the probability of each expression: Among them, W i is the weight corresponding to each expression category, and P(emotion i ∣I face ) represents the probability that the image I face belongs to expression i; Output the predicted expression category y through the classification model emotion and the predicted confidence P(emotion i ) for each expression category, where y emotion ∈{smile,angry,sad,…}, and P(emotion i ) represents the recognition probability of this expression.
4. The method for identifying a multi-modal event information dissemination strategy for a short video platform according to claim 3, wherein The hue and color analysis is specifically as follows: Convert the RGB image to the HSV color space, including: Obtain the maximum value C in RGB max = max(R, G, B) and the minimum value C min = min(R, G, B), C max and C min are the brightest and darkest color values in the three RGB channels respectively; Calculate the hue H, and the formula is as follows: If the calculated H value is less than 0, add 360 to it to ensure that the hue is in the range of [0, 360); Calculate the saturation S, and the formula is as follows: If C max equals 0 then S = 0; Calculate the luminance V, where V = C max ; Finally, implement the conversion of the RGB image to the HSV color space; Based on the image in the converted HSV color space, perform hue analysis to determine whether the image has a certain publicity strategy according to the hue distribution; saturation analysis, by analyzing the saturation changes in the image, identify whether there is a transfer of certain emotions or strategies through color harmony; brightness analysis, brightness affects the visual perception of the image, and judge whether the image wants to express a specific atmosphere through brightness analysis.
5. The method for identifying a multi-modal event information dissemination strategy for a short video platform according to claim 3, wherein The detection of social iconic elements is specifically as follows: By detecting specific social iconic elements in video images, identify the emotions, values, or cultural symbols related to the publicity strategy; Train and use a convolutional neural network CNN for the detection of social iconic elements; collect image data with iconic elements and annotate each image to mark the regions of social iconic elements in it. These data will be used to train the convolutional neural network CNN; automatically learn the features in the images through the CNN model and extract the high-level features related to the iconic elements; Use the trained CNN to predict the input image I t Perform prediction, output the category and location of the landmark elements. If a specific logo or symbol is recognized, the model outputs the following information: The bounding box (x min , y min , x max , y max ) of each detected landmark element represents the position of the landmark element in the image; the class label C label corresponding to each detected landmark element; for each recognized element, the confidence P(C label |I t ) of the output model represents the probability that the element belongs to a certain class.
6. The method for identifying a multi-modal event information dissemination strategy for a short video platform according to claim 1, wherein In step S2, the processing of audio data includes: Let the audio signal in the video be A t , representing the audio signal arriving at time point t, A t is a time-domain signal, A t ={a1,a2,…,a n}, a i ∈R, a i represents the sample point of the audio signal, and the sampling frequency is f s Hz, that is, f s samples are collected per second; For the audio signal, preprocessing is first performed, including noise suppression and spectral enhancement; for noise suppression, an improved Wiener filtering method is used. Let the original signal be x (t) , whose power spectrum is P x(f) , and the noise power spectrum is P n(f) . Then the frequency response of the filter is: H(f) = P x(f) / (P x(f) + α·P n(f) ) Among them, α is a smoothing factor, and its value range is [0, 1]; The filtered signal is denoted as x'(t), and then a speech separation process based on a deep neural network is used to extract the target speech from the mixed signal. Let the mixed signal x'(t) be expressed as the superposition of multiple signals: x'(t) = s 1(t) + s 2(t) +... + s n(t) Among them, s 1(t) , s 2(t) ,..., s n(t) respectively represent signal components from different sources; Speech separation uses the time-frequency masking method. First, calculate the short-time spectrum X(f,t) of the signal: X(f,t)=STFT(x'(t)) Among them, STFT represents the short-time Fourier transform; Then, estimate the time-frequency masking matrix M for each sound source i(f,t) , the target speech signal s 1(t) is obtained in the following way: S 1(f,t) = M 1(f,t) · X(f, t) s 1(t) = ISTFT(S 1(f,t) ) Important features of the extracted speech signal in the time domain include the zero-crossing rate, which is used to distinguish between voiced and unvoiced sounds. For the sampling point sequence {a i}, the zero-crossing rate calculation formula is as follows: where I is the indicator function, sgn(a i ) represents the sign function of the signal sample a i : To obtain the time-frequency characteristics of the signal, perform a short-time Fourier transform on the processed speech signal to obtain the time-varying spectrum: Among them, x(n) is the audio signal, and w(t-n) is the Hamming window function, which is expressed as: w(n)=0.54 - 0.46cos(2πn / N), 0≤n≤N - 1 Among them, f is the frequency, t is the time, and N is the window length; the frequency distribution characteristics of the signal at different time points are obtained through STFT, which are used for subsequent feature extraction and analysis.
7. The multimodal event information dissemination strategy recognition method for a short video platform according to claim 1, wherein In step S2, the processing of the subtitle text data includes: The text is segmented and normalized, and specific word segmentation strategies are adopted for different languages. For Chinese text, a word segmentation model based on conditional random fields is adopted. This model uses the feature function f i (y t-1 ,y t ,x,t) describes the contextual dependency of word boundaries, where y represents the tag sequence, x represents the input text sequence, and t represents the position index; the probability distribution of the model is expressed as: where λ i is the feature weight and Z(x) is the normalization factor; For English texts, a rule-based and statistics-based method is used for word segmentation, including space splitting, punctuation recognition, and special phrase processing. Subsequently, stop word filtering is performed to remove function words, conjunctions, and auxiliary words that have no actual semantic contribution, while retaining keywords with emotional colors and topic-indicating functions. For English texts, case normalization is also carried out, converting all letters to lowercase. Subsequently, the Word2Vec model based on neural networks is used to map words into vector representations. The Word2Vec model adopts the Skip-gram architecture and learns word vectors by maximizing the logarithmic probability of predicting its context word w t to predict its context word w t+j : Among them, c is the context window size, θ is the model parameter, and p(w t+j |w t ) is calculated by the softmax function: where, v w and v' w are the input and output vector representations of word w, respectively. Finally, each word w is mapped to a d-dimensional vector space to obtain a dense word vector v w ∈R d , where d ranges from 100 to 300; Let the text sequence be T t =(w1, w2,..., w L ), where L is the text length and w i is the i-th word; the feature representation of text T t adopts weighted average based on the self-attention mechanism: where α i is the attention weight of the i-th word and is calculated as follows: Among them, W a ∈R d×d , b a ∈R d and v a ∈R d are learnable parameters of the attention mechanism; The sentiment analysis of the text is based on a deep model that combines a bidirectional long short-term memory network and an attention mechanism. The input to the model is a sequence of word vectors (v w1 , v w2 ,..., v wL ). The context-related feature representations h i are extracted through Bi-LSTM: The sentiment classifier calculates the probability distribution of the sentiment label y_emotion based on the feature vector h: P(y emotion |T t ) = softmax(W e h + b e ) where, W e and b e are the parameters of the sentiment classifier, and y emotion ∈{positive, negative, neutral}; Meanwhile, a hierarchical attention network is used for topic classification, considering semantic information at both the word and sentence levels, and the topic category y is output topic and its probability distribution, where y topic ∈{politics,entertainment,sports,technology,...}.
8. The method for identifying a multi-modal event information dissemination strategy for a short video platform according to claim 1, wherein The defined publicity strategies include: Sentiment manipulation and emotional polarization strategies, false information and misleading information dissemination strategies, image and sound manipulation strategies, information overload and repeated reinforcement strategies, social proof and group effect strategies, visual hint and symbol transmission strategies, and clickbait strategies; Sentiment manipulation and emotional polarization strategies, the manifestation forms include using fear to trigger behavioral responses and using compassion to guide behaviors; False information and misleading information dissemination strategies, the manifestation forms include taking out of context and fabricating facts; Image and sound manipulation strategies, the manifestation forms include picture manipulation and audio manipulation; Information overload and repeated reinforcement strategies, the manifestation forms include quickly switching scenes and information and repeatedly emphasizing key information at a high frequency; Social proof and group effect strategies, the manifestation forms include group effect and social proof, using celebrity effect; Visual hint and symbol transmission strategies, the manifestation forms include using specific symbols or icons and color psychology; Clickbait strategies, the manifestation forms include extreme titles and emotional narratives.
9. The method for identifying a multi-modal event information dissemination strategy for a short video platform according to claim 8, characterized in that, The identification of the event information dissemination strategy is specifically: The method of step S1 and S2 is adopted for the input short video for feature extraction and processing, and a feature library including video images, audio, subtitle text, and user interaction data is established; the video feature library includes facial expression sequences, hue changes, and the position information of social iconic elements, the audio feature library contains speech signal features and audio emotion features, the text feature library stores the word segmentation results, sentiment tendencies, and keyword information, and the interaction feature library includes user comment emotions and dissemination path data; After the feature extraction is completed, targeted analysis is carried out according to the feature requirements of different publicity strategies, specifically as follows: For emotional manipulation and emotional polarization strategies, a multi-modal sentiment analysis framework is used for identification. First, facial expression features in the key frames of the video are extracted through a deep neural network model, and an expression-emotion mapping matrix E = [e ij is established, where e ij represents the association strength between the i-th expression category and the j-th emotional state. At the same time, spectral analysis and prosodic feature extraction are applied to the audio signal to construct a speech emotion feature vector S = (s pitch , s energy , s tempo , s timbre ), where each component represents pitch, energy, rhythm, and timbre features respectively; Calculate the emotional polarization index: Among them, e i is the emotion value of the i-th frame, ee is the average emotion value, and σ e is the emotion standard deviation, and w i is the weight coefficient; when the PI value exceeds the threshold θ PI it is determined that there is an emotion polarization strategy; calculate the emotion change rate ΔE / Δt, and when its value increases abnormally, it is determined that there may be an emotion incitement behavior; For the dissemination strategies of false and misleading information, a cross-source information consistency verification method is adopted. First, a set of key fact points F = {f1, f2,..., f m} is extracted from the video subtitles and the results of speech-to-text conversion through named entity recognition technology. Each fact point f i contains a subject, a predicate, and an object triple; subsequently, a trusted knowledge base K and a fact-checking API are accessed to calculate the information consistency score: Among them, the verify function returns the fact point f i The consistency score with the knowledge base K; construct the event representation vector v through the deep contrast learning method event , and with the set of representation vectors of videos of similar events {v event 1 , v event 2 ,..., v event k} calculate the semantic distance matrix D = [d ij , d ij = cos(v event i , v event j ); when there are abnormally low values in D and CS is lower than the threshold θ CS , it is determined that there is false or misleading information; For the image and sound control strategy, a temporal multi-modal consistency detection algorithm is adopted. This algorithm constructs a visual emotion sequence V = (v 1, v 2, ..., v n ) and an audio emotion sequence A = (a1, a2,..., a n ). The sequence matching degree DTW(V, A) is calculated through the dynamic time warping algorithm, and the internal correlation degree between the two modalities is evaluated based on the mutual information content MI(V; A). When DTW(V, A) > θ DTW and MI(V; A) < θ MI , it is recognized that there is an inconsistent control mode between the image and the sound. At the same time, the time lag of the modal correlation is analyzed, and the deliberate emotion mismatch phenomenon is identified by calculating the lag cross-correlation coefficient ρ lag(V,A) . For the information overload and repeated reinforcement strategy, it is analyzed through the information entropy change rate and the scene jump detection algorithm. First, the term frequency-inverse document frequency value of the key information in the video is calculated, and the term frequency distribution P = (p1, p2,..., p k ) is constructed, and the information entropy is calculated When the repeatability index R = 1 - H(P) / log(k) exceeds the threshold θ R it is determined that there is a repeated reinforcement strategy; for scene change detection, calculate the structural similarity index SSIM(I t ,I t+1 ) between adjacent frames, construct the scene change sequence SC = (sc1, sc2,..., sc n-1 ), and count the frequency of its change rate exceeding the threshold θ SC ; when there is a significant temporal correlation between high-frequency scene switching and key information repetition, it is determined that an information overload strategy is adopted; For the social proof and group effect strategies, based on the social network dissemination analysis model, first construct a user interaction graph G=(V, E), where V represents the set of interacting users and E represents the interaction relationship; calculate the user group cohesion index And the influence centrality vector BC to identify the key opinion leader nodes; analyze the sentiment consistency of the review text and calculate σ sentiment 2 / μ sentiment As the sentiment polarization index; evaluate the plasticity of group opinions through the social resilience analysis method. When the communication network shows high cohesion, low heterogeneity, and the sentiment polarization index exceeds the threshold θ SP it is determined that there is a social proof strategy; For the visual hint and symbol transmission strategy, a symbolic semantic analysis framework is adopted for recognition. First, the set of visual symbols S = {s1, s2,..., s p} in the video is extracted through a pre-trained deep convolutional neural network, and the detected symbols are mapped to the corresponding semantic space using the symbol-semantic mapping matrix M = [m ij ; the temporal distribution T = {t1, t2,..., t p} and the screen position distribution P = {p1, p2,..., p p} of the symbols are analyzed to evaluate the significance and concealment of the symbol presentation; combined with the HSV color space analysis, the distribution characteristics of colors in the emotional dimension are calculated. When a significant correlation exists between a specific symbol, a specific color, and a specific emotion, and this correlation exceeds the random expectation θ SC , it is determined that there is a visual hint strategy; For the clickbait strategy, a title-content consistency evaluation model is adopted. First, the semantic vector v of the title is extracted through natural language processing technology title and the semantic vector v of the video content content . The cosine similarity cos(v title , v content ) between them is calculated; the density ρ emotion of emotional words and the usage frequency f extreme of extreme words in the title are analyzed, and a clickbait index is constructed TBI = α(1 - cos(v title , v content )) + β·ρ emotion + γ·f extreme where α, β, and γ are weighting coefficients; when the TBI exceeds the threshold θ TBI it is determined that the clickbait strategy is adopted; at the same time, the consistency between the emotional arousal intensity of the video content and the title promise is evaluated, and the content presentation patterns of deliberate exaggeration and violation of expectations are identified.
10. A multi-modal event information dissemination strategy recognition system for short video platforms, characterized in that, The system adopts the method described in any one of claims 1-9, and the system includes a data acquisition module, a data processing module, a publicity strategy module, and a dissemination strategy identification module; The data acquisition module is used to obtain data from the short video platform, including video image data, audio data, subtitle text data, and user interaction data; The data processing module is used to process the acquired video image data, audio data, and subtitle text data, and establish a feature library including video image data, audio data, subtitle text data, and user interaction data; The publicity strategy module is used to define publicity strategies and classify them, including sentiment manipulation and emotional polarization strategies, false information and misleading information dissemination strategies, image and sound manipulation strategies, information overload and repeated reinforcement strategies, social proof and group effect strategies, visual hint and symbol transmission strategies, and clickbait strategies; The dissemination strategy recognition module conducts targeted analysis according to the characteristic requirements of different publicity strategies to achieve the recognition of event information dissemination strategies.
Citation Information
Cited By
Marketing video auditing method based on AI
CN120583273A