Misleading Short Video Detection Method and Device
Through integrated field analysis and multimodal features, an external knowledge base is constructed, and multi-head gated vectors and multi-layer perceptron classifiers are used to solve the timeliness of identifying misleading short videos in the existing methods, achieving higher accuracy detection.
Patent Information
- Application Number
- CN202510455059.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-11
AI Technical Summary
Existing misleading short video detection methods rely on content or modal correlations, making it difficult to effectively identify carefully edited fake news videos, especially misleading content for past events.
Through integrated domain analysis, multimodal integration and cross-platform fact verification, the text, audio, visual and domain characteristics of short videos are obtained, external knowledge bases are built, and multi-head gated vectors and multi-layer perceptron classifiers are used for detection.
It significantly improves the detection accuracy of misleading short videos, and can identify edited misleading content across platforms, events, and fields, enhancing the model's adaptability.
Smart Images

Figure CN119992426B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence and social network technology, and more specifically relates to a misleading short video detection method and device. Background Art
[0002] With the popularity of short video platforms, the amount of user-generated content has increased dramatically, and misleading content has also spread. These videos mislead viewers through editing, clickbait, false information, etc., affecting public cognition and judgment. The motivation behind this may be to obtain traffic, advertising revenue, or even political or social manipulation. The proliferation of misleading short videos has become a serious problem on social media platforms. Therefore, designing a misleading short video detection method suitable for social media to maintain the authenticity of information and mitigate negative impact is a hot issue that needs to be solved urgently.
[0003] Since short video content is complex and contains multiple forms of data, such as titles / text records, images, audio, comments, and user information, the task of detecting misleading content is extremely challenging. Some existing methods rely on multimodal feature analysis of the short video content itself. For content analysis, researchers focus on the relative relationship between different modalities, such as the complementarity between video metadata and comment credibility, and the semantic inconsistency between video images and text titles. These methods use the cross-modal correlation of multimodal features such as vision, audio, and text to identify potential misleading information.
[0004] Most existing methods rely heavily on the multimodal features of short videos, ignoring the fact that many fake news videos are carefully edited to be persuasive. In this case, focusing only on the content and the correlation between different modes may not be effective in detection, such as maliciously making past events into new fake news or generating misinformation. Some methods explore combining external knowledge with content features to further improve the detection accuracy of fake news videos. By introducing domain knowledge or knowledge bases, researchers are able to better reveal the contradictions between video content and facts, thereby identifying videos that have been carefully edited to try to mislead viewers. However, relying on external knowledge also faces challenges, such as the sparsity and unreliability of knowledge base information, especially in the context of different fields, it is still difficult to find relevant and high-quality knowledge.
[0005] In summary, existing misleading short video detection works rely only on content or make judgments based on the correlation between different modalities. However, these methods are not conducive to identifying fake news instances that have been maliciously tampered with and re-circulated on short video platforms, especially for events that occurred in the past. Summary of the invention
[0006] Embodiments of the present invention propose a misleading short video detection method and device. This method integrates domain analysis, multi-modal integration, and cross-platform fact verification, and solves the problem that existing short video detection ignores timeliness due to relying on content or based on the correlation between different modalities.
[0007] Embodiments of the present invention provide a misleading short video detection method, including:
[0008] According to the received target short video, obtain the text feature, audio feature, visual feature, metadata feature, and domain feature of the target short video;
[0009] Obtain a multi-modal integrated representation and a multi-domain integrated vector representation in sequence according to the text feature, audio feature, visual feature, metadata feature, and domain feature, and obtain a cross-domain multi-modal vector representation according to the multi-modal integrated representation and the multi-domain integrated vector representation;
[0010] Preprocess the text feature and construct an external knowledge base, obtain an external knowledge text related to the target short video through a text aggregation method based on word vectors, encode the external knowledge text, and obtain a vector representation corresponding to the external knowledge text; determine the cross-domain multi-modal vector representation corresponding to the target short video as an alternative video representation, and obtain a relationship representation vector between the external knowledge text and the alternative video according to the vector representation corresponding to the external knowledge text and the alternative video representation;
[0011] Obtain a fused vector according to an adaptive multi-head gating vector, a cross-domain multi-modal vector representation, and a relationship representation vector learned through a trainable process; input the fused vector into a multi-layer perceptron classifier, determine the prediction label output by the multi-layer perceptron classifier, and select the category with the largest prediction label as the final detection result.
[0012] Preferably, the step of obtaining a multi-modal integrated representation and a multi-domain integrated vector representation in sequence according to the text feature, audio feature, visual feature, metadata feature, and domain feature, and obtaining a cross-domain multi-modal vector representation according to the multi-modal integrated representation and the multi-domain integrated vector representation specifically includes:
[0013] Determine a domain vector according to the domain feature, determine a title vector and a text record vector according to the text feature, determine an audio vector according to the audio feature, determine a visual vector according to the visual feature, and determine a metadata vector according to the metadata feature;
[0014] The domain vector, title vector, and text record vector obtain a multi-domain integrated vector representation through a collaborative attention mechanism;
[0015] The title vector, text record vector, audio vector, visual information vector, and metadata vector are processed through a cross-modal Transformer mechanism to obtain a multi-modal integrated representation;
[0016] A cross-domain multi-modal vector representation is obtained based on the multi-domain integrated vector representation and the multi-modal integrated representation.
[0017] Preferably, the cross-domain multi-modal vector representation is expressed by the following formula:
[0018]
[0019] D c = co-attention{D r ,T 1r ,T 2r}
[0020] M c = cross-modal-transformer{T 1r ,T 2r ,A r ,C r ,M r}
[0021] where, e M represents the cross-domain multi-modal vector representation, M c represents the multi-modal integrated representation, D c represents the multi-domain integrated vector representation, D r represents the domain vector, T 1r represents the title vector, T 2r represents the text record vector, A r represents the audio vector, C r represents the visual information vector, M r represents the metadata vector, co-attention(·) represents the co-attention mechanism, cross-modal-transformer(·) represents the cross-modal attention encoder, and ⊕ represents the vector concatenation operation.
[0022] Preferably, the preprocessing based on text features and the construction of the external knowledge base specifically include:
[0023] Performing text preprocessing on the title features and text record features included in the target short video to obtain preprocessed text;
[0024] Collecting content related to the domain features included in the target short video to construct an external knowledge base, and performing retrieval based on the external knowledge base to obtain multiple external knowledge texts related to the target short video;
[0025] A text aggregation method based on word vectors, which converts a preprocessed text and multiple external knowledge texts into word vectors, and calculates the similarity between the word vectors of the preprocessed text and each external knowledge text according to the cosine similarity;
[0026] Convert multiple external knowledge texts into vector representations according to a pre-trained model, and generate multiple vector representations for each external knowledge text according to the attention mechanism.
[0027] Preferably, determining the cross-domain multi-modal vector representation corresponding to the target short video as an alternative video representation, and obtaining a relationship representation vector between the external knowledge text and the alternative video according to the vector representation and the alternative video representation, specifically including:
[0028] Determine the number of vector representations included in each external knowledge text and the number of weights included in each external knowledge text, and obtain a relationship representation vector through the following formula according to the vector representations corresponding to the multiple external knowledge texts, the weights corresponding to the multiple external knowledge texts, and the alternative video representation:
[0029] e F = Embed(F)
[0030]
[0031] where e F represents the relationship representation vector between the external knowledge text and the alternative video, Embed(·) represents the embedding operation of vector dimension conversion, F represents the comprehensive score of the external knowledge, e M represents the cross-domain multi-modal vector representation, j represents the loop index variable, γ j represents the weight corresponding to the j-th external knowledge text, K j represents the vector representation corresponding to the j-th external knowledge text, and N represents the number of external knowledge texts.
[0032] Preferably, the fusion vector is determined through the following formula:
[0033] e MF = g e ⊙ e M + (1 - g e ) ⊙ e F
[0034] where e MF represents the fusion vector, g e represents the adaptive multi-head gating vector, e M represents the cross-domain multi-modal vector representation, e F represents the relationship representation vector between the external knowledge text and the alternative video, and ⊙ represents the element-wise multiplication operation.
[0035] Preferably, the prediction label is determined by the following formula:
[0036]
[0037] After selecting the class with the largest predicted label as the final detection result, the loss value is determined according to the following formula:
[0038]
[0039] where represents the predicted label of the data, argmax(·) represents the function of selecting the class with the largest probability, Softmax(·) represents the normalized exponential function for calculating the probability of each class, MLP(·) represents the multi-layer perceptron classifier, i represents the class, y represents the true label of the data, e MF represents the fusion vector, and L represents the loss value.
[0040] The embodiment of the present invention provides a misleading short video detection device, including:
[0041] An acquisition unit, configured to obtain the text feature, audio feature, visual feature, metadata feature, and domain feature of the target short video according to the received target short video;
[0042] A first obtaining unit, configured to sequentially obtain a multi-modal integrated representation and a multi-domain integrated vector representation according to the text feature, audio feature, visual feature, metadata feature, and domain feature, and obtain a cross-domain multi-modal vector representation according to the multi-modal integrated representation and the multi-domain integrated vector representation;
[0043] A second obtaining unit, configured to preprocess the text feature and construct an external knowledge base, obtain an external knowledge text related to the target short video based on the text aggregation method of word vectors, encode the external knowledge text, and obtain a vector representation corresponding to the external knowledge text; determine the cross-domain multi-modal vector representation corresponding to the target short video as an alternative video representation, and obtain a relationship representation vector between the external knowledge text and the alternative video according to the vector representation corresponding to the external knowledge text and the alternative video representation;
[0044] A third obtaining unit, configured to obtain a fused vector according to the adaptively multi-headed gated vector, cross-domain multi-modal vector representation, and relationship representation vector learned during the trainable process; input the fused vector into a multi-layer perceptron classifier, and determine the predicted label of the output of the multi-layer perceptron classifier, and select the class with the largest predicted label as the final detection result.
[0045] An embodiment of the present invention provides a computer device, which includes a scenario database and a processor. The scenario database stores a computer program. When the computer program is executed by the processor, the processor executes the misleading short video detection method described in any one of the above.
[0046] An embodiment of the present invention provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor executes the misleading short video detection method described in any one of the above.
[0047] The present invention proposes a misleading short video detection method and device. Based on content understanding and external knowledge enhancement, by constructing a multi-domain external knowledge base, it effectively solves the problems of data sparsity and uneven distribution in different domains. At the same time, combined with multi-modal feature fusion and an adaptive gating mechanism, it significantly improves the detection accuracy and the adaptive ability of the model. Further, by introducing external knowledge and domain vectors, this method can more effectively identify misleading short videos that have been carefully edited and tampered with, providing a new solution for cross-platform, cross-event, and cross-domain misleading short video detection, and having important application value and promotion prospects. Description of the Drawings
[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0049] Figure 1 It is a schematic flowchart of the misleading short video detection method provided by the embodiment of the present invention;
[0050] Figure 2 It is a schematic structural diagram of the misleading short video detection device provided by the embodiment of the present invention. Detailed Embodiments
[0051] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0052] Figure 1 It is a schematic flowchart of the misleading short video detection method provided by the embodiment of the present invention. The following will be combined with Figure 1Taking the example of the present invention, the misleading short video detection method provided by the embodiment of the present invention is explained in detail. Figure 1 As shown, the method comprises the following steps:
[0053] Step 101, according to the received target short video, obtaining text features, audio features, visual features, metadata features and domain features of the target short video;
[0054] Step 102, obtaining a multimodal integrated representation and a multi-domain integrated vector representation in sequence according to the text features, audio features, visual features, metadata features, and domain features, and obtaining a cross-domain multimodal vector representation according to the multimodal integrated representation and the multi-domain integrated vector representation;
[0055] Step 103, preprocessing is performed according to text features and an external knowledge base is constructed, an external knowledge text related to the target short video is obtained based on a text aggregation method based on word vectors, the external knowledge text is encoded, and a vector representation corresponding to the external knowledge text is obtained; the cross-domain multimodal vector representation corresponding to the target short video is determined as an alternative video representation, and a relationship representation vector between the external knowledge text and the alternative video is obtained according to the vector representation corresponding to the external knowledge text and the alternative video representation;
[0056] Step 104, obtaining a fused vector based on the adaptive multi-head gating vector, cross-domain multimodal vector representation and relationship representation vector learned by the trainable process; inputting the fused vector into a multi-layer perceptron classifier, determining the predicted label of the output of the multi-layer perceptron classifier, and selecting the category with the largest predicted label as the final detection result.
[0057] The method provided by the embodiment of the present invention is described below in conjunction with a specific embodiment. It should be noted that the execution subject of the method is a processor, and the processor processes the received short video. Specifically, assuming that there is a data set containing multiple short videos, one of the short videos is taken as an example for description.
[0058] In step 101, the processor receives a target short video, which includes at least the following information:
[0059] Title: "Natural plant tea, high blood pressure can be reduced after just one drink!" This title uses exaggeration to attract attention.
[0060] Text record: In the video, the host vigorously promotes the special effects of plant tea, saying that as long as you keep drinking it, high blood pressure can be restored to normal without taking antihypertensive drugs.
[0061] Audio: The host’s audio commentary has an exaggerated and extremely misleading tone.
[0062] Vision: It shows the appearance of the plant tea, the scene of brewing, and some photos of the so-called users after their blood pressure has decreased.
[0063] Metadata: The video was posted on a health and wellness account recently.
[0064] Domain: It belongs to the field of health science popularization, which clarifies the theme category of the video.
[0065] In the embodiment of the present invention, after receiving the target short video, because the target short video includes the above information but is not directly displayed on the processor, the processor needs to obtain the feature information included in the target short video after receiving it. In the embodiment of the present invention, the feature information of the target short video mainly includes text features, audio features, visual features, metadata features, and domain features.
[0066] In practical applications, the processor can use a pre-trained language model (such as BERT (Bidirectional Encoder Representations from Transformers)) to encode the title and text records included in the target short video to obtain the title vector T 1r and the text record vector T 2r , and here the title vector and the text record vector are determined as text features; assume T 1r and T 2r are 768-dimensional vectors.
[0067] Use an audio feature extraction tool (such as Librosa) to extract the audio features of the target short video, such as pitch, timbre, rhythm, etc., to obtain the audio features including the audio vector A r with a dimension of 128.
[0068] Use a convolutional neural network (such as ResNet (Residual Network)) to process the video frames and image information of the target short video to obtain the visual features including the visual information vector C r with a dimension of 256.
[0069] Perform one-hot encoding on the metadata of the target short video to obtain the metadata features including the metadata vector M r with a dimension of 256.
[0070] Obtain the domain features including the domain vector D r through a pre-trained domain model, with a dimension of 768.
[0071] In step 102, the title vector and text record vector included in the text features, the audio vector included in the audio features, the visual information vector included in the visual features, and the metadata vector included in the metadata features are processed through a cross-modal Transformer mechanism to obtain a multi-modal integrated representation:
[0072] M c = cross-modal-transformer{T 1r ,T 2r ,A r ,C r ,M r}(1)
[0073] Wherein, M c represents the multi-modal integrated representation, T 1r represents the title vector, T 2r represents the text record vector, A r represents the audio vector, C r represents the visual information vector, M r represents the metadata vector, and cross-modal-transformer(·) represents the cross-modal attention encoder. Assume that after calculation, the multi-modal integrated representation M c is a 768-dimensional vector.
[0074] Furthermore, the domain vector included in the domain features, the title vector and text record vector included in the text features are processed through a co-attention mechanism to obtain a multi-domain integrated vector representation:
[0075] D c = co-attention{D r ,T 1r ,T 2r}(2)
[0076] Wherein, D c represents the multi-domain integrated vector representation, co-attention(·) represents the co-attention mechanism, D r represents the domain vector, T 1r represents the title vector, T 2r represents the text record vector. Assume that the finally obtained multi-domain integrated vector representation D c is a 768-dimensional vector.
[0077] In the embodiments of the present invention, after obtaining the multi-modal integrated representation and the multi-domain integrated vector representation, the multi-modal integrated representation and the multi-domain integrated vector representation can be fused through the following formula to obtain a cross-domain multi-modal representation vector:
[0078]
[0079] Among them, e M represents the cross-domain multimodal vector representation, M c represents the multimodal integration representation, D c represents the multi-domain integration vector representation, represents the vector concatenation operation. Based on the calculations of the above two steps, the cross-domain multimodal representation vector e M is obtained here, which is also a vector with a dimension of 768. Assuming e M = [0.15, 0.25,..., 0.768].
[0080] In step 103, a text aggregation method is designed to retrieve the external knowledge base most relevant to the target short video. Specifically, it includes multiple steps such as text preprocessing, constructing an external knowledge base, text aggregation, similarity calculation, and vector representation generation. The above multiple steps are introduced below in combination with specific embodiments:
[0081] Text preprocessing: Use natural language processing tools to preprocess the title and text records included in the target short video. The purpose is to remove punctuation marks and stop words in the title and text records, and through word segmentation technology, split the text into individual words or phrases to obtain the core text information, that is, the preprocessed text. The stop words here can be "of", "will", "can", etc.
[0082] For the title "Natural plant tea, can lower high blood pressure immediately!" and the text (in the video, the anchor strongly promoted the special effects of the plant tea, saying that as long as you keep drinking it and don't take blood pressure-lowering medicine, high blood pressure can return to normal) of the target short video in the previous embodiment, perform cleaning. Use natural language processing tools to remove stop words (such as "of", "will", "can", etc.), punctuation marks, etc. Through word segmentation technology, split the text into individual words or phrases to obtain the core text information "Natural plant tea drink cures high blood pressure".
[0083] In the embodiment of the present invention, preprocessing the title and text records included in the target short video aims to build a text structure, highlight key information, and facilitate subsequent matching and similarity calculation with external knowledge texts.
[0084] Constructing an external knowledge base: Collect content related to the domain features included in the target short video and construct an external knowledge base. For example, collect various information such as authoritative research reports, professional medical journal articles, expert opinions, and disease treatment guidelines in the medical field. Use web crawler technology to obtain relevant text data from well-known medical websites, academic databases and other sources, and perform cleaning, deduplication, and classification and sorting on the obtained data to construct a comprehensive and accurate external knowledge base in the medical field.
[0085] Further, retrieve based on an external knowledge base to obtain multiple external knowledge texts related to the target short video; in the embodiments of the present invention, the external knowledge base serves as the basis for continuously retrieving relevant external knowledge and provides a rich reference basis for judging the authenticity and scientificity of the content of the target short video.
[0086] Text aggregation: Use a pre-trained word vector model to convert the preprocessed text and external knowledge texts into word vectors. Assume the dimension of the word vector is 300. For the preprocessed text "Natural plant tea cures hypertension" of the target short video, each vocabulary is converted into a 300-dimensional vector, and then these vocabulary vectors are combined into a 300-dimensional vector representing the entire preprocessed text through methods such as average pooling; for the multiple external knowledge texts retrieved based on the external knowledge base, the same vocabulary vector conversion and combination operations are also performed. For the external knowledge text 1 "Medical research shows that hypertension is a chronic disease and currently cannot be cured by drinking tea. The main treatment methods are drug treatment and lifestyle adjustment", it is converted into a 300-dimensional vector; for the external knowledge text 2 "The therapeutic effect of a single tea-drinking therapy on hypertension is limited and cannot replace regular antihypertensive drug treatment", it is converted into a 300-dimensional vector.
[0087] Similarity calculation: Calculate the similarity between the preprocessing vector corresponding to the preprocessed text and the external vectors corresponding to each external knowledge text according to the cosine similarity. In the embodiments of the present invention, external knowledge texts with a similarity higher than the threshold can be screened according to the set similarity threshold. For example, if both external knowledge text 1 and external knowledge text 2 meet the conditions, they are aggregated as the external knowledge most relevant to the target short video.
[0088] Vector representation generation: Use a pre-trained BERT model to convert the screened external knowledge texts into vector representations, and then use the multi-head attention mechanism to generate multiple vector representations for each external knowledge text. For example, for external knowledge text 1, after being processed by the BERT model and the 3-head attention mechanism operation, 3 vector representations K 11 ,K 12 ,K 13 ; for external knowledge text 2, 3 vector representations K 21 ,K 22 ,K 23 .
[0089] Further, first determine the cross - domain multi - modal vector representation corresponding to the target short video as the alternative video representation, then successively determine the number of vector representations included in each external knowledge text and the number of weights included in each external knowledge text, and then obtain the relationship representation vector through the following formula based on the vector representations corresponding to multiple external knowledge texts, the weights corresponding to multiple external knowledge texts, and the alternative video representation:
[0090] e F = Embed(F) (4 - 1)
[0091]
[0092] For example, for external knowledge text 1, it includes 3 vectors, and each vector corresponds to a weight, which are respectively: γ 11 = 0.3, γ 12 = 0.3, γ 13 = 0.4; for external knowledge text 2, it also includes 3 vectors, and each vector corresponds to a weight, which are respectively: γ 21 = 0.2, γ 22 = 0.4, γ 23 = 0.4. According to formula (4 - 1), the relationship representation vector can be calculated. Suppose after calculation, F = [0.08, 0.12,..., 0.768].
[0093] Among them, e F represents the relationship representation vector between the external knowledge text and the alternative video, Embed(·) represents the embedding operation for vector dimension conversion, F represents the comprehensive score of external knowledge, e M represents the cross - domain multi - modal vector representation, j represents the loop index variable, γ j represents the weight corresponding to the j - th external knowledge text, K j represents the vector representation corresponding to the j - th external knowledge text, and N represents the number of external knowledge texts.
[0094] In step 104, according to the adaptive multi - head gating vector learned through the training process, the cross - domain multi - modal vector representation, and the relationship representation vector, obtain the fusion vector through the following formula:
[0095] e MF = g e ⊙ e M +(1 - g e )⊙ e F (5)
[0096] For example, the adaptive multi - head gating vector g e is a 768 - dimensional vector, the first 384 dimensions are 0.7, and the last 384 dimensions are 0.3; suppose after calculation, eF = [0.05, 0.08,..., 0.6]; According to the formula e M = M c ⊕ D c Calculate the fusion vector. Assume that after calculation, e M = [0.09, 0.14,..., 0.7]. Then according to formula (5), e MF , e MF = [0.3, 0.2,..., 0.2, 0.3].
[0097] Among them, e MF represents the fusion vector, g e represents the adaptive multi-head gating vector, which is used to control the contribution ratio of e M and e F when forming the fusion vector. ⊙ represents the element-wise multiplication operation, e M represents the cross-domain multi-modal vector representation, e F represents the relationship representation vector between the external knowledge text and the alternative video.
[0098] Furthermore, input the fusion vector into the multi-layer perceptron classifier, process the output of the multi-layer perceptron classifier through the Softmax function to obtain the probability of each category, and then take the category with the maximum probability as the final detection result. Specifically:
[0099]
[0100] For example, take the previously obtained e MF as the input vector and input it into the MLP classifier. The output vector of the MLP classifier is processed through the Softmax function. Then the following output vector can be obtained: [0.4375, 0.6625].
[0101] In this embodiment, 0.6625 is the category with the maximum probability in the output vector, and the category with the maximum probability can be taken as the final detection result.
[0102] Furthermore, after selecting the category with the maximum probability as the final detection result, determine the loss value according to the following formula:
[0103]
[0104] Among them, represents the predicted label of the data, argmax(·) represents the function of selecting the category with the maximum probability, Softmax(·) represents the normalized exponential function for calculating the probability of each category, MLP(·) represents the multi-layer perceptron classifier, i represents the category, y represents the true label of the data, e MFDenote the fusion vector as and the loss value as L.
[0105] For example, in the above embodiment, assume that the target short video is actually misleading, so y = 1. From the Softmax output vector in the above embodiment, it can be regarded as the probabilities of different categories. Assume that category 1 represents misleading and category 0 represents non-misleading, then (the probability corresponding to category 1).
[0106] According to the above formula (7), the cross-entropy loss function can be calculated. Substitute the values y = 1 and into formula (7) to get: L = -[1×log(0.6625)+(1 - 1)log(1 - 0.6625)]; finally, we get: L = -log(0.6625)≈0.411. At the same time, the accuracy of the model prediction can be evaluated through the value of the cross-entropy loss function L = 0.411. During the training process, the model parameters will be continuously adjusted to reduce the loss value.
[0107] In this embodiment, according to the category with the highest probability in the Softmax output vector, it is determined that the target short video is a misleading short video. Here, it is mainly used to output the detection result of the misleading short video and evaluate the accuracy of the model prediction. Specifically, first, the true label and the predicted label are compared, and the loss value is calculated through the cross-entropy loss function. This loss value can reflect the difference between the model prediction result and the actual situation, so as to evaluate the accuracy of the model prediction. During the training process, this loss value will be used as an important feedback information to continuously adjust the model parameters to reduce the loss value and improve the prediction accuracy of the model.
[0108] To introduce the method provided by the embodiments of the present invention more clearly, the following combines specific embodiments to introduce this method:
[0109] Embodiment 1
[0110] This time, a short video promoting that "soaking feet with a magic herbal medicine pack can cure diabetes" is to be detected for misleading. The operation is carried out according to the method for detecting misleading short videos based on content understanding and external knowledge enhancement, which specifically includes the following steps:
[0111] Step 201, input short video data:
[0112] This short video contains the following information:
[0113] Title: "Soak feet with a magic herbal medicine pack, diabetes will be cured immediately!"
[0114] Transcript: In the video, the anchor constantly emphasizes the amazing efficacy of the herbal medicine pack, saying that as long as you keep soaking your feet with it, you don't need to take medicine or get injections, and diabetes can be cured.
[0115] Audio: The audio of the host's explanation, with a positive and inciting tone.
[0116] Vision: It shows the appearance of the herbal medicine packet, the scene of foot soaking, and some photos of the so-called users after recovery.
[0117] Metadata: The video was posted on a certain health preservation account, and the posting time is recent.
[0118] Step 202: Extract short video feature information:
[0119] Use a pre-trained language model (such as BERT) to encode the title and transcript to obtain the title vector T 1r and the transcript vector T 2r . Assume T 1r and T 2r are 768-dimensional vectors.
[0120] Use an audio feature extraction tool to extract the features of the audio, such as pitch, timbre, rhythm, etc., to obtain the audio vector A r , with a dimension of 128.
[0121] Apply a convolutional neural network to process the video frames and image information to obtain the visual information vector C r , with a dimension of 256.
[0122] Perform one-hot encoding on the metadata to obtain the metadata vector M r , with a dimension of 32.
[0123] Domain vector D r Is obtained through a pre-trained domain model, with a dimension of 768.
[0124] Step 203: Fusion the domain representation and the multi-modal representation through the cross-domain multi-modal integration method to obtain the cross-domain multi-modal representation vector:
[0125] 1) Multi-modal feature integration:
[0126] According to the formula M c = cross-modal-transformer{T 1r ,T 2r ,A r ,C r ,M r}, use the cross-modal Transformer mechanism to integrate the features of different modalities. Assume that after calculation, the multi-modal integration representation M c is a 768-dimensional vector.
[0127] 2) Calculate the multi-domain integration vector representation:
[0128] According to the formula obtain using the co-attention mechanism Assume that the finally obtained multi-domain integrated vector representation D c is a 768-dimensional vector.
[0129] 3) Cross-domain multi-modal vector representation:
[0130] According to the formula fuse the multi-modal integrated representation and the multi-domain integrated vector representation to obtain the cross-domain multi-modal representation vector e M , with a dimension of 768. Assume e M = [0.1, 0.2,..., 0.768].
[0131] Step 204, design a text aggregation method to retrieve the external knowledge most relevant to the target short video:
[0132] 1) Text preprocessing: Clean the title of the target short video "Soaking feet with a magical herbal medicine pack, diabetes is cured immediately!" and the text (in the video, the host constantly emphasizes the magical effects of the herbal medicine pack, claiming that as long as you keep soaking your feet with it and don't need to take medicine or get injections, diabetes can be cured). Use natural language processing tools to remove stop words (such as "of", "just", "can", etc.), punctuation marks, etc. Through word segmentation technology, split the text into individual words or phrases to obtain the core text information "Soaking feet with a magical herbal medicine pack cures diabetes".
[0133] The purpose is to simplify the text structure, highlight the key information, and facilitate subsequent matching and calculating similarity with external knowledge texts.
[0134] 2) Build an external knowledge base: Collect various information such as authoritative research reports, professional medical journal articles, expert opinions, and disease treatment guidelines in the medical field. Use web crawler technology to obtain relevant text data from well-known medical websites, academic databases, and other sources. Clean, deduplicate, and classify and organize the obtained data to build a comprehensive and accurate external knowledge base in the medical field.
[0135] This knowledge base serves as the basis for subsequent retrieval of relevant external knowledge and provides rich reference basis for judging the authenticity and scientificity of the content of the target short video.
[0136] 3) Text aggregation:
[0137] 3-1) Word vector conversion:
[0138] Use a pre-trained word vector model (such as Word2Vec) to convert both the pre-processed target short video text and the external knowledge text into word vectors. Assume the dimension of the word vectors is 300. For the pre-processed text of the target short video "Soaking feet with a magical herbal medicine pack cures diabetes", each word is converted into a 300-dimensional vector, and then these word vectors are combined into a 300-dimensional vector representing the entire pre-processed text through methods such as average pooling, denoted as V target 。
[0139] For the external knowledge text in the external knowledge base, the same operations of word vector conversion and combination are performed. For example, for the external knowledge text 1 "Medical research shows that diabetes is a chronic metabolic disease and cannot be cured by soaking feet at present. The main treatment methods are drug treatment and diet control", it is converted into a 300-dimensional vector V text1 ; for the external knowledge text 2 "The therapeutic effect of a single foot soaking therapy on diabetes is minimal and cannot replace formal medical means", it is converted into vector V text2 。
[0140] 4) Similarity calculation:
[0141] 4-1) Use the cosine similarity formula to calculate the similarity scores between the pre-processed text vector and each external knowledge text vector. The cosine similarity formula is: where A and B are two vectors respectively, · represents the dot product of vectors, and ‖A‖‖B‖ represent the norms of vectors A and B respectively.
[0142] 4-2) Calculate the similarity between V target and V text1 :
[0143] Let V target =[a1,a2,...,a 300 , V text1 =[b1,b2,...,b 300 , the dot product of vectors
[0144] V target The norm of V text1 The norm of Then Assume the result is sim1 = 0.8.
[0145] 4-3) Similarly, calculate the similarity between V target and V text2 , and assume the calculation result is sim2 = 0.75.
[0146] 4-4) Filter out external knowledge texts with a similarity higher than the set similarity threshold (e.g., 0.7). In this example, both external knowledge text 1 and external knowledge text 2 meet the conditions, and they are aggregated as the external knowledge most relevant to the target short video.
[0147] 5) Vector representation generation:
[0148] Use the pre-trained BERT model to convert the filtered external knowledge texts (external knowledge text 1 and external knowledge text 2) into vector representations. The vectors output by the BERT model have a dimension of 768.
[0149] At the same time, use a 3-head attention mechanism to generate multiple vector representations for each external knowledge text. For external knowledge text 1, after being processed by the BERT model and operated by the 3-head attention mechanism, 3 vectors K of 768 dimensions are obtained 11 ,K 12 ,K 13 ; for external knowledge text 2, vectors K are obtained 21 ,K 22 ,K 23 .
[0150] These vector representations not only contain the semantic information of the external knowledge texts but also capture the potential relationships between the external knowledge texts and the target short video from different perspectives through the multi-head attention mechanism.
[0151] 6) Subsequent applications of similarity:
[0152] Application in multi-head knowledge aggregation: In the above "multi-head knowledge aggregation, learning the relationship representation vectors of different external knowledge and candidate videos", the similarity scores will affect the distribution of attention weights. Although the formula does not directly reflect the similarity, in the actual training process, the attention weights γ corresponding to the external knowledge with a high similarity to the target short video j may be adjusted to be larger. For example, since the similarity sim1 = 0.8 of the external knowledge text 1 to the target short video is higher than the similarity sim2 = 0.75 of the external knowledge text 2, then when setting the weights, the sum of γ corresponding to the external knowledge text 1 11 ,γ 12 ,γ 13 may be relatively larger, making the external knowledge text 1 have a greater impact on the result when calculating the relationship representation vector F.
[0153] Step 205: Multi-head knowledge aggregation, learning the relationship representation vectors of different external knowledge and candidate videos:
[0154] Candidate video representation v = e M, that is, the cross-domain multi-modal representation vector obtained previously. The weights of different external knowledge corresponding to multiple attentions are respectively:
[0155] For the 3 vectors of external knowledge text 1: γ 11 = 0.2, γ 12 = 0.3, γ 13 = 0.5; for the 3 vectors of external knowledge text 2: γ 21 = 0.3, γ 22 = 0.4, γ 23 = 0.3.
[0156] According to the formula Calculate the relationship representation vector F. Assume that after calculation here, F = [0.05, 0.1,..., 0.768].
[0157] Step 206: Obtain the fusion vector using the gating mechanism:
[0158] The cross-domain multi-modal vector representation g e is a 768-dimensional vector, the first 384 dimensions are 0.6, and the last 384 dimensions are 0.4. Let e F = Embed(F). Assume that after calculation, e F = [0.03, 0.06,..., 0.5], e M = [0.1, 0.2,..., 0.768], according to the formula e MF = g e ⊙ e M + (1 - g e ) ⊙ e F The fusion vector can be determined.
[0159] Step 207: Input this vector into the MLP classifier:
[0160] Input vector: Take the fusion vector e MF as the input vector and input it into the MLP classifier. Here, it is simplified to e MF = [0.2, 0.3,..., 0.1, 0.4].
[0161] Input the fusion vector into the multi-layer perceptron classifier, and process the output of the multi-layer perceptron classifier through the Softmax function to obtain the probability of each category as [0.4375, 0.6625].
[0162] Find the category index with the largest probability through the argmax function. In this example, 0.6625 is the largest probability value, and its index is 1. So corresponds to category 1, that is, it is determined that this short video is a misleading short video.
[0163] Step 208: Output the detection result of misleading short videos:
[0164] Assume that the short video is actually misleading, so the true label y = 1; according to the cross-entropy loss function Calculate the loss. Substitute y = 1 and into the formula, we get: L = -[1×log(0)+(1 - 1)log(1 - 0)], L = -log(0).
[0165] Since log(0) is mathematically undefined, but in actual calculations, when is very close to 0, it will be a very large positive number. In deep learning, to avoid this situation, usually a very small positive number ∈ (∈ = 1e - 7) is added to then L = -[1×log(0 + ∈)+(1 - 1)log(1 - (0 + ∈))], and then we get L = -log(∈) ≈ 16.118.
[0166] The present invention proposes a method and device for detecting misleading short videos. The method is based on content understanding and external knowledge enhancement. By constructing an external knowledge base in multiple domains, it effectively solves the problems of data sparsity and uneven distribution in different domains. At the same time, combined with multi-modal feature fusion and an adaptive gating mechanism, it significantly improves the detection accuracy and the adaptive ability of the model. Further, by introducing external knowledge and domain vectors, the method can more effectively identify carefully edited and tampered misleading short videos, providing a new solution for cross-platform, cross-event, and cross-domain misleading short video detection, and having important application value and promotion prospects.
[0167] Based on the same inventive concept, the embodiments of the present invention provide a device for constructing a compliance index system based on a large language model. Since the principle of the device for solving technical problems is similar to that of the method for constructing a compliance index system based on a large language model, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be elaborated.
[0168] As Figure 2 shown, the device includes an acquisition unit 201, a first obtaining unit 202, a second obtaining unit 203, and a third obtaining unit 204.
[0169] The acquisition unit 201 is configured to obtain the text feature, audio feature, visual feature, metadata feature, and domain feature of the target short video according to the received target short video;
[0170] The first obtaining unit 202 is configured to sequentially obtain a multimodal integrated representation and a multi-domain integrated vector representation according to the text feature, audio feature, visual feature, metadata feature, and domain feature, and obtain a cross-domain multimodal vector representation according to the multimodal integrated representation and the multi-domain integrated vector representation;
[0171] The second obtaining unit 203 is configured to perform preprocessing on the text feature and construct an external knowledge base, obtain an external knowledge text related to the target short video by a text aggregation method based on word vectors, encode the external knowledge text, and obtain a vector representation corresponding to the external knowledge text; determine the cross-domain multimodal vector representation corresponding to the target short video as an alternative video representation, and obtain a relationship representation vector between the external knowledge text and the alternative video according to the vector representation corresponding to the external knowledge text and the alternative video representation;
[0172] The third obtaining unit 204 is configured to obtain a fused vector according to an adaptive multi-head gating vector, a cross-domain multimodal vector representation, and a relationship representation vector learned through a trainable process; input the fused vector into a multi-layer perceptron classifier, and determine a predicted label of an output of the multi-layer perceptron classifier, and select a category with the largest predicted label as a final detection result.
[0173] It should be understood that the units included in the above misleading short video detection device are only logical divisions according to the functions implemented by the device. In actual applications, the above units can be superimposed or split. And the functions implemented by the misleading short video detection device provided in this embodiment correspond one by one to the misleading short video detection method provided in the above embodiment. For a more detailed processing flow implemented by the device, it has been described in detail in the first method embodiment above, and will not be described in detail here.
[0174] Another embodiment of the present invention further provides a computer device, which includes: a processor and a scenario database; the scenario database is used to store computer program codes, and the computer program codes include computer instructions; when the processor executes the computer instructions, the electronic device executes each step of the misleading short video detection method in the method flow shown in the above method embodiment.
[0175] Another embodiment of the present invention further provides a computer-readable storage medium, in which computer instructions are stored. When the computer instructions run on a computer device, the computer device is caused to execute each step of the misleading short video detection method in the method flow shown in the above method embodiment.
[0176] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made to these embodiments by those skilled in the art once they learn of the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention.
[0177] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A method for detecting misleading short videos, characterized in that, Including: Obtain the text feature, audio feature, visual feature, metadata feature, and domain feature of the target short video according to the received target short video; Obtain a multi-modal integrated representation and a multi-domain integrated vector representation in sequence according to the text feature, audio feature, visual feature, metadata feature, and domain feature, and obtain a cross-domain multi-modal vector representation according to the multi-modal integrated representation and the multi-domain integrated vector representation; Preprocess the text feature and construct an external knowledge base, obtain an external knowledge text related to the target short video by a text aggregation method based on word vectors, encode the external knowledge text, and obtain a vector representation corresponding to the external knowledge text; determine the cross-domain multi-modal vector representation corresponding to the target short video as an alternative video representation, and obtain a relationship representation vector between the external knowledge text and the alternative video according to the vector representation corresponding to the external knowledge text and the alternative video representation; Obtain a fused vector according to the adaptive multi-head gated vector, cross-domain multi-modal vector representation, and relationship representation vector learned through a trainable process; input the fused vector into a multi-layer perceptron classifier, determine the prediction label output by the multi-layer perceptron classifier, and select the category with the largest prediction label as the final detection result; Among them, the step of obtaining a multi-modal integrated representation and a multi-domain integrated vector representation in sequence according to the text feature, audio feature, visual feature, metadata feature, and domain feature, and obtaining a cross-domain multi-modal vector representation according to the multi-modal integrated representation and the multi-domain integrated vector representation specifically includes: Determine a domain vector according to the domain feature, determine a title vector and a text record vector according to the text feature, determine an audio vector according to the audio feature, determine a visual vector according to the visual feature, and determine a metadata vector according to the metadata feature; The domain vector, title vector, and text record vector obtain a multi-domain integrated vector representation through a collaborative attention mechanism; The title vector, text record vector, audio vector, visual information vector, and metadata vector obtain a multi-modal integrated representation through a cross-modal Transformer mechanism; Obtain a cross-domain multi-modal vector representation according to the multi-domain integrated vector representation and the multi-modal integrated representation.
2. The method according to claim 1, wherein The cross-domain multi-modal vector representation is represented by the following formula: D c = co-attention{D r ,T 1r ,T 2r} M c = cross-modal-transformer{T 1r ,T 2r ,A r ,C r ,M r} Among them, e M represents the cross-domain multi-modal vector representation, M c represents the multi-modal integrated representation, D c represents the multi-domain integrated vector representation, D r represents the domain vector, T 1r represents the title vector, T 2r represents the text record vector, A r represents the audio vector, C r represents the visual information vector, M r represents the metadata vector, co-attention(·) represents the co-attention mechanism, cross-modal-transformer(·) represents the cross-modal attention encoder, represents the vector concatenation operation.
3. The method according to claim 1, wherein The step of preprocessing the text feature and constructing an external knowledge base specifically includes: Perform text preprocessing on the title feature and text record feature included in the target short video to obtain a preprocessed text; Collect content related to the domain feature included in the target short video to construct an external knowledge base, perform retrieval based on the external knowledge base, and obtain multiple external knowledge texts related to the target short video; Based on a text aggregation method based on word vectors, convert the preprocessed text and multiple external knowledge texts into word vectors, and calculate the similarity between the word vectors of the preprocessed text and each external knowledge text according to the cosine similarity; Convert multiple external knowledge texts into vector representations according to a pre-trained model, and generate multiple vector representations for each external knowledge text according to the attention mechanism.
4. The method according to claim 1, wherein Determining the cross-domain multi-modal vector representation corresponding to the target short video as an alternative video representation, and obtaining a relationship representation vector between the external knowledge text and the alternative video according to the vector representation corresponding to the external knowledge text and the alternative video representation, specifically including: Determining the number of vector representations included in each external knowledge text and the number of weights included in each external knowledge text, and obtaining a relationship representation vector through the following formula according to the vector representations corresponding to multiple external knowledge texts, the weights corresponding to multiple external knowledge texts, and the alternative video representation: e F = Embed(F) Among them, e F represents the relationship representation vector of the external knowledge text and the alternative video, Embed(·) represents the embedding operation for vector dimension transformation, F represents the comprehensive score of the external knowledge, and e M represents the cross-domain multi-modal vector representation, j represents the loop index variable, and γ j represents the weight corresponding to the j-th external knowledge text, and K j represents the vector representation corresponding to the j-th external knowledge text, and N represents the number of external knowledge texts.
5. The method according to claim 1, wherein The fusion vector is determined through the following formula: e MF = g e ⊙ e M + (1 - g e ) ⊙ e F Among them, e MF represents the fusion vector, g e represents the adaptive multi-head gating vector, e M represents the cross-domain multi-modal vector representation, e F represents the relationship representation vector between the external knowledge text and the alternative video, and ⊙ represents the element-wise multiplication operation.
6. The method according to claim 1, wherein The predicted label is determined through the following formula: After selecting the class with the largest predicted label as the final detection result, determining the loss value according to the following formula: Among them, represents the predicted label of the data, argmax(·) represents the function of selecting the category with the highest probability, Softmax(·) represents the normalized exponential function for calculating the probability of each category, MLP(·) represents the multi-layer perceptron classifier, i represents the category, y represents the true label of the data, and e MF represents the fusion vector, and L represents the loss value.
7. Misleading short video detection device, characterized in that, Including: An acquisition unit, configured to obtain the text feature, audio feature, visual feature, metadata feature, and domain feature of the target short video according to the received target short video; A first obtaining unit, configured to sequentially obtain a multi-modal integrated representation and a multi-domain integrated vector representation according to the text feature, audio feature, visual feature, metadata feature, and domain feature, and obtain a cross-domain multi-modal vector representation according to the multi-modal integrated representation and the multi-domain integrated vector representation; A second obtaining unit, configured to preprocess the text feature and construct an external knowledge base, obtain an external knowledge text related to the target short video through a text aggregation method based on word vectors, encode the external knowledge text to obtain a vector representation corresponding to the external knowledge text; determine the cross-domain multi-modal vector representation corresponding to the target short video as an alternative video representation, and obtain a relationship representation vector between the external knowledge text and the alternative video according to the vector representation corresponding to the external knowledge text and the alternative video representation; A third obtaining unit, configured to obtain a fused vector according to the adaptive multi-head gated vector, cross-domain multi-modal vector representation, and relationship representation vector learned during the trainable process; input the fused vector into a multi-layer perceptron classifier, and determine the predicted label of the output of the multi-layer perceptron classifier, and select the class with the largest predicted label as the final detection result; The first obtaining unit is further configured to: determine a domain vector according to the domain feature, determine a title vector and a text record vector according to the text feature, determine an audio vector according to the audio feature, determine a visual vector according to the visual feature, and determine a metadata vector according to the metadata feature; The domain vector, title vector, and text record vector obtain a multi-domain integrated vector representation through a collaborative attention mechanism; The title vector, text record vector, audio vector, visual information vector, and metadata vector obtain a multi-modal integrated representation through a cross-modal Transformer mechanism; Obtaining a cross-domain multi-modal vector representation according to the multi-domain integrated vector representation and the multi-modal integrated representation.
8. A computer device, characterized in that, The computer device includes a scenario database and a processor. The scenario database stores a computer program. When the computer program is executed by the processor, the processor is caused to execute the misleading short video detection method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, There is a computer program stored. When the computer program is executed by a processor, the processor is caused to execute the misleading short video detection method according to any one of claims 1-6.
Citation Information
Patent Citations
Multi-modal event knowledge graph construction method
CN114064918A
Short video release information detection method, system and device and medium
CN119031185A