Misleading short video detection method and device
Through integrated field analysis, multimodal integration and cross-platform fact verification methods, an external knowledge base is built and feature fusion using cross-modal Transformer and collaborative attention mechanisms solves the problem that existing technology is difficult to identify carefully edited misleading short videos, and achieves more efficient detection accuracy and adaptability.
Patent Information
- Application Number
- CN202510455059.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-11
AI Technical Summary
Existing misleading short video detection methods rely on content or modal relevance, making it difficult to identify carefully edited and tampered fake news videos, especially misleading videos for past events.
Using integrated domain analysis, multimodal integration and cross-platform fact verification methods, an external knowledge base is built by acquiring multiple features of short videos (text, audio, visual, metadata and domain features), using cross-modal Transformer and collaborative attention mechanisms to feature fusion, and detection through multi-layer perceptron classifiers.
It significantly improves the accuracy of misleading short video detection and the adaptability of the model, and can more effectively identify carefully edited misleading videos, suitable for cross-platform, cross-event and cross-domain detection.
Smart Images

Figure CN119992426A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence and social network technology, and more specifically relates to a misleading short video detection method and device. Background Art
[0002] With the popularity of short video platforms, the amount of user-generated content has increased dramatically, and misleading content has also spread. These videos mislead viewers through editing, clickbait, false information, etc., affecting public cognition and judgment. The motivation behind this may be to obtain traffic and advertising revenue. Therefore, designing a misleading short video detection method suitable for social media to maintain the authenticity of information and reduce negative impact is a hot issue that needs to be solved urgently.
[0003] Since short video content is complex and contains multiple forms of data, such as titles / text records, images, audio, comments, and user information, the task of detecting misleading content is extremely challenging. Some existing methods rely on multimodal feature analysis of the short video content itself. For content analysis, researchers focus on the relative relationship between different modalities, such as the complementarity between video metadata and comment credibility, and the semantic inconsistency between video images and text titles. These methods use the cross-modal correlation of multimodal features such as vision, audio, and text to identify potential misleading information.
[0004] Most existing methods rely heavily on the multimodal features of short videos, ignoring the fact that many fake news videos are carefully edited to be persuasive. In this case, focusing only on the content and the correlation between different modes may not be effective in detection, such as maliciously making past events into new fake news or generating misinformation. Some methods explore combining external knowledge with content features to further improve the detection accuracy of fake news videos. By introducing domain knowledge or knowledge bases, researchers are able to better reveal the contradictions between video content and facts, thereby identifying videos that have been carefully edited to try to mislead viewers. However, relying on external knowledge also faces challenges, such as the sparsity and unreliability of knowledge base information, especially in the context of different fields, it is still difficult to find relevant and high-quality knowledge.
[0005] In summary, existing misleading short video detection works rely only on content or make judgments based on the correlation between different modalities. However, these methods are not conducive to identifying fake news instances that have been maliciously tampered with and re-circulated on short video platforms, especially for events that occurred in the past. Summary of the invention
[0006] The embodiments of the present invention propose a method and device for detecting misleading short videos. The method integrates domain analysis, multimodal integration and cross-platform fact verification, and solves the problem that existing short video detection ignores timeliness because it relies on content or the correlation between different modalities.
[0007] The embodiment of the present invention provides a misleading short video detection method, including:
[0008] According to the received target short video, obtain text features, audio features, visual features, metadata features and domain features of the target short video;
[0009] A multimodal integrated representation and a multi-domain integrated vector representation are obtained in sequence according to the text features, audio features, visual features, metadata features and domain features, and a cross-domain multimodal vector representation is obtained according to the multimodal integrated representation and the multi-domain integrated vector representation;
[0010] Preprocessing is performed according to text features and an external knowledge base is constructed. An external knowledge text related to the target short video is obtained based on a text aggregation method based on word vectors. The external knowledge text is encoded to obtain a vector representation corresponding to the external knowledge text; the cross-domain multimodal vector representation corresponding to the target short video is determined as an alternative video representation, and a relationship representation vector between the external knowledge text and the alternative video is obtained based on the vector representation corresponding to the external knowledge text and the alternative video representation;
[0011] A fused vector is obtained according to the adaptive multi-head gating vector, the cross-domain multimodal vector representation and the relationship representation vector that can be learned according to the training process; the fused vector is input into the multi-layer perceptron classifier, and the predicted label output by the multi-layer perceptron classifier is determined, and the category with the largest predicted label is selected as the final detection result.
[0012] Preferably, the step of sequentially obtaining a multimodal integrated representation and a multi-domain integrated vector representation according to the text features, audio features, visual features, metadata features and domain features, and obtaining a cross-domain multimodal vector representation according to the multimodal integrated representation and the multi-domain integrated vector representation specifically includes:
[0013] Determine a domain vector according to the domain features, determine a title vector and a text record vector according to the text features, determine an audio vector according to the audio features, determine a visual vector according to the visual features, and determine a metadata vector according to the metadata features;
[0014] The domain vector, title vector and text record vector are represented by a multi-domain integrated vector through a collaborative attention mechanism;
[0015] The title vector, text record vector, audio vector, visual information vector and metadata vector are expressed in a multimodal integrated manner through a cross-modal Transformer mechanism;
[0016] A cross-domain multimodal vector representation is obtained according to the multi-domain integrated vector representation and the multimodal integrated representation.
[0017] Preferably, the cross-domain multimodal vector representation is expressed by the following formula:
[0018]
[0019]
[0020]
[0021] in, represents a cross-domain multimodal vector representation, represents multimodal integrated representation, represents the multi-domain integrated vector representation, represents the domain vector, represents the title vector, represents a text record vector, represents the audio vector, represents the visual information vector, represents a metadata vector, represents the collaborative attention mechanism, represents the cross-modal attention encoder, Represents a vector concatenation operation.
[0022] Preferably, the preprocessing according to text features and constructing an external knowledge base specifically includes:
[0023] Performing text preprocessing on the title features and text record features included in the target short video to obtain a preprocessed text;
[0024] Collecting content related to the domain features included in the target short video to build an external knowledge base, and performing a search based on the external knowledge base to obtain a plurality of external knowledge texts related to the target short video;
[0025] The text aggregation method based on word vector converts the preprocessed text and multiple external knowledge texts into word vectors, and calculates the similarity between the word vectors of the preprocessed text and each external knowledge text based on the cosine similarity;
[0026] The plurality of external knowledge texts are converted into vector representations according to the pre-trained model, and a plurality of vector representations are generated for each of the external knowledge texts according to the attention mechanism.
[0027] Preferably, determining the cross-domain multimodal vector representation corresponding to the target short video as the candidate video representation, and obtaining the relationship representation vector between the external knowledge text and the candidate video according to the vector representation and the candidate video representation, specifically includes:
[0028] Determine the number of vector representations included in each of the external knowledge texts and the number of weights included in each of the external knowledge texts, and obtain the relationship representation vector according to the vector representations corresponding to the plurality of external knowledge texts, the weights corresponding to the plurality of external knowledge texts, and the candidate video representations through the following formula:
[0029]
[0030]
[0031] in, A vector representation representing the relationship between the external knowledge text and the candidate video, represents the embedding operation of vector dimension transformation, represents the comprehensive score of external knowledge, Indicates alternative video representation, Represents the loop index variable, Indicates The weight corresponding to the external knowledge text, Indicates The vector representation corresponding to the external knowledge text, Indicates the number of external knowledge texts.
[0032] Preferably, the fusion vector is determined by the following formula:
[0033]
[0034] in, represents the fusion vector, represents the adaptive multi-head gating vector, represents a cross-domain multimodal vector representation, A vector representation representing the relationship between the external knowledge text and the candidate video, Represents an element-wise multiplication operation.
[0035] Preferably, the predicted label is determined by the following formula:
[0036]
[0037] After selecting the category with the largest predicted label as the final detection result, the loss value is determined according to the following formula:
[0038]
[0039] in, represents the predicted label of the data, Represents the function of selecting the category with the highest probability, represents the normalized exponential function that calculates the probability of each category, represents a multi-layer perceptron classifier, Indicates category, represents the true label of the data, represents the fusion vector, Represents the loss value.
[0040] The embodiment of the present invention provides a misleading short video detection device, including:
[0041] An acquisition unit, used for acquiring text features, audio features, visual features, metadata features and domain features of the target short video according to the received target short video;
[0042] A first obtaining unit is used to obtain a multimodal integrated representation and a multi-domain integrated vector representation in sequence according to the text features, audio features, visual features, metadata features and domain features, and obtain a cross-domain multimodal vector representation according to the multimodal integrated representation and the multi-domain integrated vector representation;
[0043] The second obtaining unit is used to perform preprocessing according to text features and construct an external knowledge base, obtain external knowledge text related to the target short video based on a text aggregation method based on word vectors, encode the external knowledge text, and obtain a vector representation corresponding to the external knowledge text; determine the cross-domain multimodal vector representation corresponding to the target short video as an alternative video representation, and obtain a relationship representation vector between the external knowledge text and the alternative video according to the vector representation corresponding to the external knowledge text and the alternative video representation;
[0044] The third obtaining unit is used to obtain a fused vector according to the adaptive multi-head gating vector, the cross-domain multimodal vector representation and the relationship representation vector that can be learned according to the training process; the fused vector is input into the multi-layer perceptron classifier, and the predicted label of the output of the multi-layer perceptron classifier is determined, and the category with the largest predicted label is selected as the final detection result.
[0045] An embodiment of the present invention provides a computer device, which includes a scene database and a processor. The scene database stores a computer program. When the computer program is executed by the processor, the processor executes any one of the above-mentioned misleading short video detection methods.
[0046] An embodiment of the present invention provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor executes any one of the above-mentioned misleading short video detection methods.
[0047] The present invention proposes a misleading short video detection method and device. The method is based on content understanding and external knowledge enhancement. By building a multi-domain external knowledge base, it effectively solves the problem of data sparsity and uneven distribution in different fields. At the same time, it combines multimodal feature fusion and adaptive gating mechanism to significantly improve the detection accuracy and the adaptive ability of the model. Furthermore, by introducing external knowledge and domain vectors, the method can more effectively identify misleading short videos that have been carefully edited and tampered with, providing a new solution for the detection of misleading short videos across platforms, events, and fields, and has important application value and promotion prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0049] Figure 1 A schematic diagram of a process flow of a misleading short video detection method provided by an embodiment of the present invention;
[0050] Figure 2 A schematic diagram of the structure of a misleading short video detection device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0051] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0052] Figure 1 A schematic diagram of a misleading short video detection method provided by an embodiment of the present invention is shown below. Figure 1 Taking the example of the present invention, the misleading short video detection method provided by the embodiment of the present invention is explained in detail. Figure 1 As shown, the method comprises the following steps:
[0053] Step 101, according to the received target short video, obtaining text features, audio features, visual features, metadata features and domain features of the target short video;
[0054] Step 102, obtaining a multimodal integrated representation and a multi-domain integrated vector representation in sequence according to the text features, audio features, visual features, metadata features, and domain features, and obtaining a cross-domain multimodal vector representation according to the multimodal integrated representation and the multi-domain integrated vector representation;
[0055] Step 103, preprocessing is performed according to text features and an external knowledge base is constructed, an external knowledge text related to the target short video is obtained based on a text aggregation method based on word vectors, the external knowledge text is encoded, and a vector representation corresponding to the external knowledge text is obtained; the cross-domain multimodal vector representation corresponding to the target short video is determined as an alternative video representation, and a relationship representation vector between the external knowledge text and the alternative video is obtained according to the vector representation corresponding to the external knowledge text and the alternative video representation;
[0056] Step 104, obtaining a fused vector according to the adaptive multi-head gating vector, the cross-domain multimodal vector representation and the relationship representation vector that can be learned according to the training process; inputting the fused vector into the multi-layer perceptron classifier, determining the predicted label of the output of the multi-layer perceptron classifier, and selecting the category with the largest predicted label as the final detection result.
[0057] The method provided by the embodiment of the present invention is described below in conjunction with a specific embodiment. It should be noted that the execution subject of the method is a processor, and the processor processes the received short video. Specifically, assuming that there is a data set containing multiple short videos, one of the short videos is taken as an example for description.
[0058] In step 101, the processor receives a target short video, which includes at least the following information:
[0059] Title: "Natural plant tea, high blood pressure can be reduced after just one drink!" This title uses exaggeration to attract attention.
[0060] Text record: In the video, the host vigorously promotes the special effects of plant tea, saying that as long as you keep drinking it, high blood pressure can be restored to normal without taking antihypertensive drugs.
[0061] Audio: The host’s audio commentary has an exaggerated and extremely misleading tone.
[0062] Visual: Shows what the herbal tea looks like, how it is brewed, and some photos of alleged users experiencing lowered blood pressure.
[0063] Metadata: The video was posted on a health and wellness account recently.
[0064] Field: Belongs to the field of health science popularization, which clarifies the subject scope of the video.
[0065] In an embodiment of the present invention, after receiving the target short video, because the target short video includes the above information, but it is not directly displayed on the processor, therefore, after receiving the target short video, the processor needs to obtain the feature information included in the target short video. In an embodiment of the present invention, the feature information of the target short video mainly includes text features, audio features, visual features, metadata features and domain features.
[0066] In practical applications, the processor can use a pre-processed language model (such as BERT (Bidirectional Encoder Representations from Transformers)) to encode the title and text records included in the target short video to obtain a title vector and text record vector , where the title vector and the transcript vector are identified as text features; assuming and is a 768-dimensional vector.
[0067] Use audio feature extraction tools (such as Librosa) to extract audio features of the target short video, such as pitch, timbre, rhythm, etc., to obtain audio vectors The audio features of , the dimension is 128.
[0068] Use convolutional neural networks (such as ResNet (Residual Network)) to process the video frames and image information of the target short video to obtain a visual information vector The visual features of , the dimension is 256.
[0069] Perform one-hot encoding on the metadata of the target short video to obtain a metadata vector The metadata feature of , with a dimension of 256.
[0070] The domain vector is obtained through the pre-trained domain model The domain features of , the dimension is 768.
[0071] In step 102, the title vector and text record vector included in the text features, the audio vector included in the audio features, the visual information vector included in the visual features, and the metadata vector included in the metadata features are obtained through the cross-modal Transformer mechanism to obtain a multimodal integrated representation:
[0072] (1)
[0073] in, represents multimodal integrated representation, represents the title vector, represents a text record vector, represents the audio vector, represents the visual information vector, represents a metadata vector, represents the cross-modal attention encoder. Assume that after calculation, the multimodal integration representation is a 768-dimensional vector.
[0074] Furthermore, the domain vectors included in the domain features, the title vectors and the text record vectors included in the text features are represented by multi-domain integrated vectors through a collaborative attention mechanism:
[0075] (2)
[0076] in, represents the multi-domain integrated vector representation, represents the collaborative attention mechanism, represents the domain vector, represents the title vector, represents the text record vector. Assume that the final multi-domain integrated vector representation is a 768-dimensional vector.
[0077] In the embodiment of the present invention, after the multimodal integrated representation and the multi-domain integrated vector representation are obtained, the multimodal integrated representation and the multi-domain integrated vector representation can be fused by the following formula to obtain a cross-domain multimodal representation vector:
[0078] (3)
[0079] in, represents a cross-domain multimodal vector representation, represents multimodal integrated representation, represents the multi-domain integrated vector representation, Representation vector concatenation operation. Based on the calculations in the above two steps, the cross-domain multimodal representation vector is obtained here , which is also a vector with a dimension of 768. Assume .
[0080] In step 103, a text aggregation method is designed to retrieve the external knowledge base that is most relevant to the target short video. Specifically, it includes multiple steps of text preprocessing, building an external knowledge base, text aggregation, similarity calculation, and vector representation generation. The above steps are introduced below in conjunction with specific embodiments:
[0081] Text preprocessing: Use natural language processing tools to preprocess the title and transcript included in the target short video. The purpose is to remove punctuation marks and stop words in the title and transcript, and through word segmentation technology, split the text into individual words or phrases to obtain the core text information, that is, the preprocessed text. The stop words here can be "of", "will", "can", etc.
[0082] Clean the title "Natural plant tea, reduces high blood pressure immediately after drinking!" and the text (in the video, the host strongly promoted the special effects of the plant tea, saying that as long as you keep drinking it and don't need to take blood pressure-lowering medicine, high blood pressure can return to normal) of the target short video in the previous embodiment. Use natural language processing tools to remove stop words (such as "of", "will", "can", etc.), punctuation marks, etc. Through word segmentation technology, split the text into individual words or phrases to obtain the core text information "Natural plant tea drink cure high blood pressure".
[0083] In the embodiment of the present invention, preprocess the title and transcript included in the target short video, the purpose is to build the text structure, highlight the key information, and facilitate subsequent matching with external knowledge texts and calculating the similarity.
[0084] Build an external knowledge base: Collect content related to the domain features included in the target short video and build an external knowledge base. For example, collect authoritative research reports, professional medical journal articles, expert opinions, disease treatment guidelines, etc. in the medical field. Use web crawler technology to obtain relevant text data from well-known medical websites, academic databases and other sources, clean, deduplicate and classify the obtained data to build a comprehensive and accurate external knowledge base in the medical field.
[0085] Furthermore, perform a search based on the external knowledge base to obtain multiple external knowledge texts related to the target short video; in the embodiment of the present invention, the external knowledge base serves as the basis for subsequent retrieval of relevant external knowledge, providing a rich reference basis for judging the authenticity and scientificity of the content of the target short video.
[0086] Text aggregation: Use the pre-trained word vector model to convert the preprocessed text and external knowledge text into word vectors. Assume that the dimension of the word vector is 300 dimensions. For the preprocessed text of the target short video "Natural plant tea cures high blood pressure", each word is converted into a 300-dimensional vector, and then these word vectors are merged into a 300-dimensional vector representing the entire preprocessed text through methods such as average pooling; for multiple external knowledge texts retrieved based on external knowledge bases, the same word vector conversion and merging operations are performed. For external knowledge text 1 "Medical research shows that hypertension is a chronic disease that cannot be cured by drinking tea. The main treatment methods are drug therapy and lifestyle adjustments", it is converted into a 300-dimensional vector; for external knowledge text 2 "The single tea therapy has limited therapeutic effect on hypertension and cannot replace regular antihypertensive drug treatment", it is converted into a 300-dimensional vector.
[0087] Similarity calculation: The similarity between the preprocessed vector corresponding to the preprocessed text and the external vector corresponding to each external knowledge text is calculated based on the cosine similarity. In an embodiment of the present invention, external knowledge texts with a similarity higher than the threshold can be screened out based on a set similarity threshold. For example, if both external knowledge text 1 and external knowledge text 2 meet the conditions, they are aggregated as the external knowledge most relevant to the target short video.
[0088] Vector representation generation: Use the pre-trained BERT model to convert the selected external knowledge texts into vector representations, and then use the multi-head attention mechanism to generate multiple vector representations for each external knowledge text. For example, for external knowledge text 1, after being processed by the BERT model and the 3-head attention mechanism, 3 vector representations are obtained. ; For external knowledge text 2, 3 vector representations can also be obtained .
[0089] Furthermore, the cross-domain multimodal vector representation corresponding to the target short video is first determined as the candidate video representation, and then the number of vector representations included in each external knowledge text and the number of weights included in each external knowledge text are determined in turn, and then the relationship representation vector is obtained according to the vector representations corresponding to the multiple external knowledge texts, the weights corresponding to the multiple external knowledge texts and the candidate video representation through the following formula:
[0090] (4-1)
[0091] (4-2)
[0092] For example, for external knowledge text 1, it includes 3 vectors, each of which corresponds to a weight, namely: ; For external knowledge text 2, it also includes 3 vectors, each of which corresponds to a weight, namely: According to formula (4-1), the relationship representation vector can be calculated. Assuming that .
[0093] in, A vector representation representing the relationship between the external knowledge text and the candidate video, represents the embedding operation of vector dimension transformation, represents the comprehensive score of external knowledge, Indicates alternative video representation, Represents the loop index variable, Indicates The weight corresponding to the external knowledge text, Indicates The vector representation corresponding to the external knowledge text, Indicates the number of external knowledge texts.
[0094] In step 104, according to the adaptive multi-head gating vector, the cross-domain multimodal vector representation and the relationship representation vector that can be learned according to the training process, the fusion vector is obtained by the following formula:
[0095] (5)
[0096] For example, adaptive multi-head gate vector is a 768-dimensional vector, with the first 384 dimensions being 0.7 and the last 384 dimensions being 0.3; assuming that after calculation According to the formula Calculate the fusion vector, assuming that after calculation According to formula (5), we can get , .
[0097] in, represents the fusion vector, Represents the adaptive multi-head gating vector, used to control and The contribution ratio in forming the fusion vector, represents an element-wise multiplication operation, represents a cross-domain multimodal vector representation, A relationship vector representation is provided between the external knowledge text and the candidate videos.
[0098] Furthermore, the fusion vector is input into the multi-layer perceptron classifier. The function processes the output of the multi-layer perceptron classifier to obtain the probability of each category, and then takes the category with the highest probability as the final detection result. Specifically:
[0099] (6)
[0100] For example, the previously obtained As the input vector, it is input into the MLP classifier. The output vector of the MLP classifier is Function processing. Then we can get the following output vector: .
[0101] In this embodiment, To output the category with the highest probability in the vector, the category with the highest probability can be used as the final detection result.
[0102] Furthermore, after selecting the category with the highest probability as the final detection result, the loss value is determined according to the following formula:
[0103] (7)
[0104] in, represents the predicted label of the data, Represents the function of selecting the category with the highest probability, represents the normalized exponential function that calculates the probability of each category, represents a multi-layer perceptron classifier, Indicates category, represents the true label of the data, represents the fusion vector, Represents the loss value.
[0105] For example, in the above embodiment, it is assumed that the target short video is actually misleading, so there is From the above embodiments The output vector can be viewed as the probability of different categories. Assuming that category 1 represents misleading and category 0 represents non-misleading, then (corresponding to the probability of class 1).
[0106] According to the above formula (7), the cross entropy loss function can be calculated. Substitute the value into formula (7) and , we can get: ; Finally get: At the same time, the value of the cross entropy loss function The accuracy of the model's predictions can be evaluated, and during the training process the model parameters are continuously adjusted to reduce the loss value.
[0107] In this embodiment, according to The category with the highest probability in the output vector is judged as a misleading short video. This is mainly used to output the detection results of misleading short videos and evaluate the accuracy of model predictions. Specifically, the real label and the predicted label are first compared, and the loss value is calculated through the cross entropy loss function. The loss value can reflect the difference between the model prediction result and the actual situation, thereby evaluating the accuracy of the model prediction. During the training process, this loss value will serve as an important feedback information to continuously adjust the parameters of the model to reduce the loss value and improve the prediction accuracy of the model.
[0108] In order to more clearly describe the method provided by the embodiment of the present invention, the method is described below in conjunction with a specific embodiment:
[0109] Embodiment 1
[0110] This time, we will conduct misleading detection on a short video that promotes "Magic herbal bag foot bath can cure diabetes". We will follow the misleading short video detection method based on content understanding and external knowledge enhancement, which includes the following steps:
[0111] Step 201, input short video data:
[0112] This short video contains the following information:
[0113] Title: "Bath with the magic herbal bag, diabetes is cured in one bath!"
[0114] Transcript: In the video, the host repeatedly emphasizes the miraculous effects of the herbal bag, saying that as long as you keep using it to soak your feet, you can cure diabetes without taking any medicine or injections.
[0115] Audio: The host’s explanation audio, with an affirmative and provocative tone.
[0116] Visual: Shows the appearance of herbal bags, foot bathing scenes, and some photos of alleged users after recovery.
[0117] Metadata: The video was posted on a health and wellness account recently.
[0118] Step 202: Extract short video feature information:
[0119] Use a pre-trained language model (such as BERT) to encode the title and transcript to obtain the title vector and text record vector Assumptions and is a 768-dimensional vector.
[0120] Use audio feature extraction tools to extract audio features, such as pitch, timbre, rhythm, etc., to obtain audio vectors , with a dimension of 128.
[0121] Use a convolutional neural network to process video frames and image information to obtain a visual information vector , with a dimension of 256.
[0122] Perform one-hot encoding on the metadata to obtain a metadata vector , with a dimension of 32.
[0123] Domain vector Obtained through a pre-trained domain model, with a dimension of 768.
[0124] Step 203: Use the cross-domain multimodal integration method to fuse the domain representation with the multimodal representation to obtain a cross-domain multimodal representation vector:
[0125] 1) Multimodal feature integration:
[0126] According to the formula , use the cross-modal Transformer mechanism to integrate features of different modalities. Assume that after calculation, the multimodal integration representation is a 768-dimensional vector.
[0127] 2) Calculate the multi-domain integration vector representation:
[0128] According to the formula , use the co-attention mechanism to obtain . Assume that the final multi-domain integration vector representation is a 768-dimensional vector.
[0129] 3) Cross-domain multimodal vector representation:
[0130] According to the formula , fuse the multimodal integration representation and the multi-domain integration vector representation to obtain a cross-domain multimodal representation vector , with a dimension of 768. Assume .
[0131] Step 204, design a text aggregation method to retrieve the external knowledge most relevant to the target short video:
[0132] 1) Text preprocessing: Clean the title of the target short video "Foot bath with a magical herbal medicine pack, diabetes cured instantly!" and the text (the host in the video constantly emphasizes the magical effect of the herbal medicine pack, claiming that as long as you keep using it to soak your feet, you don't need to take medicine or get injections, and diabetes can be cured). Use natural language processing tools to remove stop words (such as "of", "just", "can", etc.), punctuation marks, etc. Through word segmentation technology, split the text into individual words or phrases to obtain the core text information "magical herbal medicine pack foot bath cure diabetes".
[0133] The purpose is to simplify the text structure, highlight key information, and facilitate subsequent matching and similarity calculation with external knowledge texts.
[0134] 2) Build an external knowledge base: Collect authoritative research reports, professional medical journal articles, expert opinions, disease treatment guidelines and other information in the medical field. Use web crawler technology to obtain relevant text data from well-known medical websites, academic databases and other sources. Clean, deduplicate and classify the acquired data to build a comprehensive and accurate external knowledge base in the medical field.
[0135] This knowledge base serves as the basis for subsequent retrieval of relevant external knowledge and provides rich reference for judging the authenticity and scientificity of the target short video content.
[0136] 3) Text aggregation:
[0137] 3-1) Word vector conversion:
[0138] Use a pre-trained word vector model (such as Word2Vec (word vector)) to convert the pre-processed target short video text and external knowledge text into word vectors. Assume that the dimension of the word vector is 300. For the pre-processed text of the target short video "Magic herbal bag foot bath cures diabetes", each word is converted into a 300-dimensional vector, and then these word vectors are merged into a 300-dimensional vector representing the entire pre-processed text through methods such as average pooling, recorded as .
[0139] For external knowledge texts in the external knowledge base, the vocabulary vector conversion and merging operations are also performed. For example, for external knowledge text 1 "Medical research shows that diabetes is a chronic metabolic disease that cannot be cured by foot bathing at present. The main treatment methods are drug therapy and diet control", it is converted into a 300-dimensional vector ; For external knowledge text 2 "Single foot bath therapy has little effect on the treatment of diabetes and cannot replace regular medical treatment", convert it into a vector .
[0140] 4) Similarity calculation:
[0141] 4-1) Use the cosine similarity formula to calculate the similarity score between the preprocessed text vector and each external knowledge text vector. The cosine similarity formula is: ,in and are two vectors, represents the vector dot product, Represents vectors and Model.
[0142] 4-2) Calculation and Similarity:
[0143] set up , , vector dot product .
[0144] Model , Model .but , assuming the result is .
[0145] 4-3) Calculate in the same way and The similarity of .
[0146] 4-4) According to the set similarity threshold (for example, 0.7), filter out external knowledge texts with similarity higher than the threshold. In this example, both external knowledge text 1 and external knowledge text 2 meet the conditions, and aggregate them as the most relevant external knowledge to the target short video.
[0147] 5) Vector representation generation:
[0148] The pre-trained BERT model is used to convert the filtered external knowledge texts (external knowledge text 1 and external knowledge text 2) into vector representations. The vector dimension output by the BERT model is 768 dimensions.
[0149] At the same time, the 3-head attention mechanism is used to generate multiple vector representations for each external knowledge text. For external knowledge text 1, after being processed by the BERT model and the 3-head attention mechanism, three 768-dimensional vectors are obtained. ; For external knowledge text 2, get vector .
[0150] These vector representations not only contain the semantic information of the external knowledge text, but also capture the potential relationship between the external knowledge text and the target short video from different perspectives through a multi-head attention mechanism.
[0151] 6) Subsequent application of similarity:
[0152] Application in multi-head knowledge aggregation: In the above “multi-head knowledge aggregation, learning the relationship representation vector between different external knowledge and candidate videos”, the similarity score will affect the allocation of attention weights. The similarity is not directly reflected in the training, but in the actual training process, the attention weight corresponding to the external knowledge with high similarity to the target short video may be adjusted to be larger. For example, due to the similarity between the text external knowledge 1 and the target short video Higher similarity than external knowledge text 2 , then when setting the weight, the external knowledge text 1 corresponds to The sum of may be relatively large, making the external knowledge text 1 in calculating the relationship representation vector have a greater impact on the results.
[0153] Step 205: Aggregate multi-head knowledge to learn the relationship representation vectors between different external knowledge and candidate videos:
[0154] Alternative video representation , which is the cross-domain multimodal representation vector obtained above. The weights of different external knowledge corresponding to multiple attentions are:
[0155] For the three vectors of external knowledge text 1: ; For the three vectors of external knowledge text 2: .
[0156] According to the formula , calculate the relationship representation vector . Assume that after calculation .
[0157] Step 206: Obtain a fusion vector using a gating mechanism:
[0158] Adaptive multi-head gate vector is a 768-dimensional vector, with the first 384 dimensions being 0.6 and the last 384 dimensions being 0.4. . Assuming that after calculation , according to the formula A fusion vector may be determined.
[0159] Step 207: Input the vector into the MLP classifier:
[0160] Input vector: The fused vector As the input vector is input to the MLP classifier, it is simplified here to .
[0161] The fusion vector is input into the multi-layer perceptron classifier. The function processes the output of the multilayer perceptron classifier and obtains the probability of each category as .
[0162] pass The function finds the index of the class with the highest probability, in this case, is the maximum probability value, and its index is 1, so , corresponding to category 1, that is, the short video is judged to be misleading.
[0163] Step 208: Output the misleading short video detection result:
[0164] Assume that the short video is actually misleading, so the true label ; According to the cross entropy loss function Calculate the loss. and Substituting into the formula, we get: , .
[0165] because There is no mathematical definition, but in practical calculations, when Very close to 0, will be a large positive number. In deep learning, to avoid this situation, it is usually Add a very small positive number to ,but , then we can get .
[0166] The present invention proposes a misleading short video detection method and device. The method is based on content understanding and external knowledge enhancement. By building a multi-domain external knowledge base, it effectively solves the problem of data sparsity and uneven distribution in different fields. At the same time, it combines multimodal feature fusion and adaptive gating mechanism to significantly improve the detection accuracy and the adaptive ability of the model. Furthermore, by introducing external knowledge and domain vectors, the method can more effectively identify misleading short videos that have been carefully edited and tampered with, providing a new solution for cross-platform, cross-event, and cross-domain misleading short detection, which has important application value and promotion prospects.
[0167] Based on the same inventive concept, an embodiment of the present invention provides a device for constructing a compliance indicator system based on a large language model. Since the principle of the device for solving the technical problem is similar to the method for constructing a compliance indicator system based on a large language model, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be repeated.
[0168] like Figure 2 As shown, the device includes an acquisition unit 201 , a first obtaining unit 202 , a second obtaining unit 203 and a third obtaining unit 204 .
[0169] The acquisition unit 201 is used to acquire text features, audio features, visual features, metadata features and domain features of the target short video according to the received target short video;
[0170] A first obtaining unit 202 is used to obtain a multimodal integrated representation and a multi-domain integrated vector representation in sequence according to the text features, audio features, visual features, metadata features and domain features, and obtain a cross-domain multimodal vector representation according to the multimodal integrated representation and the multi-domain integrated vector representation;
[0171] The second obtaining unit 203 is used to perform preprocessing according to text features and construct an external knowledge base, obtain external knowledge text related to the target short video based on a text aggregation method based on word vectors, encode the external knowledge text, and obtain a vector representation corresponding to the external knowledge text; determine the cross-domain multimodal vector representation corresponding to the target short video as an alternative video representation, and obtain a relationship representation vector between the external knowledge text and the alternative video according to the vector representation corresponding to the external knowledge text and the alternative video representation;
[0172] The third obtaining unit 204 is used to obtain a fused vector based on the adaptive multi-head gating vector, the cross-domain multimodal vector representation and the relationship representation vector that can be learned according to the training process; the fused vector is input to the multi-layer perceptron classifier, and the predicted label of the output of the multi-layer perceptron classifier is determined, and the category with the largest predicted label is selected as the final detection result.
[0173] It should be understood that the units included in the above-mentioned misleading short video detection device are only logical divisions based on the functions implemented by the device. In practical applications, the above-mentioned units can be superimposed or split. And the functions implemented by the misleading short video detection device provided in this embodiment correspond one-to-one to the misleading short video detection method provided in the above-mentioned embodiment. The more detailed processing flow implemented by the device has been described in detail in the above-mentioned method embodiment 1, and will not be described in detail here.
[0174] Another embodiment of the present invention also provides a computer device, which includes: a processor and a scene database; the scene database is used to store computer program code, and the computer program code includes computer instructions; when the processor executes the computer instructions, the electronic device executes each step of the misleading short video detection method in the method flow shown in the above method embodiment.
[0175] Another embodiment of the present invention further provides a computer-readable storage medium, which stores computer instructions. When the computer instructions are executed on a computer device, the computer device executes each step of the misleading short video detection method in the method flow shown in the above method embodiment.
[0176] Although the preferred embodiments of the present invention have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0177] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A misleading short video detection method, characterized in that: include: According to the received target short video, obtain text features, audio features, visual features, metadata features and domain features of the target short video; A multimodal integrated representation and a multi-domain integrated vector representation are obtained in sequence according to the text features, audio features, visual features, metadata features and domain features, and a cross-domain multimodal vector representation is obtained according to the multimodal integrated representation and the multi-domain integrated vector representation; Preprocessing is performed according to text features and an external knowledge base is constructed. An external knowledge text related to the target short video is obtained based on a text aggregation method based on word vectors. The external knowledge text is encoded to obtain a vector representation corresponding to the external knowledge text; the cross-domain multimodal vector representation corresponding to the target short video is determined as an alternative video representation, and a relationship representation vector between the external knowledge text and the alternative video is obtained based on the vector representation corresponding to the external knowledge text and the alternative video representation; A fused vector is obtained according to the adaptive multi-head gating vector, the cross-domain multimodal vector representation and the relationship representation vector that can be learned according to the training process; the fused vector is input into the multi-layer perceptron classifier, and the predicted label output by the multi-layer perceptron classifier is determined, and the category with the largest predicted label is selected as the final detection result.
2. The method according to claim 1, characterized in that The method of sequentially obtaining a multimodal integrated representation and a multi-domain integrated vector representation according to the text features, audio features, visual features, metadata features, and domain features, and obtaining a cross-domain multimodal vector representation according to the multimodal integrated representation and the multi-domain integrated vector representation specifically includes: Determine a domain vector according to the domain features, determine a title vector and a text record vector according to the text features, determine an audio vector according to the audio features, determine a visual vector according to the visual features, and determine a metadata vector according to the metadata features; The domain vector, title vector and text record vector are represented by a multi-domain integrated vector through a collaborative attention mechanism; The title vector, text record vector, audio vector, visual information vector and metadata vector are expressed in a multimodal integrated manner through a cross-modal Transformer mechanism; A cross-domain multimodal vector representation is obtained according to the multi-domain integrated vector representation and the multimodal integrated representation.
3. The method according to claim 2, characterized in that The cross-domain multimodal vector representation is expressed by the following formula: in, represents a cross-domain multimodal vector representation, represents multimodal integrated representation, represents the multi-domain integrated vector representation, represents the domain vector, represents the title vector, represents a text record vector, represents the audio vector, represents the visual information vector, represents a metadata vector, represents the collaborative attention mechanism, represents the cross-modal attention encoder, Represents a vector concatenation operation.
4. The method according to claim 1, characterized in that The preprocessing according to the text features and building the external knowledge base specifically includes: Performing text preprocessing on the title features and text record features included in the target short video to obtain a preprocessed text; Collecting content related to the domain features included in the target short video to build an external knowledge base, and performing a search based on the external knowledge base to obtain a plurality of external knowledge texts related to the target short video; The text aggregation method based on word vector converts the preprocessed text and multiple external knowledge texts into word vectors, and calculates the similarity between the word vectors of the preprocessed text and each external knowledge text based on the cosine similarity; The plurality of external knowledge texts are converted into vector representations according to the pre-trained model, and a plurality of vector representations are generated for each of the external knowledge texts according to the attention mechanism.
5. The method according to claim 1, characterized in that The step of determining the cross-domain multimodal vector representation corresponding to the target short video as the candidate video representation, and obtaining a relationship representation vector between the external knowledge text and the candidate video according to the vector representation corresponding to the external knowledge text and the candidate video representation, specifically includes: Determine the number of vector representations included in each of the external knowledge texts and the number of weights included in each of the external knowledge texts, and obtain the relationship representation vector according to the vector representations corresponding to the plurality of external knowledge texts, the weights corresponding to the plurality of external knowledge texts, and the candidate video representations through the following formula: in, A vector representation representing the relationship between the external knowledge text and the candidate video, represents the embedding operation of vector dimension transformation, represents the comprehensive score of external knowledge, Indicates alternative video representation, Represents the loop index variable, Indicates The weight corresponding to the external knowledge text, Indicates The vector representation corresponding to the external knowledge text, Indicates the number of external knowledge texts.
6. The method according to claim 1, characterized in that The fusion vector is determined by the following formula: in, represents the fusion vector, represents the adaptive multi-head gating vector, represents a cross-domain multimodal vector representation, A vector representation representing the relationship between the external knowledge text and the candidate video, Represents an element-wise multiplication operation.
7. The method according to claim 1, characterized in that The predicted label is determined by the following formula: After selecting the category with the largest predicted label as the final detection result, the loss value is determined according to the following formula: in, represents the predicted label of the data, Represents the function of selecting the category with the highest probability, represents the normalized exponential function that calculates the probability of each category, represents a multi-layer perceptron classifier, Indicates category, represents the true label of the data, represents the fusion vector, Represents the loss value.
8. A misleading short video detection device, characterized in that: include: An acquisition unit, used for acquiring text features, audio features, visual features, metadata features and domain features of the target short video according to the received target short video; A first obtaining unit is used to obtain a multimodal integrated representation and a multi-domain integrated vector representation in sequence according to the text features, audio features, visual features, metadata features and domain features, and obtain a cross-domain multimodal vector representation according to the multimodal integrated representation and the multi-domain integrated vector representation; The second obtaining unit is used to perform preprocessing according to text features and construct an external knowledge base, obtain external knowledge text related to the target short video based on a text aggregation method based on word vectors, encode the external knowledge text, and obtain a vector representation corresponding to the external knowledge text; determine the cross-domain multimodal vector representation corresponding to the target short video as an alternative video representation, and obtain a relationship representation vector between the external knowledge text and the alternative video according to the vector representation corresponding to the external knowledge text and the alternative video representation; The third obtaining unit is used to obtain a fused vector according to the adaptive multi-head gating vector, the cross-domain multimodal vector representation and the relationship representation vector that can be learned according to the training process; the fused vector is input into the multi-layer perceptron classifier, and the predicted label of the output of the multi-layer perceptron classifier is determined, and the category with the largest predicted label is selected as the final detection result.
9. A computer device, characterized in that: The computer device includes a scene database and a processor, wherein the scene database stores a computer program, and when the computer program is executed by the processor, the processor executes the misleading short video detection method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that: A computer program is stored, and when the computer program is executed by a processor, the processor executes the misleading short video detection method as described in any one of claims 1-7.
Citation Information
Patent Citations
Multi-mode intelligent monitoring system and method
CN101753992A
Multi-modal event knowledge graph construction method
CN114064918A
False news confrontation detection system and method linked with external knowledge base
CN118606788A
Knowledge base construction method based on video content reading analysis
CN118966329A
Short video release information detection method, system and device and medium
CN119031185A