A Misleading Short Video Detection Method, System, Device and Medium

Generating diverse comments through large language models and replacing low-quality titles, the problem of modal imbalance in misleading short video detection is solved, and more efficient multimodal feature fusion is achieved, which improves detection accuracy and robustness.

CN120126057BActive Publication Date: 2025-07-18NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510577996.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-07-18
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

When existing misleading short video detection methods face the problem of modal distribution imbalance, especially when comment data is missing and title quality is low, it is difficult to effectively identify misleading content, resulting in insufficient detection accuracy and robustness.

Method used

Comments based on different user portraits are generated through a large language model, semantic correlation and entropy increments between comment features, title features and video features are obtained, diversified comment features are selected using greedy algorithms, and low-quality titles are replaced by calculating the similarity between titles and video features, ensuring the consistency of features, and using a collaborative attention mechanism for fusion.

Benefits of technology

It effectively solves the problems of missing comment data and low title quality, improves the accuracy and robustness of misleading short video detection, and can accurately identify misleading content under multimodal conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126057B_ABST
    Figure CN120126057B_ABST
Patent Text Reader

Abstract

The present invention discloses a misleading short video detection method, system, device and medium, which relates to the technical field of video detection, and includes the steps of: obtaining the content information and dissemination information of the short video; simulating different user portraits based on the original comments of the short video, and generating comments reflecting different user views and emotions according to different user portraits; extracting the title features, video features and comment features of the short video, obtaining the semantic relevance between the comment features and the title features and video features, obtaining a new comment score according to the weighted combination of the semantic relevance and the entropy increment, and selecting comment features based on the comment score; obtaining the relevance between the title features and the video features, replacing the title with a quality lower than the threshold through the relevance, and fusing the replaced title features with the selected comment features to obtain the short video detection result. The present invention can effectively address the problem of uneven modal distribution in the dataset and improve the performance of misleading short video detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video detection, and particularly to a method, system, device and medium for detecting misleading short videos. Background Art

[0002] With the popularity of short video platforms, such platforms have become an important channel for information dissemination. Short videos, with their multimodal presentation forms of text, audio, and visual information, are more likely to attract users' trust and dissemination compared to the traditional "picture + text" information dissemination method. Misleading short videos bring serious negative impacts to society by misleading the public. With the development of generative AI and short video platforms, the dissemination forms of misleading short videos are becoming increasingly diverse, and their confusion and concealment are significantly enhanced, making it difficult for traditional unimodal detection methods to cope with complex multimodal misleading short video scenarios. The research on multimodal misleading short video detection can not only capture the inconsistencies between modalities by integrating multi-source information such as text, audio, and video, improve the accuracy and robustness of detection, but also cope with the dissemination characteristics of cross-modal misleading information, provide technical support for the content governance of emerging media platforms such as short videos, and reduce the social harm of false information.

[0003] The detection of misleading short videos involves the complex fusion of communication features such as text, audio, visual information, and comments. Existing methods for detecting misleading short videos use some or all of these features for multimodal fusion classification. However, in practical applications, it often faces the major challenge of unbalanced modality distribution in the training data. For example, SV-FEND emphasizes using title text features to align video and audio features, but it ignores the problems of missing or garbled title data, resulting in negative guidance when using low-quality titles to align other features and affecting the detection process. With the development of social media and the growth of the number of Internet users, more and more research incorporates communication features into short video detection. Communication features are mainly based on user comments, which usually contain information about opinions, emotions, and positions.

[0004] However, due to different user participation levels, comment data may be missing, making it insufficient to fully represent public opinions. Some algorithms use large language models to supplement comments, but they fail to emphasize positions, often resulting in ambiguous generated comments, that is, in reality, misleading short video datasets often face the problem of modality imbalance, such as due to limited user participation, missing comment data, and blank or garbled titles. Summary of the Invention

[0005] The purpose of the present invention is to provide a method, system, device and medium for detecting misleading short videos to solve the problems in the prior art in view of the above-mentioned deficiencies of the prior art.

[0006] The present invention specifically provides the following technical solutions:

[0007] A misleading short video detection method, comprising:

[0008] Obtain the content information and dissemination information of the short video, where the content information includes the title and the video, and the dissemination information includes the original comments;

[0009] Based on the original comments of the short video, simulate different user portraits through a large language model, and generate comments reflecting different user views and emotions according to different user portraits;

[0010] Use the comments generated to reflect different user views and emotions as the new comments of the short video, extract the title features, video features and comment features of the short video, obtain the semantic relevance between the comment features and the title features and video features, and obtain the new comment score according to the weighted combination of the semantic relevance and the entropy increment, and select the comment features with the new comment score;

[0011] Obtain the correlation between the title features and the video features, use the correlation as the quality judgment criterion, replace the title features with a quality lower than the threshold with the title features with a quality higher than the threshold, fuse the video features, the replaced title features and the selected comment features, and classify the fusion result to obtain the short video detection result.

[0012] Preferably, the extraction of the title features, video features and comment features of the short video is specifically:

[0013] Input the title into the CLIP model text encoder to extract the title features ; Input the video frames into the CLIP image encoder to extract the video features ; Input the comments into the CLIP text encoder to extract the comment features , is the comment index.

[0014] Preferably, the obtaining of the semantic relevance between the comment features and the title features and video features, and the obtaining of the new comment score according to the weighted combination of the semantic relevance and the entropy increment, and the selection of the comment features with the new comment score include:

[0015] For each comment, obtain the cosine similarity between the comment features and the title features and video features, and the specific expression is:

[0016] ;

[0017] ;

[0018] ;

[0019] In the formula, is the Cosine similarity between a comment and a title For the th comment and the cosine similarity of the video is the title weight is the video weight is the semantic relevance between the comment and the short video content information, and the comments are sorted according to the level of semantic relevance;

[0020] Use the greedy algorithm to select a comment subset of size ; where the entropy of the comment subset is defined as: entropy is defined as:

[0021] ;

[0022] Among them, is the comment probability distribution of The specific expression is:

[0023] ;

[0024] When a new comment is added to the selected set , use the entropy increment to measure the contribution of the new comment ; where the entropy increment The specific expression is:

[0025] ;

[0026] Obtain the new comment score through the weighted combination of semantic relevance and entropy increment. The specific expression is:

[0027] ;

[0028] Among them, is the new comment score, and are the weights of semantic relevance and diversity respectively, is the union;

[0029] The greedy algorithm selects the comment feature with the highest score in each iteration, adds the comment feature with the highest score to the selected set , and updates the entropy until , the set represents the comment features sorted and filtered according to the level of semantic relevance.

[0030] Preferably, the correlation between the title features and the video features is obtained, and the correlation is used as the quality evaluation criterion. The title features with quality higher than the threshold are used to replace the title features with quality lower than the threshold, specifically as follows:

[0031] Construct a similarity matrix by calculating the cosine similarity between the title features and video features of different short videos , and the specific expression is:

[0032] ;

[0033] In the formula, is the title feature of the th short video, is the video feature of the th short video, is the cosine similarity calculation, is the th short video's title feature and the th short video's video feature's cosine similarity, where the cosine similarity here is the correlation;

[0034] Taking the correlation as the quality evaluation criterion, the title features with quality higher than the threshold are used to replace the title features with quality lower than the threshold, and the specific expression is:

[0035] ;

[0036] In the formula, is the new title feature of the th short video after being replaced, is the title feature of the th short video, is the th short video's video feature, is the function used to find the parameter that makes the function value the largest.

[0037] Preferably, the video features, the replaced title features and the selected comment features are fused, and the specific expression of the feature fusion is:

[0038] ;

[0039] In the formula, is the fused feature, is the collaborative attention mechanism, is the multi-layer perceptron model, is the title feature after replacing the low-quality title, is the comment feature after sorting by semantic relevance level and diversity screening, is the video feature.

[0040] Preferably, the simulation of different user portraits through the large language model includes:

[0041] Design prompt words using the original comments of the short video;

[0042] Based on the prompt words, use the large language model to extract the attribute information of the user, and simulate different user portraits for different comments; the attribute information includes the user's age, gender, occupation, location, and tone.

[0043] Preferably, the short video detection results include misleading videos and non-misleading videos.

[0044] The present invention provides a misleading short video detection system, including:

[0045] An acquisition module for obtaining the content information and dissemination information of the short video, where the content information includes the title and the video, and the dissemination information includes the original comments;

[0046] A comment generation module for simulating different user portraits through the large language model based on the original comments of the short video, and generating comments reflecting different user views and emotions according to different user portraits;

[0047] A selection module for using the generated comments reflecting different user views and emotions as the new comments of the short video, extracting the title features, video features, and comment features of the short video, obtaining the semantic relevance between the comment features and the title features and video features, obtaining the new comment score according to the weighted combination of the semantic relevance and the entropy increment, and selecting the comment features with the new comment score;

[0048] A detection module for obtaining the relevance between the title features and the video features, using the relevance as the quality judgment standard, replacing the title features with a quality lower than the threshold with the title features with a quality higher than the threshold, fusing the video features, the replaced title features, and the selected comment features, and classifying the fusion result to obtain the short video detection result.

[0049] The present invention provides a computer device, including a memory and a processor. When a program stored in the memory is executed by the processor, the processor executes the steps of the above-mentioned misleading short video detection method.

[0050] The present invention provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned misleading short video detection method are implemented.

[0051] Compared with the prior art, the present invention has the following remarkable advantages:

[0052] The present invention generates comments based on different user portraits through a large language model, and the comments can reflect the positions of different users. It also obtains the semantic relevance between the comment features, the title features, and the video features, and obtains a new comment score based on the weighted combination of the semantic relevance and the entropy increment to ensure the diversity of the selected comments. The comment features are selected based on the new comment score, effectively solving the problem of missing comment data. By calculating the similarity between the title and the video features, the low-quality title is automatically replaced to ensure the consistency between the title and the video content, effectively avoiding the risk of misjudgment caused by inconsistent titles. It can effectively address the problem of uneven modal distribution in the dataset and improve the performance of misleading short video detection by fusing multiple modal features. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 It is a flowchart of a method for detecting misleading short videos according to the present invention;

[0054] Figure 2 It is a framework diagram of a method for detecting misleading short videos according to the present invention;

[0055] Figure 3 It is an application case diagram of a method for detecting misleading short videos according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] The following clearly and completely describes the technical solutions of the embodiments of the present invention with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0057] The present invention provides a method for detecting misleading short videos. Refer to Figure 1 , including:

[0058] Step S1: Obtain the content information and dissemination information of the short video, where the content information includes the title and the video, and the dissemination information includes the original comments.

[0059] Step S2: Based on the original comments of the short video, simulate different user portraits through a large language model, and generate comments reflecting different user views and emotions according to different user portraits.

[0060] Simulating different user portraits through a large language model includes:

[0061] Design prompt words using the original comments of the short video.

[0062] Based on the prompt, use the large language model to extract the user's attribute information and simulate different user portraits for different comments; the attribute information includes the user's age, gender, occupation, location, and tone.

[0063] According to different user portraits, design prompts that emphasize that the comments should not only conform to the characteristics of the user portraits but also have a clear stance and tone, and generate comments that reflect different user views and emotions.

[0064] Step S3: Generate comments that reflect different user views and emotions as new comments for the short video, extract the title features, video features, and comment features of the short video, obtain the semantic relevance between the comment features and the title features and video features, and obtain the new comment score according to the weighted combination of the semantic relevance and entropy increment, and select comment features based on the new comment score.

[0065] Extract the title features, video features, and comment features of the short video, specifically:

[0066] Input the title into the text encoder of the contrastive language-image pre-training CLIP (Contrastive Language-Image Pre-Training) model to extract the title features ; Input the video frames into the CLIP image encoder to extract the video features ; Input the comment into the CLIP text encoder to extract the comment features , as the comment index.

[0067] Obtain the semantic relevance between the comment features and the title features and video features, and obtain the new comment score according to the weighted combination of the semantic relevance and entropy increment, and select diverse comment features based on the comment score, including:

[0068] For each comment, obtain the cosine similarity between the comment features and the title features and video features. The specific expression is:

[0069] ;

[0070] ;

[0071] ;

[0072] In the formula, is the cosine similarity between the th comment and the title, is the cosine similarity between the th comment and the video, is the title weight, is the video weight, For the semantic relevance of comment and short video content information, sort the comments according to the level of semantic relevance.

[0073] Use the greedy algorithm to select a comment subset of size , maximizing the diversity of the selected comments, where the diversity is measured by the entropy of the selected comment set; where the entropy of the comment subset is defined as:

[0074] ;

[0075] where is the probability distribution of comment , obtained by using the Softmax function on its features, and the specific expression is:

[0076] ;

[0077] When a new comment is added to the selected set , use the entropy increment to measure the contribution of the new comment ; where the entropy increment is specifically expressed as:

[0078] ;

[0079] Obtain the new comment score through a weighted combination of semantic relevance and entropy increment, and the specific expression is:

[0080] ;

[0081] where, is the new comment score, and are the weights of semantic relevance and diversity respectively, is the union. The greedy algorithm selects the comment feature with the highest score in each iteration, adds the comment feature with the highest score to the selected set , and updates the entropy until , and the set represents the comment features sorted and filtered according to the level of semantic relevance.

[0082] Step S4: Obtain the relevance between the title features and the video features, use the relevance as the quality judgment criterion, replace the title features with quality lower than the threshold with the title features with quality higher than the threshold, fuse the video features, the replaced title features and the selected comment features, and classify the fusion result to obtain the short video detection result.

[0083] Obtain the correlation between the title features and the video features, and use the correlation as the quality evaluation criterion. Replace the title features with a quality lower than the threshold with the title features with a quality higher than the threshold, including:

[0084] Construct a similarity matrix by calculating the cosine similarity between the title features and the video features of different short videos , , where is a predefined threshold, and the specific expression is:

[0085] ;

[0086] In the formula, is the title feature of the th short video, is the video feature of the th short video, is the cosine similarity calculation, is the cosine similarity between the title feature of the th short video and the video feature of the th short video, where the cosine similarity here is the correlation.

[0087] Use the correlation as the quality evaluation criterion, and replace the title features with a quality lower than the threshold with the title features with a quality higher than the threshold. The specific expression is:

[0088] ;

[0089] In the formula, is the new title feature after replacement of the th short video, is the title feature of the th short video, is the video feature of the th short video, is the function used to find the parameter that makes the function value the largest.

[0090] Fuse the title features after replacing the titles with low quality with the selected comment features to obtain the detection result, specifically:

[0091] The specific expression of feature fusion is:

[0092] ;

[0093] In the formula, is the fused feature, is the collaborative attention mechanism, is the multi-layer perceptron model, is the title feature after replacing the low-quality title, They are comment features sorted by semantic relevance and screened for diversity. They are video features.

[0094] Use a classifier to output classification results. The input to the classifier is video features, title features after replacing low-quality titles, and comment features, and the output is a binary classification result: misleading videos and non-misleading videos.

[0095] To highlight the superiority of the present invention, comparative experiments and ablation experiments are conducted for testing, as follows.

[0096] The present invention is tested using the FakeSV and FVC datasets. FakeSV is the largest Chinese misleading short video detection dataset, collected from Douyin and Kuaishou. FVC is an English dataset covering multiple topics such as politics, sports, and accidents, collected from YouTube.

[0097] To evaluate the effectiveness of the proposed misleading short video detection MMDA method of the present invention, comparative experiments are conducted with various baseline methods. The baseline methods are divided into the following three categories: (1) Single-modal methods: Traditional single-modal methods detect misleading short videos by analyzing single-modal features. The present invention evaluates methods such as VGGish (audio), BERT (comments), VGG19 (images), and C3D (videos). (2) Multi-modal methods: Multi-modal methods improve the performance of misleading short video detection by integrating information from multiple modalities, including methods such as Hou et al., TikTec, FANVN, SV-FEND, etc. These methods mainly focus on the integration of cross-modal content but usually fail to explicitly address low-quality or missing modalities and lack explicit modeling of propagation feature interactions. In addition, GenFEND enhances the detection ability by using large language models to generate user-based comments, but the generated comments often lack semantic diversity and clear stances. (3) Large language models: The present invention designs a specific method for detection based on prompts and conducts experiments using the APIs of GPT3.5-turbo, GPT4.0-turbo, and GPT4.0-vision.

[0098] The results of the comparative experiments are shown in Table 1. MMDA significantly outperforms all baseline methods on both the FakeSV and FVC datasets. Single-modal methods are limited in performance because they rely on single-modal information and cannot capture the richness of multimodal interactions. Although multimodal methods utilize the complementary characteristics of different modalities, they usually struggle when dealing with missing or low-quality inputs and lack fine-grained cross-modal reasoning capabilities. For example, TikTec, FANVN, and SV-FEND mainly focus on the integration of cross-modal content but fail to explicitly address low-quality or incomplete inputs. GenFEND attempts to enhance detection by using large language models to generate user-based comments, but the comments it generates often lack semantic diversity and a clear stance, limiting its effectiveness. In contrast, MMDA introduces a comment generation module (CAM) and a title enhancement module (TEM), which can effectively supplement low-quality or missing modal information and generate diverse and clearly-stanced comments. This enables MMDA to integrate enhanced multimodal features and achieve state-of-the-art performance even under challenging data conditions.

[0099] Table 1 Results of the comparative experiments

[0100]

[0101] The results of the ablation experiments are shown in Table 2. Each input and module has a significant impact on the overall performance.

[0102] Table 2 Results of the ablation experiments

[0103]

[0104] Removing video, text, or comment inputs leads to a performance drop, which demonstrates the importance of utilizing multimodal information. The comment generation module CAM plays a key role in improving the performance of detecting misleading short videos. Removing CAM leads to a significant performance drop, and replacing CAM with comments generated by GenFEND further causes a performance drop. This shows the superiority of the proposed CAM in generating high-quality, semantically diverse, and user-stance-compliant comments. Similarly, the title enhancement module TEM also proves its effectiveness in improving the model performance. By optimizing low-quality or missing news titles, TEM ensures the consistency and informativeness of text inputs, thus promoting better multimodal alignment. These findings validate the effectiveness of MMDA in addressing data incompleteness and enhancing multimodal feature representation.

[0105] Figure 3It demonstrates the effectiveness of the present invention in enhancing the detection effect of misleading short videos by using the comment generation module CAM and the title enhancement module TEM. In the first example, due to the lack of reliable comments, CAM generated comments with diverse semantics and stances, helping to identify the short video as misleading. On the other hand, since the provided title was misleading, TEM could not function, highlighting the importance of context in title enhancement. In the second example, CAM effectively utilized the generated comments to support the detection process. Additionally, TEM successfully enhanced the meaningless title, improving the quality of the title and thus achieving correct classification. Overall, the case study shows that combining CAM and TEM enables MMDA to compensate for missing or low-quality modal information and achieve stronger performance in the detection of misleading short videos.

[0106] Based on the above method, the present invention provides a misleading short video detection system, including: an acquisition module, a comment generation module, a selection module, and a detection module.

[0107] Among them, the acquisition module is used to obtain the content information and dissemination information of the short video, where the content information includes the title and the video, and the dissemination information includes the original comments; the comment generation module is used to simulate different user portraits through a large language model based on the original comments of the short video, and generate comments reflecting different user opinions and emotions according to different user portraits; the selection module is used to use the generated comments reflecting different user opinions and emotions as the new comments of the short video, extract the title features, video features, and comment features of the short video, and obtain the semantic relevance between the comment features and the title features and video features, and obtain a new comment score according to the weighted combination of the semantic relevance and the entropy increment, and select diverse comment features based on the comment score; the detection module is used to obtain the relevance between the title features and the video features, use the relevance as the quality judgment criterion, replace the title features with a quality lower than the threshold with the title features with a quality higher than the threshold, fuse the video features, the replaced title features, and the selected comment features, and classify the fusion result to obtain the short video detection result.

[0108] The present invention also provides a computer device, including a memory and a processor. When a program stored in the memory is executed by the processor, the processor executes the steps of a misleading short video detection method.

[0109] According to the disclosed embodiments, the computer device can communicate with one or more external devices (such as a keyboard, a pointing device, Bluetooth communication, etc.), or communicate with any device (such as a router, a demodulator, etc.) that enables the computing device to communicate with one or more other computing devices.

[0110] The present invention also provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of a misleading short video detection method are implemented.

[0111] According to the disclosed embodiments, the storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device.

[0112] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. For those skilled in the art of the present invention, without departing from the concept of the present invention, several simple deductions or substitutions can be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A misleading short video detection method, characterized in that, including: Obtain the content information and dissemination information of the short video, where the content information includes the title and the video, and the dissemination information includes the original comments; Based on the original comments of the short video, simulate different user portraits through a large language model, and generate comments reflecting different user views and emotions according to different user portraits; Use the comments reflecting different user views and emotions as the new comments of the short video, extract the title features, video features and comment features of the short video, obtain the semantic relevance between the comment features and the title features and video features, and obtain the new comment score according to the weighted combination of the semantic relevance and the entropy increment, and select the comment features with the new comment score; The entropy increment is specifically: Select a comment subset of size using a greedy algorithm ; where the entropy of the comment subset is defined as: ​ ; Among them, is the probability distribution of comments , and the specific expression is: Specific expression is: ; When a new comment is added to the selected set , the entropy increment is used to measure the contribution of the new comment ; where the entropy increment has the specific expression as follows: ; Obtain the relevance between the title features and the video features, use the relevance as the quality judgment criterion, replace the title features with a quality lower than the threshold with the title features with a quality higher than the threshold, fuse the video features, the replaced title features and the selected comment features, and classify the fusion result to obtain the short video detection result; The obtaining of the relevance between the title features and the video features, using the relevance as the quality judgment criterion, and replacing the title features with a quality lower than the threshold with the title features with a quality higher than the threshold is specifically: Construct a similarity matrix by calculating the cosine similarity between different short video title features and video features , and the specific expression is as follows: ; Wherein, is the title feature of the th short video, is the video feature of the th short video, is the cosine similarity calculation, is the cosine similarity between the title feature of the th short video and the video feature of the th short video, where the cosine similarity here is the correlation; Using the relevance as the quality judgment criterion, replacing the title features with a quality lower than the threshold with the title features with a quality higher than the threshold, the specific expression is: ; In the formula, is the new title feature after the th short video is replaced, is the title feature of the th short video, is the video feature of the th short video, is a function used to find the parameter that maximizes the function value.

2. The misleading short video detection method according to claim 1, wherein The extraction of the title features, video features and comment features of the short video is specifically: Input the title into the CLIP model text encoder to extract title features ; Input the video frame into the CLIP image encoder to extract video features ; Input the comment into the CLIP text encoder to extract comment features , is the comment index.

3. The misleading short video detection method according to claim 2, wherein, The obtaining of the semantic relevance between the comment features and the title features and video features, and obtaining the new comment score according to the weighted combination of the semantic relevance and the entropy increment, and selecting the comment features with the new comment score includes: For each comment, obtain the cosine similarity between the comment features and the title features and video features, and the specific expression is: ; ; ; Wherein, is the cosine similarity between the th comment and the title, is the cosine similarity between the th comment and the video, is the title weight, is the video weight, is the semantic relevance between the comment and the short video content information, and the comments are sorted according to the level of semantic relevance; Obtain the new comment score through the weighted combination of the semantic relevance and the entropy increment, and the specific expression is: ; Among them, is the new comment score, and are the weights of semantic relevance and diversity respectively, is the union set; The greedy algorithm selects the comment feature with the highest score in each iteration, adds the comment feature with the highest score to the selected set , and updates the entropy until , the set represents the comment features after sorting and screening by semantic relevance level.

4. The misleading short video detection method according to claim 1, wherein The fusion of the video features, the replaced title features and the selected comment features, where the specific expression of the feature fusion is: ; Wherein, is the fused feature, is the collaborative attention mechanism, is the multi-layer perceptron model, is the title feature after replacing the low-quality title, is the comment feature after sorting by semantic relevance and diversity screening, is the video feature.

5. The misleading short video detection method according to claim 1, wherein The simulation of different user portraits through the large language model includes: Use the original comments of the short video to design prompt words; Based on the prompt words, use the large language model to extract the attribute information of the user, and simulate different user portraits for different comments; The attribute information includes the user's age, gender, occupation, location and tone.

6. The misleading short video detection method according to claim 1, wherein The short video detection result includes misleading videos and non-misleading videos.

7. A misleading short video detection system, applied to a misleading short video detection method according to any one of claims 1-6, characterized in that including: A collection module for obtaining the content information and dissemination information of the short video, where the content information includes the title and the video, and the dissemination information includes the original comments; A comment generation module for simulating different user portraits through a large language model based on the original comments of the short video, and generating comments reflecting different user views and emotions according to different user portraits; A selection module for using the comments reflecting different user views and emotions as the new comments of the short video, extracting the title features, video features and comment features of the short video, obtaining the semantic relevance between the comment features and the title features and video features, and obtaining the new comment score according to the weighted combination of the semantic relevance and the entropy increment, and selecting the comment features with the new comment score; A detection module, which is used to obtain the correlation between the title features and the video features, uses the correlation as a quality evaluation criterion, replaces the title features with a quality lower than the threshold with the title features with a quality higher than the threshold, fuses the video features, the replaced title features and the selected comment features, and classifies the fusion result to obtain a short video detection result.

8. A computer device, characterized in that, It includes a memory and a processor. A program is stored in the memory. When the program is executed by the processor, the processor is caused to execute the steps of a misleading short video detection method according to any one of claims 1 to 6.

9. A storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by the processor, the steps of a misleading short video detection method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Method and apparatus for authenticating video content

    CN104205865A

  • Multi-modal visual language understanding and positioning method and device, terminal and medium

    CN116091836A