Video review auditing method and apparatus, computer device, and storage medium
By constructing a multimodal model and combining the correlation between video frames and comment data, key information and sentiment information are extracted, similarity and conflict degree are calculated, and implicit information in video comments is identified. This solves the problem that existing technologies cannot identify the satirical meaning of video comments and enables effective review of video comments.
Patent Information
- Application Number
- CN202411975512.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2044-12-30
AI Technical Summary
Existing technologies lack methods for reviewing the satirical meanings implied in video comments, making it impossible to effectively identify the potential risk of negative energy spreading in scenarios where videos and comments are combined.
By constructing a multimodal model, combining the correlation between video frames and comment data, key information and sentiment information are extracted, similarity and conflict levels are calculated, negative comments are identified, and a large multimodal model is used for fine-tuning training to identify implicit information in video comments.
It enables effective review of implicit information in video comments, reduces the risk of public opinion dissemination, and improves the stability of comment quality and review efficiency.
Smart Images

Figure CN119903392B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of broadcast media content review, in particular to a video comment review method and device, computer equipment and a storage medium. BACKGROUND
[0002] In the process of information dissemination such as text, sound, image and video in the field of audiovisual media, content review can help filter illegal, malicious and other bad content, and ensure that the dissemination content is legal, compliant and correct in value orientation.
[0003] Current content review mainly combines manual content review and intelligent content review. In the field of intelligent review, different intelligent review classification methods are used according to different data types. In the text field, natural processing technology is used to automatically analyze and understand the text; in the image field, computer vision technology is used to automatically identify and filter illegal images through image segmentation and target detection; in the audio field, audio is converted into text for judgment; in the video field, video is decomposed into video frames and audio for review.
[0004] In related technologies, a semantic recognition method and system for AI automatic review of scripts and pictures are provided. The system receives script data and picture data to be published, cleanses all data, performs semantic understanding and sentiment analysis on the script data, and performs picture processing on the picture data, including feature extraction, target detection and semantic segmentation. Finally, the processed data is subjected to content review. A multi-modal data fusion classification method based on a large model and an attention mechanism is also provided. This method uses a cross-attention fusion model to fuse image and text data, and directly fuses different image features, which not only fuses the feature information between different images, but also avoids the risk of overfitting caused by excessive fusion, reduces information redundancy and noise, and can better balance the text and image modalities to improve the accuracy of the classification results.
[0005] However, in related technologies, natural language processing technology is used for text analysis, and computer vision technology is used for image analysis, which cannot directly identify negative information content in multi-modal data.
[0006] Currently, with the development of the network, people's way of obtaining information is becoming more and more popular. Content carriers such as short videos have become the main body of people discussing and sharing life. The emergence of online comments allows users to discuss any topic and share their views on video content, while network language has become an interesting way for young people to express their views on the network.
[0007] The implicit information of the video comment refers to the sentiment, attitude and opinion not explicitly expressed in the comment. In the scenario of video combined with comments, usually these comments themselves have no sarcastic meaning problem as words, but combined with the image or video context, they have sarcastic meaning and can cause the spread of negative energy. Therefore, the implicit meaning of such comments needs to be audited in advance to avoid the risk of public opinion spread.
[0008] In the related art, there is no method directly used for sarcastic meaning auditing of video comments. In the academic field, the sentiment analysis technology is usually used to detect the implicit information of the comments. At present, there are mainly three methods: dictionary-based technology, machine learning-based technology and hybrid method. The dictionary-based technology generally uses statistical analysis, such as nearest neighbor and conditional random field, which is easier to detect when the sentiment words are explicitly expressed. The machine learning-based technology is divided into traditional machine learning and deep learning technology, which can understand the lexical features by using recurrent neural network, etc. These methods are usually converted into document, sentence level classification problems. The hybrid method is to mix the dictionary-based and deep learning methods to enhance the accuracy of text sentiment analysis. The current sentiment analysis is mainly for sentiment intensity and polarity analysis of the text itself, and there is no implicit information auditing of text comments combined with images.
[0009] In the related art, a multi-dimensional comment auditing method is provided, which audits the comments from multiple dimensions. The repetition degree is used to detect the novelty of the comment to be audited. The text richness is used to detect whether the content of the comment to be audited is single. The sentiment recognition is used to detect the positive sentiment of the comment to be audited. The timeliness is used to measure the publishing timeliness of the comment to be audited. Thus, the comment is audited from the above four dimensions, the quality of the comment is quantified, and the stability and quality of the comment quality are ensured. A news comment auditing method is also provided, which includes obtaining a comment initiated by a user end, identifying the text and pictures in the comment; extracting the text and elements in the picture, matching and identifying the elements to determine whether the elements contain illegal elements, if so, removing the picture; performing semantic monitoring on the text and picture text, if the monitoring result is a sensitive comment, a spam comment or an excessive comment, determining that the comment is an illegal comment, removing the comment and recording the comment in the database; if it is unable to determine whether the comment is an illegal comment, obtaining the user situation of the user publishing the comment, further determining the comment according to the user situation to determine whether the comment is illegal. This method has the effect of increasing the auditing efficiency and reducing the passing rate of illegal comments.
[0010] Therefore, in the related art, the auditing is mainly performed on the text and image, and there is no method for auditing the video. In addition, the implicit meaning of the comment cannot be determined, that is, the comment itself has no risk, and the situation of the comment producing content risk in the video scenario cannot be determined. SUMMARY
[0011] The embodiment of the application provides a video comment auditing method and device, computer equipment and a storage medium.
[0012] In a first aspect, the embodiment of the application provides a video comment auditing method, comprising:
[0013] obtaining a target video and corresponding comment data thereof;
[0014] annotating the comment data corresponding to the target video based on the correlation between the target video and the corresponding comment data thereof, and constructing a comment data set corresponding to the target video;
[0015] using the target video and the comment data set corresponding thereto as a training set, fine-tuning a preset multi-modal model to obtain a trained first multi-modal model;
[0016] receiving a specified video and corresponding comment data thereof, inputting the specified video and the corresponding comment data thereof into the first multi-modal model, obtaining an auditing result of the comment data of the specified video, and processing the comment data with a negative comment.
[0017] In an optional embodiment of the application, the annotation of the comment data corresponding to the target video based on the correlation between the target video and the corresponding comment data thereof comprises:
[0018] extracting key information and emotional information in the comment data corresponding to the target video, and preliminarily determining whether the comment data is a negative comment according to the emotional information;
[0019] in the case of preliminarily determining that the comment data is a negative comment, labeling the comment data as a negative comment;
[0020] in the case of preliminarily determining that the comment data is a positive comment, further determining whether the comment data is a negative comment based on the correlation between the target video and the corresponding comment data thereof;
[0021] in the case of further determining that the comment data is a negative comment, labeling the comment data as a negative comment;
[0022] in the case of further determining that the comment data is a positive comment, labeling the comment data as a positive comment.
[0023] In an optional embodiment of the application, the further determination of whether the comment data is a negative comment based on the correlation between the target video and the corresponding comment data thereof comprises:
[0024] extracting key information and emotional information in the target video, and calculating the similarity between the key information and the emotional information in the target video and the corresponding comment data thereof.
[0025] determine a conflict degree between the target video and the text content in the comment data corresponding to the target video;
[0026] determine a conflict score between the target video and the comment data corresponding to the target video according to the similarity between the key information and the sentiment information in the target video and the comment data corresponding to the target video and the conflict degree between the text content;
[0027] determine the comment data with a conflict score exceeding a preset conflict score threshold as a negative comment.
[0028] In an optional embodiment of the present application, the similarity between the key information and the sentiment information in the target video and the comment data corresponding to the target video is calculated by the following expression:
[0029]
[0030] wherein A is a first text vector composed of the key information and the sentiment information in the target video, B is a second text vector composed of the key information and the sentiment information in the comment data corresponding to the target video, a i is the i-th vector element in the first text vector, b i is the i-th vector element in the second text vector, and w1 and w2 are weight factors.
[0031] In an optional embodiment of the present application, the constructing of the comment data set corresponding to the target video comprises:
[0032] collecting each comment data corresponding to the target video and the annotation information of each comment data into the comment data set corresponding to the target video.
[0033] In an optional embodiment of the present application, the fine-tuning training of the preset multi-modal model using the target video and the comment data set corresponding to the target video as a training set to obtain the trained first multi-modal model comprises:
[0034] using the video frame of the target video and each comment data as input, using the annotation information of the current comment data as output, training the network structure in the preset multi-modal model for processing the video frame in a lora fine-tuning manner of a supervised task, training the network structure for processing the comment data in a qlora fine-tuning manner of a supervised task, and obtaining the trained first multi-modal model.
[0035] In an optional embodiment of the present application, the inputting of the specified video and the comment data corresponding to the specified video into the first multi-modal model to obtain the review result of the comment data of the specified video comprises:
[0036] input the video frame of the specified video into a network structure in the first multi-modal model for processing the video frame to obtain a video frame feature;
[0037] input the comment data corresponding to the specified video into a network structure in the first multi-modal model for processing the comment data to obtain a comment data feature;
[0038] input the video frame feature and the comment data feature into a fusion network structure in the first multi-modal model to obtain an audit result of the comment data of the specified video.
[0039] A second aspect of the embodiments of the present application provides a video comment audit device, comprising:
[0040] an acquisition module configured to acquire a target video and comment data corresponding to the target video;
[0041] a construction module configured to label the comment data corresponding to the target video based on the correlation between the target video and the comment data corresponding to the target video, and construct a comment data set corresponding to the target video;
[0042] a training module configured to use the target video and the comment data set corresponding to the target video as a training set to fine-tune a preset multi-modal model to obtain a trained first multi-modal model;
[0043] an input module configured to receive a specified video and comment data corresponding to the specified video, input the specified video and the comment data corresponding to the specified video into the first multi-modal model, obtain an audit result of the comment data of the specified video, and process the comment data with a negative comment.
[0044] A third aspect of the embodiments of the present application provides a computer device, comprising a memory and a processor, the memory stores a computer program, and the processor implements the steps of any one of the above methods when executing the computer program.
[0045] A fourth aspect of the embodiments of the present application provides a computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the steps of any one of the above video comment audit methods.
[0046] The above technical solutions provided by the embodiments of the present application have at least some or all of the following advantages compared with the prior art:
[0047] The video comment auditing method provided by the embodiment of the application comprises the following steps: obtaining a target video and comment data corresponding to the target video; labeling the comment data corresponding to the target video based on the correlation between the target video and the comment data corresponding to the target video, and constructing a comment data set corresponding to the target video; taking the target video and the comment data set corresponding to the target video as a training set, and fine-tuning a preset multi-modal model to obtain a first multi-modal model after training; receiving a specified video and comment data corresponding to the specified video, inputting the specified video and the comment data corresponding to the specified video into the first multi-modal model, obtaining an auditing result of the comment data of the specified video, and processing the comment data with a negative comment result. The video frame of the target video is understood, the comment of the target video is labeled, the preset multi-modal model is trained and fine-tuned in combination with the video frame and the comment and the label of the comment, and the first multi-modal model after training is used for identifying the hidden information of the video comment. BRIEF DESCRIPTION OF DRAWINGS
[0048] The accompanying drawings, which are included to provide a further understanding of the application, constitute a part of the application and serve to explain the application, but should not be construed as limiting the application. In the drawings:
[0049] Figure 1 The flowchart of the video comment auditing method provided by an embodiment of the application;
[0050] Figure 2 The flowchart of the video comment auditing method provided by another embodiment of the application;
[0051] Figure 3 The detailed flowchart of the video comment data labeling step provided by an embodiment of the application;
[0052] Figure 4 The flowchart of the content auditing method provided by an embodiment of the application;
[0053] Figure 5 The schematic diagram of the content auditing method based on a large model provided by an embodiment of the application;
[0054] Figure 6 The schematic diagram of the text content auditing method provided by an embodiment of the application;
[0055] Figure 7 The schematic diagram of the image scene content auditing method provided by an embodiment of the application;
[0056] Figure 8 The schematic diagram of the video comment auditing device structure provided by an embodiment of the application;
[0057] Figure 9A computer device structure schematic diagram provided by an embodiment of the present application. DETAILED DESCRIPTION
[0058] In order to make the technical solutions and advantages of the embodiments of the present application clearer, the exemplary embodiments of the present application are further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.
[0059] Please refer to Figure 1 and Figure 2 The video comment auditing method provided by the embodiments of the present application includes the following steps 100-400:
[0060] Step 100, obtaining a target video and its corresponding comment data;
[0061] Step 200, based on the association between the target video and its corresponding comment data, labeling the comment data corresponding to the target video, and constructing a comment data set corresponding to the target video;
[0062] Step 300, taking the target video and its corresponding comment data set as a training set, fine-tuning the pre-set multi-modal model, and obtaining a trained first multi-modal model;
[0063] Step 400, receiving a specified video and its corresponding comment data, and inputting the specified video and its corresponding comment data into the first multi-modal model to obtain an auditing result of the comment data of the specified video, and processing the comment data whose auditing result is negative comment.
[0064] In an optional embodiment of the present application, referring to Figure 3 , in step 200, based on the association between the target video and its corresponding comment data, the comment data corresponding to the target video is labeled, including:
[0065] Extracting key information and emotional information in the comment data corresponding to the target video, and preliminarily determining whether the comment data is a negative comment according to the emotional information;
[0066] In the case of preliminarily determining that the comment data is a negative comment, the comment data is labeled as a negative comment;
[0067] In the case of preliminarily determining that the comment data is a positive comment, based on the association between the target video and its corresponding comment data, further determine whether the comment data is a negative comment;
[0068] In a case where it is further determined that the comment data is a negative comment, the comment data is labeled as a negative comment.
[0069] In a case where it is further determined that the comment data is a positive comment, the comment data is labeled as a positive comment.
[0070] In an optional embodiment of the present application, the extraction of the key information and the sentiment information in the comment data corresponding to the target video can be achieved by a multi-modal large model or a pre-trained second multi-modal model.
[0071] In the present application, in a case where it is preliminarily determined that the comment data is a positive comment, based on the relevance between the target video and the comment data corresponding thereto, it is further determined whether the comment data is a negative comment, which can identify the implicit information in the comment data.
[0072] In an optional embodiment of the present application, the further determination of whether the comment data is a negative comment based on the relevance between the target video and the comment data corresponding thereto includes:
[0073] extracting the key information and the sentiment information in the target video, and calculating the similarity between the key information and the sentiment information in the target video and the comment data corresponding thereto;
[0074] determining the conflict degree between the target video and the text content in the comment data corresponding thereto;
[0075] determining the conflict score between the target video and the comment data corresponding thereto according to the similarity between the key information and the sentiment information in the target video and the comment data corresponding thereto and the conflict degree between the text contents;
[0076] determining the comment data with a conflict score exceeding a preset conflict score threshold as a negative comment.
[0077] In an optional embodiment of the present application, the extraction of the key information and the sentiment information in the target video can be achieved by a multi-modal large model or a pre-trained second multi-modal model. For example, there is a transport vehicle loaded with a car on the left side of the video, which has overturned, causing traffic interruption. The key information includes: snowy day, highway, vehicle queue, traffic accident, and the sentiment information or emotional color is negative.
[0078] In an optional embodiment of the present application, the calculation of the similarity between the key information and the sentiment information in the target video and the comment data corresponding thereto is performed by using cosine similarity and Euclidean distance (which needs to be normalized) algorithms.
[0079] In an optional embodiment of the present application, the similarity between the target video and the key information and sentiment information in the corresponding comment data thereof is calculated by the following expression:
[0080]
[0081] wherein A is a first text vector composed of the key information and sentiment information in the target video, B is a second text vector composed of the key information and sentiment information in the corresponding comment data of the target video, a i is the i-th vector element in the first text vector, b i is the i-th vector element in the second text vector, and w1 and w2 are weight factors.
[0082] In an optional embodiment of the present application, the conflict degree between the target video and the literal content in the corresponding comment data thereof is determined, which is realized by using the natural language processing Bert model, conflict detection is performed by using logical consistency detection and semantic analysis, and a conflict degree score is given. There is no conflict at all, which is 0, and there is complete opposition, which is 1. For example, the viewpoint that the relationship between economic development and environmental protection is dialectically unified and economic development is more important has a conflict degree of about 0.5; advocating "hard work and heavy reading" and "reading is not important" conflict with each other, and the conflict degree is about 0.9.
[0083] In an optional embodiment of the present application, the conflict score between the target video and the corresponding comment data thereof is determined, which is realized by using weighted scoring, the similarity and the conflict degree are respectively assigned weights for weighting, and a comprehensive conflict score is calculated.
[0084] In the present application, since the implicit information is caused by the conflict between the comment content and the video frame content, the extracted video frame content and the comment meaning content are compared. If they are completely opposite or conflict, negative marking is performed according to the text content, and a graph-text data set about negative marking is obtained, that is, the image features of the video are extracted by using the video understanding algorithm; the text features are obtained by analyzing the comment by using the text analysis method; and after the two features are aligned in the feature space, the feature similarity is calculated to evaluate the consistency or opposition of the features; then a rule model and a text analysis model similar to BERT are applied for conflict detection to identify opposite viewpoints and give a conflict degree score; the similarity and the conflict degree are comprehensively evaluated by using the weighted scoring method to evaluate the conflict level, which is divided into three levels of low, medium and high. For the medium and high conflict comments, it is considered that there may be implicit sarcastic meaning, so negative marking is performed. Finally, professional auditors review to ensure data accuracy and quality. After the completion of the labeled data set, it can be used for testing and training of large models in the video comment detection scene. After repeated verification, whether the new data comment has implicit meaning is judged by the inference of the large model, so as to achieve the purpose of comment content review.
[0085] In an optional embodiment of the present application, in step 200, the comment data set corresponding to the target video is constructed, including:
[0086] Each comment data corresponding to the target video and the annotation information of each comment data are collected as the comment data set corresponding to the target video.
[0087] In an optional embodiment of the present application, in step 300, the target video and the comment data set corresponding thereto are used as a training set to fine-tune the preset multi-modal model to obtain the first trained multi-modal model, including:
[0088] The video frame of the target video and each comment data are used as input, the annotation information of the current comment data is used as output, the network structure in the preset multi-modal model for processing the video frame is trained in a supervised task lora fine-tuning manner, the network structure for processing the comment data is trained in a supervised task qlora fine-tuning manner, and the first trained multi-modal model is obtained.
[0089] In an optional embodiment of the present application, for text data such as comment data, the supervised task training fine-tuning adopts a qlora fine-tuning manner, experiments are completed on an Ubuntu 22.04 operating system using an 8-card NVIDIAGPX H800 GPU, all experiments use a Pytorch framework, the batch size is set to 4 during training, and a total of 20 training cycles are performed. The gradient update number of each batch is 1, the learning rate is set to 2e -4 , the learning rate is constant with a warm up strategy, that is, the learning rate is 0 at first, then gradually increases to the set value of the learning rate in the warm-up stage, and the set value of the learning rate is used during formal learning. The maximum length sequence length of the input accepted by the model is 128. For video frame data, the supervised task training fine-tuning adopts a lora fine-tuning manner, experiments are completed on an Ubuntu 22.04 operating system using an 8-card NVIDIAGPX H800 GPU, all experiments use a Pytorch framework, the batch size is set to 4 during training, a total of 20 training cycles are performed, the learning rate is set to 0.00001, the learning rate is a Cosine Annealing LR strategy, and the learning rate is adjusted based on the curve shape of the cosine function. When training starts, the learning rate is large, which can help the model converge quickly. As the training progresses, the learning rate gradually decreases to ensure that the model can accurately search the parameter space. The image resolution of the input accepted by the model is 1024x1024.
[0090] In an optional embodiment of the present application, in step 400, the input of the specified video and its corresponding comment data into the first multi-modal model to obtain the review result of the comment data of the specified video comprises:
[0091] The video frame of the specified video is input into the network structure for processing the video frame in the first multi-modal model to obtain the video frame feature.
[0092] The comment data corresponding to the specified video is input into the network structure for processing the comment data in the first multi-modal model to obtain the comment data feature.
[0093] The video frame feature and the comment data feature are input into the fusion network structure in the first multi-modal model to obtain the review result of the comment data of the specified video.
[0094] In an optional embodiment of the present application, the preset multi-modal model is a pre-trained second multi-modal model or a multi-modal large model. In the case of the preset multi-modal model being the pre-trained second multi-modal model, the target video and its corresponding comment data set are used as a training set to fine-tune the pre-trained second multi-modal model to obtain the trained first multi-modal model. The training and use process of the pre-trained second multi-modal model is as shown in Figure 4 . Referring to Figure 4 , the training and use process of the pre-trained second multi-modal model comprises the following steps: first, constructing a news text data set, the annotation method of the data set is a label annotation method for text; second, using a multi-modal large model to convert the text into a feature vector and learn the joint distribution features of the text vector; third, searching and matching the user input text data, and predicting using the trained multi-modal large model; fourth, reviewing and evaluating the actual text content; fifth, establishing a feedback mechanism of the evaluation result and the artificial review result to continuously optimize the accuracy of the model evaluation.
[0095] In an optional embodiment of the present application, in the first step, the news text data set is constructed, including news text data acquisition and news text data annotation. The social and political news and text with positive and negative information are selected as data sources from the correct value orientation, the news text data set is constructed, and the data set is classified into four scenes: scene one negative article review scene, scene two negative article and comment review scene, scene three ugly character image review scene, and scene four text bad hidden reflection scene. For the news text data set constructed according to the above scenes, the text and image are divided into news text data set and news image data set, and the data set for the four kinds of subdivided scenes is constructed respectively for content review model training.
[0096] For text data, the public social and political news headlines, news content and news comments of authoritative sources are selected as the main data source, the reasons are mainly as follows: 1) social and political news covers positive and negative events, and public news has the advantage of easy access. 2) Social news involves social events, social problems, social features and other aspects of people's daily life, and has the characteristics of authenticity, timeliness and accuracy. For the positive and negative discrimination task in the Chinese context, news texts including domestic and foreign news are collected, and JAVA natural language processing toolkit is used to filter out social news with positive and negative meanings, such as THULAC (THU Lexical Analyzer for Chinese, Chinese lexical analysis toolkit), which can accurately perform Chinese word segmentation and part-of-speech tagging, and can realize user-defined text classification tasks. In the content review task, in order to enable the model to review negative news text, the sampled news text is kept as much as possible to have more negative than positive. The collected news text dataset is used to train the model's ability to distinguish between positive and negative, and the above collected news is filtered to retain news text related to negative information, and text without negative information is considered as positive news text, and text with format errors is deleted. After data processing, 4519 news headlines and news content are used to construct the required news text dataset for scenario one. Table 1 shows an example of scenario one dataset. 4627 news headlines, content and comments are used to construct the required news text dataset for scenario two. Table 2 shows an example of scenario two.
[0097] Table 1
[0098]
[0099]
[0100] Table 2
[0101]
[0102]
[0103] For image data, the negative of the image generally appears in the public figure image tampering, and the positive image of the figure is used for satire and implicit scene. Therefore, the selected images are mainly selected in the social public figures. Because the number of positive images of the figure image is more than that of negative images in China, leading to the problem of sample imbalance in automatically collected image data, in order to solve this problem, the main figure images are further crawled to expand the negative image data.
[0104] The news image dataset constructed above is used to train the ability of positive and negative discrimination of the model, therefore, the collected images are screened, the images and related image comments containing positive and negative information are retained, and the neutral images not containing positive and negative information, the images with format errors and the like are deleted, after data processing, 9243 images are used to construct the required news image dataset of scene three. Table 3 is an example of scene three dataset. 4252 images and image comments are used to construct the required news image dataset of scene four. Table 4 is an example of scene four.
[0105] Table 3
[0106]
[0107] Table 4
[0108]
[0109]
[0110] In the process of labeling, positive data refers to positive and positive reports and comments formed on the Internet and social media. Negative data opposite to positive data contains negative comments, malicious criticism and even defamatory reports. The labeling of the review scene is mainly aimed at negative data, therefore, the negative data is mainly labeled, and the rest of the data is positive data, which does not contain neutral data. Specifically, four scenes are labeled, for positive scenes, labeled as positive; negative scenes, labeled as negative, and negative reason analysis. According to the definition of positive and negative, the data in the scene is labeled as 1 (positive), 2 (negative and negative reason).
[0111] In an optional embodiment of the present application, in the second step, the four scenes are fine-tuned and trained using a multi-modal large model and a news image dataset to learn the joint distribution of image and text vectors, and a content review large model is obtained. Referring to Figure 5 , mainly using the large language model of ChatGLM (a generative language model) and the multi-modal large model of CogVLM (an open source visual language base model), the input of the model is two categories of text and image, scene one as shown in ①, input text content; scene two as shown in ②, input text and text comment content; scene three as shown in ③, input image; scene four as shown in ④, input image and comment content about the image. GLM mainly uses the decoder part of the Transformer (a neural network model based on self-attention mechanism), in natural language understanding and natural language generation tasks, GLM uses the generation principle for reasoning. The main principle is to generate possible fill-in content according to known part of text content, GLM can be used in automatic text completion, question and answer system, semantic understanding and generation and other natural language processing tasks. For example,Figure 6 shown, is the principle part of text content understanding. After the text content is input, the word embedding is formed to form the text features. Assuming that a given text content t = {t1,..., tn} is input, the original text is segmented by word, and multiple texts x = [x1,..., xn] are selected, each x represents a continuous text token (the smallest text unit), and the word embedding output with semantic features is represented. The GLM pre-training structure is called autoregressive blank filling. The original text token is [x1,..., xn], and the original text is randomly sampled, assuming that [x3], [x5, x6] are sampled texts. The sampled part of the original text is replaced with [M] as Part A, and the GLM generates Part B part autoregressively. Each sampled text is added with [S] as input before it, and [E] is added after it to represent the output. Part A and Part B are added with two-dimensional position encoding, in which position encoding Pos1 represents the position of the original text, and position encoding Pos2 represents the position of the sampled text itself. After adding the position encoding, the input is input to the sub-attention mask part of the GLM. Each text in the Part A part can be noticed, and the text in the Part B part can be noticed in the Part A and Part B that have been generated. Finally, the output text is obtained, and the output text is as close to the original text as possible. The GLM understands the text meaning in this way of autoregressive blank filling. After understanding the text meaning, the model can be fine-tuned in combination with specific tasks. In the text review scene, the text content meaning is understood, and the detailed text content of the text scene is judged to be positive or negative. In the natural language understanding task of GLM fine-tuning, a labeled example (x, y) is given, and the input text x is converted into a cloze problem by including a single mask label. The conditional probability of predicting y given x is given. Then the positive and negative labels are mapped to the words good (good) and bad (bad), and the cross-entropy loss function is used for back propagation to achieve fine-tuning. In this application, given a text content x, the labeled positive and negative meanings are used as y labels, and the model uses the autoregressive blank filling method to generate the self-judged positive and negative meanings every time. The CogVLM large model uses the task form of visual question answering, which consists of four basic components: a visual encoder ViT (Vision Transformer, visual transformer) encoder, an MLP (Multilayer Perceptron, multilayer perceptron) adapter, a pre-trained large language model LLM, and a visual specialist module. As shown in Figure 7The understanding principle of image and question-answer text is shown. Specifically, first, the text question-answer description about the image is converted into text features through word embedding; the image is converted into feature representation through the ViT encoder, and then mapped into the same space as the text features through the MLP adapter. The two features are combined through concat (a function for concatenating multiple strings) and combined with position coding input into a pre-trained large language model structure, which is mainly based on the Vicuna-7B model, which is a pre-training architecture of GPT. Each layer in the structure adds a visual expert module to realize the deep alignment of vision and language, so as to realize the deep understanding of the model to the image data. Finally, similar to the text review scene, the image review scene also needs to combine specific tasks to fine-tune the model. In the process of training the news image-text data set of the large model, different question-asking methods are adopted as prompt words to obtain the review results. The question-asking methods are as follows, Scene one question-asking method: please judge whether the following text contains negative comments. Negative comments usually involve derogatory or misleading comments, which can cause social dissatisfaction or damage social stability. If the text contains negative comments, please return "negative" and give the reason, otherwise return "positive". Scene two question-asking method: please judge whether the following comment contains negative comments according to the text title and content. Negative comments usually involve derogatory or misleading comments, which can cause social dissatisfaction or damage social stability. If the article content contains negative information or the comment contains negative comments, please return "negative" and give the reason, otherwise return "positive". Scene three question-asking method: does the picture belong to the ugly image of the person? If the picture belongs to the ugly image of the person, please return "negative" and give the reason, otherwise return "positive". Scene four question-asking method: does the picture and comment involve sensitive comments? If the picture and comment involve sensitive comments, please return "negative" and give the reason, otherwise return "positive".
[0112] During fine-tuning training, each time the text and image to be detected are input, the system will automatically combine the question mode of each scene, input to the model, and obtain the final review result. In the text scene, the GLM model is used to understand the meaning of the text by filling in the blanks in an autoregressive manner to determine whether the text content is positive or negative. In the image review scene, the understanding principle of image and text is used in CogVLM. First, the text question and answer description about the image are converted into text features through word embedding; the image is converted into feature representation through the ViT encoder, and then mapped into the same space as the text features through the MLP adapter. The two features are combined through concat and combined with position encoding to input into the pre-trained large language model structure. This structure is mainly based on the Vicuna-7B model, which is a pre-training architecture of GPT. Each layer in this structure adds a visual expert module to achieve deep alignment of vision and language, thereby achieving deep understanding of image data by the model. Finally, similar to the text review scene, the image review scene also needs to combine specific tasks to fine-tune the model.
[0113] In an optional embodiment of the present application, in the third step, the trained large model is used for result prediction. A micro-service is established through the large model, and a content review system based on a multi-modal large model is formed. The architecture of the content review system based on the multi-modal large model includes hardware, software, databases, and functional modules. The system hardware is a server with high computing power. The system database includes a text database and an image database for storing the uploaded content for review. The system functional modules include a system data conversion module, a content review algorithm packaging module, a system output module, an unstructured data processing module, a multi-modal data integration module, a content analysis module, a content anomaly quantification module, and an intelligent control module.
[0114] The data conversion module is used to judge the type of data input by the user into the system, determine the scene category from the input data, combine the different question modes of each scene after preprocessing the data, and then transmit it to the multi-modal large model for reasoning;
[0115] The unstructured data processing module is used to use the multi-modal large model to perform semantic analysis on unstructured data such as news text, audio, and video, generate a semantic association graph, extract the core context and information correlation in the text and image through deep analysis, and build a preliminary review dataset;
[0116] The multi-modal data integration module is used to integrate the text data, audio data, and video data in a unified context, convert them into a text association vector space, and transmit them to the content analysis module for deep analysis of the content context and emotion;
[0117] The content analysis module is configured to preliminarily determine the matching degree of the text and the image content according to the semantic correlation graph generated by the unstructured data processing module, generate a context factor Qycs, and analyze the potential sensitive content in each modality data to generate a sensitive factor Gmyz; then, according to the context factor Qycs and the sensitive factor Gmyz, the potential emotional factors, the social sensitive information and the potential misleading factors in the image-text content are hierarchically analyzed to generate a multi-modal correlation factor Mmgyz.
[0118] The content anomaly quantification module is configured to calculate an audit weight value Hqwz according to the context factor Qycs, the sensitive factor Gmyz and the multi-modal correlation factor Mmgyz; then, by calculating the system audit risk factor Rxz and associating the audit weight value Hqwz, a final content anomaly index Cxzs is generated.
[0119] The calculation formula of the audit weight value Hqwz is:
[0120] Hqwz = b1 x Qycs + b2 x Gmyz + b3 x Mmgyz + B.
[0121] Wherein, b1, b2 and b3 represent the preset weight coefficients of the context factor Qycs, the sensitive factor Gmyz and the multi-modal correlation factor Mmgyz respectively, and B represents an audit correction coefficient.
[0122] The intelligent control module is configured to set a compliance threshold T and an audit risk threshold R, and compare the audit weight value Hqwz with the compliance threshold T to determine whether the current content meets the audit standard, and compare the audit risk threshold R with the content anomaly index Cxzs to obtain the final audit result, and selectively trigger the control scheme of manual review or automatic release according to the audit result.
[0123] In an optional embodiment of the present application, the unstructured data processing module is configured to first receive the unstructured data of the news image-text, audio and video to be audited, and perform format standardization processing on the unstructured data, including image resolution adjustment, text language normalization processing and audio-video transcoding.
[0124] Then, the pre-training network of the multi-modal large model is called to extract multi-level features including visual features, semantic features and emotional features from the news image-text, audio and video respectively, and the extracted multi-level features are normalized to generate a single-modal feature matrix; based on the single-modal feature matrix, the multi-level features of the news image-text, audio and video are mapped in a multi-modal correlation vector space, and a cross-modal semantic correlation graph is generated through the semantic embedding mechanism of the multi-modal large model.
[0125] Secondly, the cross-modal semantic association graph is used to analyze the deep association between text and image, extract the core context elements in text and image, and synchronously analyze the emotional state and expression trend in audio and video, so as to identify the potential sensitive information, misleading information and social influence factors in news content in real time, and form a multi-modal association semantic index.
[0126] Finally, the preliminary review data set is generated.
[0127] In an optional embodiment of the present application, the multi-modal data integration module is used to receive the preliminary review data set generated by the unstructured data processing module, and perform time sequence and content unified alignment processing on the text and image data, audio data and video data therein through frame synchronization algorithm;
[0128] Then, the multi-modal feature mapping algorithm is applied to fuse the context, emotion and semantic features of text, audio and video into multi-modal feature vectors, and generate an association degree matrix according to the association degree between each mode, so as to perform hierarchical mapping on information density, context relevance and semantic consistency;
[0129] Subsequently, the text and image association vector space is constructed based on the association degree matrix, and the multi-modal data is quantized in the unified semantic space to generate an integrated semantic association representation vector, and the semantic deviation caused by the mode difference is corrected through the context adjustment algorithm;
[0130] Finally, the integrated text and image association vector space is transmitted to the content analysis module.
[0131] In an optional embodiment of the present application, the content analysis module includes a preliminary discrimination unit, a multi-layer sensitivity analysis unit and an association factor generation unit.
[0132] The preliminary discrimination unit identifies semantic deviation related data and content matching related data by analyzing the semantic similarity and context consistency between text and image; extracts the semantic similarity Ycy, context consistency Sjy, context association deviation degree Qjl and text and image emotion consistency Twq in the semantic deviation related data and content matching related data, and performs dimensionless processing, and then calculates the context factor Qycs through the following formula:
[0133]
[0134] The context threshold Q is compared with the context factor Qycs to evaluate the context association of text and image in the review scene, and the specific comparison and evaluation content is as follows:
[0135] If the context factor Qycs is greater than or equal to the context threshold Q, it indicates that the context matching between text and image is consistent, and the semantic, emotion and context of the two are consistent, and the review is passed and marked as "context association normal";
[0136] If the context factor Qycs is less than the context threshold Q, it indicates that the context between the text and the image is not matched, and the semantics, emotions and context of the two are not consistent, the review is not passed and the specific deviation reason is further analyzed, including potential sensitive content, wherein the semantic similarity Ycy is used to measure the similarity of the content expressed by the text and the image in the semantic level, and the text vector representation is obtained by inputting the text into the pre-trained multi-modal model; the context consistency Sjy mainly measures the consistency degree of multi-dimensional information such as time, place, person or event between the text context and the image context, and is obtained by comparing the context feature set obtained from the text with the context feature set recognized in the image or video; the context correlation deviation degree Qjl is used to quantify the potential context mismatching degree between the text and the image or video, and is obtained by identifying the inconsistent part of the "core scene or main object" after aligning the text and the image, and then counting the inconsistent categories or quantities; the text-image emotion consistency Twq is used to measure the consistency degree between the emotion conveyed by the text and the emotion conveyed by the image, and is obtained by calculating the similarity of the text emotion distribution and the image emotion distribution.
[0137] In an optional embodiment of the present application, the multi-layer sensitivity analysis unit is further used for in-depth analysis of the potential sensitive content of the image-text data, including analyzing the emotion-related data, social sensitivity-related data and potential misleading-related data in the text and the image, and constructing a sensitive content data set after dimensionless processing of the analyzed emotion-related data, social sensitivity-related data and potential misleading-related data in the text and the image; and extracting the sensitive content data set to generate a sensitive factor Gmyz through the following formula:
[0138]
[0139] In the formula, Qmq represents the emotion intensity coefficient in the sensitive content data set, Mgc represents the sensitive word index in the sensitive content data set, and Xwg represents the information misleading possibility in the sensitive content data set, wherein the emotion intensity coefficient Qmq is used to reflect the emotion or emotion intensity in the news content, and is obtained by text emotion intensity analysis; the sensitive word index Mgc is mainly used to measure the frequency and intensity of the appearance of sensitive information, sensitive terms, prohibited words or sensitive symbols in the text, and is obtained by sensitive dictionary and rule matching; and the information misleading possibility Xwg is used to evaluate whether there is a potential tendency to "confuse the audience", "fake information" and "mislead the public" in the news content, and is obtained by cross-modal comparison.
[0140] In an optional embodiment of the present application, the correlation factor generation unit is configured to quantify the correlation degree between the sentiment, context and sensitive content by fusing multi-level information in the image-text correlation vector space based on the context factor Qycs and the sensitive factor Gmyz, and to generate the multi-modal correlation factor Mmgyz according to the following formula:
[0141]
[0142] In an optional embodiment of the present application, the content anomaly quantification module comprises an audit calculation unit and an anomaly index acquisition unit.
[0143] The content evaluation unit is configured to construct a system audit risk factor Rxz, to obtain a semantic deviation degree Ycp and a content consistency index Nry by extracting semantic deviation related data and content matching related data, and to obtain a high sensitivity trigger rate Gmg by extracting a sensitive content data set, and to calculate the system audit risk factor Rxz according to the following formula:
[0144]
[0145] In an optional embodiment of the present application, the anomaly index acquisition unit is configured to calculate the final content anomaly index Cxzs according to the following formula:
[0146]
[0147] In an optional embodiment of the present application, the intelligent control module compares and evaluates the audit weight value Hqwz and the compliance threshold T respectively, and compares and evaluates the audit risk threshold R and the content anomaly index Cxzs, and specifically generates the following evaluation contents:
[0148] Compliance standard comparison:
[0149] If the audit weight value Hqwz is greater than or equal to the compliance threshold T, it indicates that the content meets the audit standard, and the system is marked as "compliant",
[0150] If the audit weight value Hqwz is less than the compliance threshold T, it indicates that the content does not meet the audit standard, and there is a compliance deficiency, which needs to be further reviewed or adjusted.
[0151] Abnormal risk assessment:
[0152] If the content anomaly index Cxzs is greater than or equal to the audit risk threshold R, it indicates that the content has an abnormal risk, and the system generates an "abnormal content" mark to trigger an artificial review process.
[0153] If the content anomaly index Cxzs is less than the audit risk threshold R, it indicates that the content does not have an abnormal risk, and the system generates a "normal content" mark, and enters an automatic publishing process.
[0154] When the audit weight value Hqwz and the evaluation result of the content anomaly index Cxzs obtain the labels of "compliance" and "normal content" at the same time, the content is automatically labeled as "pass" and directly published;
[0155] When the audit weight value Hqwz and the evaluation result of the content anomaly index Cxzs are content that does not meet the audit standard or has abnormal risk, the content is automatically labeled as "to be reviewed", and a selective review process is triggered.
[0156] A control scheme for setting the compliance threshold T and the audit risk threshold R, judging whether the current content meets the audit standard, obtaining the final audit result, and selectively triggering manual review or automatic publishing according to the audit result.
[0157] In an optional embodiment of the present application, a content audit algorithm packaging module is used to package all algorithms in the system.
[0158] In an optional embodiment of the present application, a system output module is used to extract key information of the result according to the result obtained by the large model, and output the final content audit category.
[0159] In an optional embodiment of the present application, the first sub-model of the first multi-modal model is used to receive and obtain the audit result of the specified video according to the specified video and its corresponding comment data.
[0160] In an optional embodiment of the present application, the second sub-model of the first multi-modal model is also used to receive and determine the target application system failure scenario according to the target application system failure graph.
[0161] In an optional embodiment of the present application, the target application system failure scenario is used to determine the query prompt word of the target application system failure graph, extract the target application system failure graph text according to the query prompt word, and decompose the target application system failure text description and the target application system failure graph text into multiple sub-tasks based on the agent and the large language model, and determine the target application system failure factor according to the information returned by the sub-tasks.
[0162] In an optional embodiment of the present application, the second sub-model of the first multi-modal model is obtained by the following steps:
[0163] The historical data of the application system failure graph under each application system failure scenario is collected respectively, wherein the application system failure includes video playback error, sequence error, data loss and content error;
[0164] The neural network model is trained by taking the application system failure graph under each application system failure scenario as input and taking the application system failure scenario type as output, to obtain a second sub-model of the first multi-modal model.
[0165] In an optional embodiment of the present application, the target application system failure scenario is used to determine a query prompt word of the target application system failure graph, including:
[0166] Based on the preset correspondence between the application system failure scenario and the query prompt word, the query prompt word corresponding to the target application system failure scenario is determined.
[0167] The query prompt word is taken as the query prompt word of the target application system failure graph.
[0168] In an optional embodiment of the present application, the extraction of the target application system failure graph text according to the query prompt word includes:
[0169] The query prompt word and the target application system failure graph are input into a preset multi-modal large model.
[0170] The text extracted from the target application system failure graph according to the query prompt word is output from the preset multi-modal large model, and is taken as the target application system failure graph text.
[0171] In an optional embodiment of the present application, the target application system failure text description and the target application system failure graph text are decomposed into a plurality of sub-tasks based on the agent and the large language model, and the target application system failure factor is determined according to the information returned by the sub-tasks, including:
[0172] The target application system failure text description and the target application system failure graph text are input into a scheduling agent.
[0173] The scheduling agent determines the failure processing state according to the target application system failure text description and the target application system failure graph text, and sends the target application system failure text description and the target application system failure graph text to a task decomposition agent in the case that the failure is not processed.
[0174] The task decomposition agent decomposes the failure into a plurality of sub-tasks according to a preset prompt word template based on the functions of the search engine, the interface and the database in the preset knowledge base, and sends the plurality of sub-tasks to a task execution agent, wherein each sub-task corresponds to at least one of the search engine, the interface and the database in the preset knowledge base.
[0175] The task execution agent generates an executable program for each sub-task according to the prompt word template through a preset large language model.
[0176] The task execution intelligent agent executes each executable program by using at least one of the search engine, the interface and the database according to the corresponding relationship between each executable program and the search engine, the interface and the database respectively, and takes the execution result of each executable program as a fault factor of the target application system.
[0177] In an optional embodiment of the present application, before taking the execution result of each executable program as a fault factor of the target application system, the method further comprises:
[0178] sending the execution result of each executable program to the task decomposition intelligent agent;
[0179] The task decomposition intelligent agent determines whether the execution result of each executable program meets a preset condition, and adjusts a preset prompt word template in a case where the execution result of each executable program does not meet the preset condition, and re-executes a plurality of executable programs by decomposing the fault into a plurality of sub-tasks according to the adjusted prompt word template until the execution result of each executable program meets the preset condition.
[0180] In an optional embodiment of the present application, the third sub-model of the first multi-modal model is further configured to determine whether the specified file is a target file to be deleted according to file title, time and text information of the specified file.
[0181] In an optional embodiment of the present application, the third sub-model of the first multi-modal model is trained by the following steps:
[0182] The third sub-model of the first multi-modal model is trained by taking file title, time and text information of the historical file as input and taking whether the file is a target file to be deleted as output.
[0183] In an optional embodiment of the present application, the file title, time and text information of the specified file are obtained by the following steps:
[0184] The file transmission module is used to receive the target file, the content distribution network module is used to send the received target file, and the information of the target file is sent to the message server;
[0185] The file parsing module is used to obtain the information of the target file on the message server, obtain the target file, parse the target file, and push the parsed file title, time and text information to the index module.
[0186] In an optional embodiment of the present application, the file parsing module is used to obtain the information of the target file on the message server, obtain the target file, parse the target file, and push the parsed file title, time and text information to the index module, which comprises:
[0187] obtaining the file type of the target file;
[0188] determining whether the target file needs to be preprocessed according to a file type of the target file;
[0189] In the case that the target file needs to be preprocessed, preprocessing the target file, and then parsing the preprocessed target file;
[0190] In the case that the target file does not need to be preprocessed, directly parsing the target file;
[0191] pushing the parsed file title, time and text information to an index module.
[0192] In an optional embodiment of the present application, determining whether the target file needs to be preprocessed according to a file type of the target file, comprising:
[0193] In the case that the file type of the target file is a text file, the target file does not need to be preprocessed;
[0194] In the case that the file type of the target file is a JSON file or an HTML file, the target file needs to be preprocessed.
[0195] In an optional embodiment of the present application, in the case that the target file needs to be preprocessed, preprocessing the target file, comprising:
[0196] In the case that the file type of the target file is a JSON file, deleting a JSON label in the JSON file;
[0197] In the case that the file type of the target file is an HTML file, in the case that the HTML file contains at least one of audio and video, obtaining a globally unique identifier of the audio and video, obtaining recognition text of the audio and video in the HTML file through the globally unique identifier of the audio and video, and then taking the recognition text as the HTML file, extracting a maximum text block of the HTML file; in the case that the HTML file does not contain at least one of audio and video, extracting a maximum text block of the HTML file.
[0198] In an optional embodiment of the present application, extracting a maximum text block of the HTML file, comprising:
[0199] parsing a label in the HTML file into a DOM tree, and obtaining a total number of labels of the HTML file and position information of each label;
[0200] sorting all nodes according to a number of line break labels and a number of paragraph labels in each node of the HTML file to obtain nodes ranked in front N;
[0201] For each node in the top N, all the nodes in the top N are sorted again according to the link quantity, the Chinese punctuation mark quantity, the total word quantity and the link-bearing word quantity of the current node, the total tag quantity of the HTML file and the position information of each tag, and the node in the first place is taken as the maximum body text block;
[0202] The title of the HTML file is extracted from the node position of the maximum body text block, the publishing time of the HTML file is extracted through a regular expression, in the case of successful extraction, the body of the node in the first place is taken as the body of the HTML file, in the case of unsuccessful extraction, the node in the second place in the secondary sorting is taken as the maximum body text block, the title and the publishing time of the HTML file are extracted until successful extraction is achieved;
[0203] In the case of taking the node in the first place as the maximum body text block, if the body of the node in the first place cannot be extracted or the extracted body word quantity is less than a preset threshold, the html tag in the HTML file is directly deleted, and the HTML file with the deleted html tag is taken as the body of the HTML file.
[0204] In an optional embodiment of the present application, all the nodes are sorted according to the line break tag quantity and the paragraph tag quantity in each node on the DOM tree, and the nodes in the top N are obtained, including:
[0205] The first weight value of each node is determined through the following expression:
[0206] W1 = brCount + 2 * pCount
[0207] wherein W1 is the first weight value of each node, brCount is the line break tag quantity of each node, pCount is the paragraph tag quantity of each node,
[0208] All the nodes are sorted according to the first weight value of all the nodes from large to small, and the nodes in the top N are obtained.
[0209] In an optional embodiment of the present application, for each node in the top N, all the nodes in the top N are sorted again according to the link quantity, the Chinese punctuation mark quantity, the total word quantity and the link-bearing word quantity of the current node, the total tag quantity of the HTML file and the position information of each tag, including:
[0210] The second weight value of each node is determined through the following expression:
[0211] W2 = (WN-LWL) * dl * f
[0212] wherein W2 is a second weight value of each node, WN is a total word number of each node, LWL is a word number with link of each node, dl is position information of each node, and f is a weight factor of each node,
[0213] performing secondary sorting on all nodes ranked in the front N in descending order of the second weight values of all nodes ranked in the front N,
[0214] wherein the position information of each node is obtained by the following expression:
[0215]
[0216] wherein dl is the position information of each node, NodeCount and NodePos are respectively a total number of tags of the HTML file and position information of each tag,
[0217] the weight factor of each node is obtained by the following expression:
[0218]
[0219] wherein f is the weight factor of each node, dLinkPower is a link number of each node, and nCDotNum is a Chinese punctuation symbol number of each node.
[0220] In an optional embodiment of the present application, the functions implemented by the system include three aspects of the main functions of the system, namely, text content review, picture content review, and history record management. The text content review is mainly divided into article content review and article comment content review. In the main function page of the system, the content display area on the right side of the text screening page is the overall content. The text screening page includes two parts, the left side is the content input area, including two switching pages of article screening and comment screening, and the right side is the screening result display area. After inputting the article title and article content, clicking the "screening" button, the article can be screened. The screening result will be displayed in the "screening result" area on the right. If the user does not agree with the review result, the "manual review" button can be clicked. After clicking, a pop-up window will be popped up, and the platform analysis can be corrected by inputting the review opinion. The user agrees with the review result, and clicks OK. Clicking Cancel will directly close the pop-up window. The picture content review is similar to the text content review. The picture content review is mainly divided into ugly image of a person and bad hidden reflection of text and image. The text and image screening page includes two parts, the left side is the content upload and input area, including two switching pages of ugly image of a leader and bad hidden reflection of text and image, and the right side is the screening result display area. In the picture screening page, the "upload picture" button is clicked to select the picture to be screened. The picture area will display the picture content screen to be screened. Clicking the "screening" button, the picture can be screened. The screening result will be displayed in the "screening result" area on the right. If the user does not agree with the review result, the "manual review" button can be clicked. The history record management includes text review and picture review content. The risk screening record page is in the form of a table. The table header has screening date, material content, material type, scene classification, screening result, manual screening result, and operation. Clicking the date, material type, screening result, and manual screening result of the table can be filtered. Clicking the "edit" button can enter the manual review pop-up window page, and clicking the "delete" button can delete the screening record. The record page has date, material content, material type, abstract, keyword, and label. Clicking the date and material type of the table can be filtered. The history management is provided. The user can view all the uploaded historical images and texts, view the content that has been uploaded, and avoid repetitive processing of the same file.
[0221] In an optional embodiment of the present application, in the fourth step, the prediction result is manually reviewed and evaluated. An audit feedback mechanism is established. For the prediction (screening) result that is not satisfied, the manual review button can be clicked to re-evaluate the result. The background will record this result, re-label the data set, and join the fine-tuning of the model.
[0222] In an optional embodiment of the present application, in the fifth step, a feedback mechanism is established between the evaluation results and the manual review results to continuously optimize the accuracy of model evaluation. The data with inconsistent results in the historical records will be re-labeled and put into the model for fine-tuning to continuously improve the accuracy of model evaluation.
[0223] The present application provides a content review method based on a multi-modal large model. The method addresses the four scenarios that need to be reviewed in the news field, i.e., negative article review scenario, negative article comment review scenario, ugly character image review scenario, and text and image bad hidden reflection review scenario, solves the problem of negative information that cannot be identified by traditional intelligent technology, improves the accuracy of content review, realizes the automation of text and image review, detects multi-modal review problems that cannot be solved by traditional intelligent review through a large model, and improves the review efficiency.
[0224] In an optional embodiment of the present application, in the model training in the second step, for text data, the data magnitudes of scene one and scene two are consistent according to the experimental parameter settings, and there is no need to consider the mixed matching problem of different task data, so all the data are used for supervised training and fine-tuning combined with the questioning methods of different scenes. For image data, since the data of scene four is less than that of scene three, the data set of scene four is fixed, and the data amount of data three is increased. When the data matching ratio of scene three to scene four is 2:1, the model performance is best, so the data mixed matching ratio of image scene is 2:1, and the training and fine-tuning are combined with the questioning methods of different scenes. There are mainly three training strategies, i.e., multi-task mixed training, multi-task sequential training, and mixed general corpus combined with data set training. In the experimental process, sequential learning in the experiments of the two scenes will cause learning forgetting, and mixing in general corpus will lose general ability, thereby causing catastrophic forgetting. However, training combined with general corpus can alleviate the occurrence of the problem. Therefore, the general corpus CogVLM-SFT-311K bilingual visual instruction data set is combined with the data set of specific scenes for fine-tuning. The trained model will make result prediction on the validation set. The validation set has a total of 910 text data and 1350 image data.
[0225] The experimental results first need to calculate the TP, TN, FP, and FN of each scene; wherein TP represents the total number of accurately detected negative information contained in the data; TN represents the total number of accurately detected positive information; FP represents the total number of false detection of negative information as positive information; and FN represents the total number of false detection of positive information as negative information. Then the accuracy, precision, recall, and F1 of each scene can be calculated according to the following formula, and the calculation results are shown in Table 5:
[0226]
[0227] Table 5
[0228] Scenario Accuracy Precision Recall [F1] 1 93.8% 94.9% 92.5% 93.1% 2 94.2% 97.8% 93.2% 93.7% 3 97.9% 99.3% 97.3% 97.6% 4 96.8% 100% 96.1% 96.4%
[0229] From Table 5, the scene accuracy, precision and recall rate are more than 90%. F1 is more than 90%, which indicates that the experimental method is ideal. Among them, the picture-text scene is more difficult in the overall task, and there is still room for further improvement. But the whole basically meets the needs of the auditors, and can reduce the auditing burden of the auditors in practical application.
[0230] It should be understood that although each step in the flowchart is shown in sequence according to the direction of the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless explicitly stated in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least part of the steps in the figure can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages is not necessarily sequential, but can be executed in rotation or alternation with at least part of other steps or sub-steps or stages of other steps.
[0231] See Figure 8 An embodiment of the present application provides a video comment auditing device 800, which comprises:
[0232] An acquisition module 810 is configured to acquire a target video and corresponding comment data of the target video.
[0233] A construction module 820 is configured to label the comment data corresponding to the target video based on the association between the target video and the comment data corresponding to the target video, and construct a comment data set corresponding to the target video.
[0234] A training module 830 is configured to use the target video and the comment data set corresponding to the target video as a training set to fine-tune a preset multi-modal model, and obtain a trained first multi-modal model.
[0235] An input module 840 is configured to receive a specified video and corresponding comment data of the specified video, input the specified video and the corresponding comment data into the first multi-modal model, obtain an auditing result of the comment data of the specified video, and process the comment data with a negative comment.
[0236] Specific limitations on the apparatus 800 can be found in the above description of the method for video comment review, which will not be repeated here. Each module in the apparatus 800 can be implemented by software, hardware, and a combination thereof in whole or in part. Each module can be embedded in or independent of a processor in the computer device in hardware form, or stored in a memory in the computer device in software form, so as to be called and executed by the processor to perform the operations corresponding to each module.
[0237] In one embodiment, a computer device is provided, and an internal structure diagram of the computer device can be as shown in Figure 9 The computer device includes a processor, a memory, a network interface, and a database connected by a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The database of the computer device is configured to store data. The network interface of the computer device is configured to communicate with an external terminal through a network connection. The computer program is executed by the processor to implement the above method for video comment review. The computer device includes a memory and a processor. The memory stores a computer program, and the processor implements any step in the above method for video comment review when executing the computer program.
[0238] In one embodiment, a computer readable storage medium is provided, and the computer readable storage medium stores a computer program. The computer program is executed by a processor to implement any step in the above method for video comment review.
[0239] Those skilled in the art will appreciate that embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage media, etc.) containing computer usable program code.
[0240] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure One one or more flowcharts and / or blocks Figure One means for functionally implementing the steps listed in the flowchart block or blocks.
[0241] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the function specified in the flowchart block or blocks. Figure One one or more flowcharts and / or blocks Figure One means for functionally implementing the steps listed in the flowchart block or blocks.
[0242] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure One one or more flowcharts and / or blocks Figure One means for functionally implementing the steps listed in the flowchart block or blocks.
[0243] While the preferred embodiments of the application have been described, additional variations and modifications can be employed by those skilled in the art. Therefore, the appended claims intend to cover all such modifications and variations as fall within the true spirit and scope of the application.
[0244] Obviously, numerous modifications and variations of the present application are possible in light of the above teachings. It is therefore to be understood that within the scope of the appended claims and their equivalents, the application can be practiced otherwise than as specifically described.
Claims
1. A video comment review method, characterized in that: include: Get the target video and its corresponding comment data; Based on the correlation between the target video and its corresponding comment data, the comment data corresponding to the target video is annotated to construct a comment dataset corresponding to the target video; The target video and its corresponding comment dataset are used as a training set to fine-tune the preset multimodal model to obtain a trained first multimodal model; Receive a specified video and its corresponding comment data, input the specified video and its corresponding comment data into a first multimodal model, obtain a review result of the comment data of the specified video, and process the comment data with a negative review result, The preset multimodal model is a pre-trained second multimodal model. The training and use process of the pre-trained second multimodal model includes the following steps: the first step is to construct a news image and text data set, and the data set is annotated in the way of adding labels to the image and text; the second step is to use the multimodal large model to convert the image and text into feature vectors and train and learn the joint distribution characteristics of the image and text vectors; the third step is to search and match the image and text data input by the user, and use the trained multimodal large model for prediction; the fourth step is to review and evaluate the actual image and text content; the fifth step is to establish a feedback mechanism between the evaluation results and the manual review results, and continuously optimize the accuracy of the model evaluation. In the third step, the trained big model is used to predict the results, microservices are established through the big model, and a content review system based on the multimodal big model is formed. The content review system architecture based on the multimodal big model includes hardware, software, database and functional modules. The system functional modules include system data conversion module, content review algorithm encapsulation module, system output module, unstructured data processing module, multimodal data integration module, content analysis module, content anomaly quantification module and intelligent control module. The content analysis module is used to preliminarily determine the degree of match between text and image content based on the semantic association graph generated by the unstructured data processing module, generate a contextual factor Qycs, and analyze the potential sensitive content in each modal data to generate a sensitivity factor Gmyz; then, based on the contextual factor Qycs and the sensitivity factor Gmyz, perform a hierarchical analysis of the potential emotional factors, socially sensitive information, and potential misleading factors in the image and text content to generate a multimodal association factor Mmgyz; The content anomaly quantification module calculates the audit weight value Hqwz based on the context factor Qycs, the sensitivity factor Gmyz, and the multimodal correlation factor Mmgyz; then obtains the system audit risk factor Rxz by calculation and associates it with the audit weight value Hqwz, thereby generating the final content anomaly index Cxzs; The calculation formula for the audit weight value Hqwz is: ; Among them, b1, b2 and b3 represent the preset weight coefficients of the situational factor Qycs, the sensitive factor Gmyz and the multimodal correlation factor Mmgyz respectively, and B represents the audit correction coefficient; The intelligent control module is used to set the compliance threshold T and the audit risk threshold R, and compare the audit weight value Hqwz with the compliance threshold T to determine whether the current content meets the audit standards, and compare the audit risk threshold R with the content anomaly index Cxzs to obtain the final audit result, and selectively trigger manual review or automatic release control plans based on the audit results.
2. The method according to claim 1, characterized in that The tagging of the comment data corresponding to the target video based on the correlation between the target video and the comment data corresponding to the target video includes: Extract key information and sentiment information from the comment data corresponding to the target video, and preliminarily determine whether the comment data is a negative comment based on the sentiment information; If it is preliminarily determined that the review data is a negative review, marking the review data as a negative review; If it is preliminarily determined that the comment data is a positive comment, further determining whether the comment data is a negative comment based on the correlation between the target video and its corresponding comment data; If it is further determined that the comment data is a negative comment, marking the comment data as a negative comment; If it is further determined that the comment data is a positive comment, the comment data is marked as a positive comment.
3. The method according to claim 2, characterized in that The further determining whether the comment data is a negative comment based on the correlation between the target video and its corresponding comment data includes: Extract key information and sentiment information from the target video, and calculate the similarity between the key information and sentiment information in the target video and its corresponding comment data; Determining the degree of conflict between the target video and the textual content in its corresponding comment data; Determine the conflict score between the target video and its corresponding comment data based on the similarity between the key information and emotional information in the target video and its corresponding comment data, as well as the degree of conflict between the text contents; Review data whose conflict score exceeds a preset conflict score threshold is determined as a negative review.
4. The method according to claim 3, characterized in that The similarity between the key information and sentiment information in the target video and its corresponding comment data is calculated using the following expression: ; Among them, A is the first text vector composed of key information and emotional information in the target video, B is the second text vector composed of key information and emotional information in the comment data corresponding to the target video, a i is the i-th vector element in the first text vector, b i is the i-th vector element in the second text vector, w1 and w2 are both weight factors.
5. The method according to claim 1, wherein The step of constructing a comment dataset corresponding to the target video includes: Each comment data corresponding to the target video and the annotation information of each comment data are collected into the comment data set corresponding to the target video.
6. The method according to claim 1, characterized in that The target video and its corresponding comment dataset are used as a training set to fine-tune the preset multimodal model to obtain a trained first multimodal model, including: The video frame of the target video and each comment data are taken as input, and the annotation information of the current comment data is taken as output. The network structure used to process the video frames in the preset multimodal model is trained using the Lora fine-tuning method of the supervised task, and the network structure used to process the comment data is trained using the QLoRa fine-tuning method of the supervised task to obtain the first multimodal model after training.
7. The method according to claim 6, characterized in that The step of inputting the designated video and its corresponding comment data into the first multimodal model to obtain a review result of the comment data of the designated video includes: Inputting a video frame of a specified video into a network structure for processing the video frame in a first multimodal model to obtain a video frame feature; Inputting comment data corresponding to a specified video into a network structure for processing comment data in a first multimodal model to obtain comment data features; The video frame features and the comment data features are input into the fusion network structure in the first multimodal model to obtain the review results of the comment data of the specified video.
8. A video comment review device, characterized in that: include: The acquisition module is used to obtain the target video and its corresponding comment data; A construction module is used to annotate the comment data corresponding to the target video based on the correlation between the target video and its corresponding comment data, and to construct a comment dataset corresponding to the target video; A training module is used to use the target video and its corresponding comment dataset as a training set to fine-tune the preset multimodal model to obtain a trained first multimodal model; An input module is configured to receive a specified video and its corresponding comment data, input the specified video and its corresponding comment data into a first multimodal model, obtain an audit result of the comment data of the specified video, and process the comment data with a negative audit result. The preset multimodal model is a pre-trained second multimodal model. The training and use process of the pre-trained second multimodal model includes the following steps: the first step is to construct a news image and text data set, and the data set is annotated in the way of adding labels to the image and text; the second step is to use the multimodal large model to convert the image and text into feature vectors and train and learn the joint distribution characteristics of the image and text vectors; the third step is to search and match the image and text data input by the user, and use the trained multimodal large model for prediction; the fourth step is to review and evaluate the actual image and text content; the fifth step is to establish a feedback mechanism between the evaluation results and the manual review results, and continuously optimize the accuracy of the model evaluation. In the third step, the trained big model is used to predict the results, microservices are established through the big model, and a content review system based on the multimodal big model is formed. The content review system architecture based on the multimodal big model includes hardware, software, database and functional modules. The system functional modules include system data conversion module, content review algorithm encapsulation module, system output module, unstructured data processing module, multimodal data integration module, content analysis module, content anomaly quantification module and intelligent control module. The content analysis module is used to preliminarily determine the degree of match between text and image content based on the semantic association graph generated by the unstructured data processing module, generate a contextual factor Qycs, and analyze the potential sensitive content in each modal data to generate a sensitivity factor Gmyz; then, based on the contextual factor Qycs and the sensitivity factor Gmyz, perform a hierarchical analysis of the potential emotional factors, socially sensitive information, and potential misleading factors in the image and text content to generate a multimodal association factor Mmgyz; The content anomaly quantification module calculates the audit weight value Hqwz based on the context factor Qycs, the sensitivity factor Gmyz, and the multimodal correlation factor Mmgyz; then obtains the system audit risk factor Rxz by calculation and associates it with the audit weight value Hqwz, thereby generating the final content anomaly index Cxzs; The calculation formula for the audit weight value Hqwz is: ; Among them, b1, b2 and b3 represent the preset weight coefficients of the situational factor Qycs, the sensitive factor Gmyz and the multimodal correlation factor Mmgyz respectively, and B represents the audit correction coefficient; The intelligent control module is used to set the compliance threshold T and the audit risk threshold R, and compare the audit weight value Hqwz with the compliance threshold T to determine whether the current content meets the audit standards, and compare the audit risk threshold R with the content anomaly index Cxzs to obtain the final audit result, and selectively trigger manual review or automatic release control plans based on the audit results.
9. A computer device comprising: A memory and a processor, wherein the memory stores a computer program, and wherein the processor implements the steps of the video comment review method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the video comment review method described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Social media negative emotion recognition method based on generative artificial intelligence
CN117493973A
Bullet screen classification method and device, computer equipment and storage medium
CN118781586A