A multi-modal content review method, system and computer storage medium
Through real-time acquisition and self-supervised learning, the multimodal large model base is generated, combined with dynamic audit rules and situational adaptability scoring functions, and the problems of inefficient and insufficient accuracy of multimodal data audit in the existing technology are solved, and efficient and accurate multimodal content audit is achieved.
Patent Information
- Application Number
- CN202510115433.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-01-24
AI Technical Summary
When processing multimodal data, the prior art cannot fully explore the correlation between the various modalities, resulting in inaccurate audit results and lack of real-time performance, unable to adapt to the rapidly changing network environment, resulting in inefficient auditing.
By collecting text, pictures, audio and video resources on the Internet in real time, using self-supervised learning technology to generate multimodal large model bases, integrating the pre-training process of multimodal data, building a multimodal audit model, and dynamically generating audit rules and context adaptive scoring functions, the context-aware analysis of real-time audit requests is realized.
It realizes the efficiency and accuracy of multimodal content audits, can fully understand the correlation between different modes, dynamically generate accurate audit strategies, adapt to the rapidly changing network environment, and improves the intelligence and automation level of audits.
Smart Images

Figure CN119577460B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of content review, and particularly to a multi-modal content review method, system and computer storage medium. Background Art
[0002] Today, with the rapid development of information technology, content review technology has become an important means to ensure the security and health of the network environment. Especially on social media and video sharing platforms, with the rapid growth of user-generated content, traditional review methods are facing huge challenges. Early reviews mainly relied on manual reviews and rule-based automated systems, which could provide a certain level of efficiency when dealing with single-modal data, but were insufficient in terms of efficiency and accuracy when facing multi-modal data (such as text, pictures, audio, and video). With the development of deep learning technology, review models based on deep learning have gradually emerged. However, existing models mostly focus on single-modal data and lack the ability to comprehensively understand and process multi-modal data.
[0003] The deficiencies of the existing technology are mainly reflected in two aspects: First, existing single-modal review systems cannot fully explore the correlations between modalities when dealing with complex multi-modal content, resulting in inaccurate review results; Second, existing models often lack real-time performance and cannot adapt to the rapidly changing network environment, leading to low review efficiency. These deficiencies make it difficult for content review to meet industry standards in terms of efficiency and accuracy. Therefore, in response to the real-time review requirements for multi-modal data, there is an urgent need for a new review method to improve the intelligence and automation level of review. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides a multi-modal content review method to solve the problems of low efficiency and insufficient accuracy in multi-modal content review.
[0006] To solve the above technical problems, the present invention provides the following technical solutions:
[0007] In a first aspect, the present invention provides a multi-modal content review method, which includes: real-time collecting text, picture, audio and video resources on the Internet;
[0008] Through self-supervised learning, pre-training the text, picture, audio and video resources to generate a multi-modal large model base;
[0009] The multi-modal large model base is adapted to the review task through fine-tuning training to obtain a multi-modal model;
[0010] Using prompt engineering to guide the multi-modal model to generate an understanding Q&A dataset as the fine-tuning dataset;
[0011] Based on the fine-tuning dataset, a multi-modal audit model is trained and constructed through an iterative fine-tuning process;
[0012] Receive an audit request through the multi-modal audit model, activate the context-aware intelligent engine, and dynamically generate audit rules in combination with the audit task to perform audit determination;
[0013] Real-time collect text, pictures, audio, and video resources on the Internet. The specific steps are as follows.
[0014] Select social media platforms and news websites as data sources;
[0015] Based on the data sources, use Kafka to coordinate the distributed crawler nodes of the Scrapy framework to perform the scraping of text, pictures, audio, and video resources;
[0016] Preprocess the scraped resources through OpenCV and FFmpeg;
[0017] Encrypt the preprocessed text, pictures, audio, and video resources and store them in Ceph to form a data stream containing text, images, audio, and video.
[0018] As a preferred solution of the multi-modal content audit method described in the present invention, wherein: through self-supervised learning, pre-train the text, pictures, audio, and video resources to generate a multi-modal large model base. The specific steps are as follows.
[0019] Integrate the text, image, audio, and video data streams and perform annotation to generate a multi-modal dataset with annotation information;
[0020] Based on the multi-modal dataset, use the CLIP model as the basis, utilize the self-supervised learning strategy, and use the contrastive learning method for pre-training;
[0021] After completing the pre-training, adjust the hyperparameters to optimize the performance, and evaluate the generalization ability of the CLIP model through the validation set;
[0022] Save the evaluated and optimized CLIP model as the multi-modal large model base.
[0023] As a preferred solution of the multi-modal content audit method described in the present invention, wherein: the multi-modal large model base adapts to the audit task through fine-tuning training to obtain a multi-modal model. The specific steps are as follows.
[0024] Initialize the multi-modal large model base, starting with the CLIP model;
[0025] Based on the CLIP model and multimodal datasets, an adaptive parametric audit function is created by combining exponential decay integral normalization and complex summation functions, and the definition is as follows:
[0026] ;
[0027] where, is the feature vector of the content to be audited, is the corresponding expected output, is the number of samples, is the learning rate parameter of the th sample, is the sample index variable, is the decay factor is used to measure the th sample's feature vector of the content to be audited and the expected output the distance metric function between them, represents the integral of the infinitesimal increment of time ;
[0028] Using the learning rate parameter and the decay factor , calculate the adaptive parametric audit function scores for each sample in the multimodal dataset;
[0029] According to the adaptive parametric audit function scores, fine-tune and train the CLIP model using the multimodal dataset;
[0030] In each fine-tuning, recalculate the adaptive parametric audit function scores, evaluate the adaptability of the CLIP model to the audit task, and dynamically adjust the learning rate and the decay factor parameters;
[0031] The CLIP model after multiple fine-tuning trainings and evaluations is used as the multimodal model.
[0032] As a preferred solution of the multimodal content audit method described in the present invention, among them: using prompt engineering to guide the multimodal model to generate an understanding Q&A dataset as the fine-tuning dataset, the specific steps are as follows,
[0033] Create a prompt template for task instructions, context information, and expected answer forms through prompt engineering;
[0034] Apply the prompt template to the multimodal model to generate question-answer pairs;
[0035] Apply the graph neural network GNN to capture the relationships between different modal data in the multimodal dataset and construct a graph structure;
[0036] Optimize the question-and-answer pairs through the graph structure to generate a question-and-answer pair dataset;
[0037] Introduce an evaluation mechanism to evaluate the question-and-answer pair dataset using a validation set;
[0038] Adjust the evaluated question-and-answer pair dataset through feedback to a high-quality question-and-answer pair dataset;
[0039] Integrate the high-quality question-and-answer pair dataset as a fine-tuning dataset.
[0040] As a preferred solution of the multimodal content review method described in the present invention, wherein: based on the fine-tuning dataset, train and construct a multimodal review model through an iterative fine-tuning process, and the specific steps are as follows,
[0041] Select the current review task base model and the fine-tuning dataset according to the multimodal large model base;
[0042] Load the fine-tuning dataset into the training environment;
[0043] Use the loaded fine-tuning dataset to perform iterative fine-tuning on the multimodal review model;
[0044] Evaluate the generalization ability of the fine-tuned multimodal review model using an independent validation set.
[0045] As a preferred solution of the multimodal content review method described in the present invention, wherein: receive a review request through the multimodal review model, activate the context-aware intelligent engine and dynamically generate review rules in combination with the review task for review determination, and the specific steps are as follows,
[0046] The multimodal review model receives a review request and parses the text, image, audio, and video metadata of the review request;
[0047] According to the parsed review request metadata, activate the context-aware intelligent engine and analyze the context characteristics of the review request;
[0048] Introduce a context adaptability scoring function , to guide the generation and adjustment of review rules, and the expression is:
[0049] ;
[0050] Wherein, is the context adaptability scoring function, is the time range, is the time decay coefficient, is the information filtering function, is the reference time point;
[0051] Dynamically generate audit rules using an adaptive parameterized audit function based on the results of the context adaptability scoring function and the audit task objectives;
[0052] Apply the audit rules to conduct an audit judgment on the audit request, and refine the application of the audit rules in combination with the context information provided by the context adaptability scoring function.
[0053] In a second aspect, the present invention provides a multi-modal content audit system, including a data collection module, a model training module, a task fine-tuning module, a prompt generation module, an iterative optimization module, and an intelligent audit module; the data collection module is used to collect text, pictures, audio, and video resources on the Internet in real time; the model training module is used to pre-train the text, pictures, audio, and video resources through self-supervised learning to generate a multi-modal large model base; the task fine-tuning module is used to adapt the multi-modal large model base to the audit task through fine-tuning training to obtain a multi-modal model; the prompt generation module is used to use prompt engineering to guide the multi-modal model to generate an understanding Q&A data set as a fine-tuning data set;
[0054] The iterative optimization module is used to train and construct a multi-modal audit model based on the fine-tuning data set through an iterative fine-tuning process; the intelligent audit module is used to receive the audit request through the multi-modal audit model, activate the context-aware intelligent engine, and dynamically generate audit rules in combination with the audit task to conduct an audit judgment.
[0055] In a third aspect, the present invention provides a computer device, including a memory and a processor, where the memory stores a computer program, and: when the computer program is executed by the processor, any step of the multi-modal content audit method described in the first aspect of the present invention is implemented.
[0056] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and: when the computer program is executed by the processor, any step of the multi-modal content audit method described in the first aspect of the present invention is implemented.
[0057] The beneficial effects of the present invention are as follows: By collecting various types of data (text, pictures, audio, and video) in real time, the present invention uses self-supervised learning technology to generate the foundation of a multimodal large model, integrating the pre-training process of multimodal data, enabling the multimodal model to fully understand the correlation between different modalities and applying it to the review task. Secondly, by dynamically generating review rules and context-adaptive scoring functions, the multimodal review model can perform context-aware analysis on real-time review requests and generate more accurate review strategies. In addition, the introduction of iterative fine-tuning and feedback mechanisms enables the multimodal review model to continuously optimize in practical applications and further improve its performance. These series of innovative steps give the present invention significant technical advantages in the field of multimodal content review and contribute to the realization of an efficient and secure network environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0059] Figure 1 It is a flowchart of the multimodal content review method in Embodiment 1;
[0060] Figure 2 It is a module diagram of the multimodal content review system in Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0061] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will provide a detailed description of the specific embodiments of the present invention with reference to the accompanying drawings of the specification.
[0062] In the following description, many specific details are set forth to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0063] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that can be included in at least one implementation manner of the present invention. The "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor is it a separate or selectively exclusive embodiment from other embodiments.
[0064] Embodiment 1, refer to Figure 1 and Figure 2, which is the first embodiment of the present invention. This embodiment provides a multi-modal content review method, including the following steps:
[0065] S1. Real-time collect text, picture, audio, and video resources on the Internet.
[0066] Furthermore, select social media platforms and news websites as data sources;
[0067] It should be noted that the characteristics of these platforms are fast information update, wide coverage, and high user participation. Therefore, they can provide a large amount of diverse raw data. This selection can ensure the freshness and relevance of the data, and at the same time reflect the changes in social hotspots and public opinions. It achieves the effect of obtaining high-quality and time-sensitive data resources, providing a basis for subsequent data analysis and applications.
[0068] Based on the data sources, use Kafka to coordinate the distributed crawler nodes of the Scrapy framework to perform the crawling of text, picture, audio, and video resources;
[0069] Specifically, by introducing Kafka to coordinate the distributed crawler nodes under the Scrapy framework, an efficient and scalable multi-type resource crawling mechanism is achieved. As a message middleware, Kafka receives upstream crawler tasks and reasonably distributes these tasks to each Scrapy node. Each task includes the target URL and the text, picture, audio, or video to be crawled. It ensures the reliability and order of message delivery, enabling the crawler tasks to reasonably distribute the load among multiple nodes and maintaining stability and efficiency even in the face of massive data.
[0070] Preprocess the crawled resources through OpenCV and FFmpeg;
[0071] Among them, the preprocessing includes the conversion of the original materials to a standardized format. OpenCV is used for image recognition and processing, such as removing noise and adjusting the resolution; FFmpeg is used for audio and video encoding, decoding, and format conversion. This step helps to eliminate redundant information in the data, improve the subsequent storage and retrieval efficiency, and at the same time provide a unified standard data format for data analysis.
[0072] Encrypt and store the preprocessed text, picture, audio, and video resources in Ceph to form a data stream containing text, images, audio, and video.
[0073] Preferably, Ceph is known for its high scalability, reliability, and flexibility, making it particularly suitable for large-scale data storage requirements. Data encryption ensures the privacy and integrity of information, preventing unauthorized access or leakage. In addition, a data stream integrating various media types such as text, images, audio, and video is constructed to facilitate subsequent data integration and analysis, supporting cross-media content association research.
[0074] S2. Through self-supervised learning, pre-train text, picture, audio, and video resources to generate a multi-modal large model base.
[0075] Furthermore, integrate text, image, audio, and video data streams and perform annotation to generate a multi-modal dataset with annotation information;
[0076] Based on the multi-modal dataset, using the CLIP model as the basis, adopt a self-supervised learning strategy and use the contrastive learning method for pre-training;
[0077] Specifically, adopt the CLIP model and introduce a self-supervised learning strategy, especially the contrastive learning method for pre-training, to effectively extract features even without a large amount of labeled data. The CLIP model can handle text and images simultaneously. Self-supervised learning reduces the dependence on expensive labeled data, and the contrastive learning method helps to learn to distinguish similar and different samples.
[0078] After completing pre-training, adjust the hyperparameters to optimize the performance and evaluate the generalization ability of the CLIP model through the validation set;
[0079] Among them, hyperparameter adjustment is a key process to find the optimal configuration, and validation set evaluation provides important feedback on the actual performance. Make the CLIP model perform well on the training data and have excellent prediction ability on new data. Improve the practical application value and enable the CLIP model to better serve real-world problem-solving.
[0080] Save the evaluated and optimized CLIP model as the multi-modal large model base.
[0081] Preferably, the multi-modal large model base can be used for transfer learning in the future, quickly adapting to new application scenarios or specific domain problems, greatly shortening the R & D cycle. In addition, it can also serve as a research platform to promote further exploration and development in the academic and industrial fields. Ultimately, it lays a solid foundation for subsequent application innovation and technological progress, promoting the development of multi-modal artificial intelligence technology.
[0082] S3. The multi-modal large model base adapts to the review task through fine-tuning training to obtain a multi-modal model.
[0083] Furthermore, initialize the multi-modal large model base using the CLIP model as a starting point;
[0084] It should be noted that selecting the pre-trained CLIP model as the initial base of the multi-modal large model enables a quick start and utilization of a CLIP model that has been trained on a large number of text-image pairs. Due to its powerful cross-modal representation ability, the CLIP model can serve as an ideal starting point for handling text-image association tasks.
[0085] Based on the CLIP model and the multi-modal dataset, create an adaptive parametric auditing function by combining exponential decay integral normalization and complex summation function. The expression is:
[0086] ;
[0087] where, is the feature vector of the content to be audited, is the corresponding expected output, is the number of samples, is the learning rate parameter of the th sample, is the sample index variable, is the decay factor, is to measure the th sample's feature vector of the content to be audited and the expected output the distance metric function between them, represents integrating the infinitesimal increment of time ;
[0088] where, is the feature vector of the content to be audited. The content to be audited refers to various types of media files or data that need to be evaluated and inspected during the audit. This content can include, but is not limited to, text, images, audio, and video. By introducing the custom adaptive parametric auditing function , a refined evaluation of the CLIP model output is achieved. This function comprehensively considers the learning rate parameters of all samples, the decay factor and the distance metric between the feature vector of the content to be audited and the expected output.
[0089] Preferably, this design allows the CLIP model to dynamically adjust the audit criteria according to specific task requirements, thus more accurately reflecting the importance of different samples. This step provides a flexible and precise mechanism to measure the actual performance and provides a quantitative basis for subsequent optimization.
[0090] Using the learning rate parameter and attenuation factor , calculate the adaptive parameterized audit function score for each sample in the multimodal dataset;
[0091] Based on the adaptive parameterized audit function scores, the CLIP model is fine-tuned using multimodal datasets;
[0092] Among them, fine-tuning training uses a deep learning framework to build a training environment, ensures support for distributed training and GPU acceleration, and sets key parameters such as the initial learning rate, batch size, optimizer type, and loss function. The loss value is calculated based on the score of the adaptive parameterized audit function. Samples with higher scores in the adaptive parameterized audit function will be given greater weights, thereby affecting the loss function.
[0093] In each fine-tuning, the adaptive parameterized audit function score is recalculated to evaluate the adaptability of the CLIP model to the audit task and to dynamically adjust the learning rate. and attenuation factor parameter;
[0094] Specifically, the adaptability of the CLIP model is evaluated based on the adaptive parameterized audit function score, and the learning rate is dynamically adjusted accordingly. and attenuation factor , realizing intelligent management of the training process. In this step, the CLIP model can automatically adjust key parameters according to actual performance to find the best configuration. This not only improves training efficiency, but also enhances the pertinence and accuracy of specific audit tasks. The dynamic adjustment strategy helps to overcome the problem of overfitting or underfitting caused by fixed parameters, ensuring that the CLIP model can maximize the fit of the target task while maintaining good generalization ability.
[0095] The CLIP model after multiple fine-tuning training and evaluation is used as a multimodal model.
[0096] It should be noted that the CLIP model gradually improves its prediction ability and decision-making logic by repeatedly learning feedback from the real world. Each fine-tuning is based on the results of the previous one to ensure continuous progress towards the optimal solution. After multiple fine-tuning training and evaluation, the multimodal model will have excellent audit capabilities and is suitable for various complex multimodal scenarios.
[0097] S4. Use prompt engineering to guide the multimodal model to generate a comprehension question answering dataset as a fine-tuning dataset.
[0098] Furthermore, by improving engineering, prompt templates are created in the form of task instructions, contextual information, and expected responses;
[0099] It should be noted that by enhancing the engineering creation task instructions and context information, it is ensured that the generated question-answer pairs have clear goals and structures. This process makes the multimodal model more accurate in understanding the user's intention, reduces ambiguity, and thus improves the quality and relevance of the generated question-answer pairs in subsequent steps.
[0100] Apply the prompt template to the multimodal model to generate question-answer pairs;
[0101] Apply the graph neural network GNN to capture the relationships between multimodal data and construct a graph structure;
[0102] Among them, applying the graph neural network GNN can efficiently capture the complex relationships between different modal data and construct a comprehensive graph structure. The graph structure not only enhances the connectivity of multimodal data but also provides a basis for subsequent optimization, helping to improve the understanding and reasoning abilities of the multimodal model.
[0103] Optimize the question-answer pairs through the graph structure to generate a question-answer pair dataset;
[0104] Introduce an evaluation mechanism and use the validation set to evaluate the question-answer pair dataset;
[0105] It should be noted that using BlenderBot as an evaluation tool combined with automatic evaluation metrics, combining the output of BlenderBot with the automatic evaluation metric ROUGE to form a hybrid evaluation mechanism, improves the quality control level of the question-answer pair dataset and provides reliable feedback information for training. By combining the automatic evaluation metrics with the evaluation of the dialogue model, retaining the efficiency of automatic evaluation and introducing the flexibility and in-depth understanding ability of the dialogue model, the balance between objectivity and subjectivity of evaluation is achieved. This method makes full use of the advantages of existing technologies, reflects the application of innovative thinking, and has high practical value.
[0106] Preferably, introduce an evaluation mechanism and use the validation set to evaluate the question-answer pair dataset to ensure that the generated dataset meets high-quality standards. This enables the development team to identify and correct potential problems in a timely manner, improving the effectiveness and usability of the question-answer pairs, and thus providing better services for users.
[0107] Adjust the evaluated question-answer pair dataset through feedback to a high-quality question-answer pair dataset;
[0108] Integrate the high-quality question-answer pair dataset as a fine-tuning dataset.
[0109] Specifically, by integrating high-quality Q&A pair datasets that have undergone strict screening and optimization as the basis for fine-tuning the multimodal model, the performance of the multimodal model has been refined. During the fine-tuning process, the multimodal model can learn the nuances and deep features contained in these high-quality data, thereby significantly enhancing its understanding and generation capabilities. Eventually, the effect of optimized performance is achieved, enabling the multimodal model to have higher precision and adaptability in specific tasks, making full preparations for practical applications.
[0110] S5. Based on the fine-tuning dataset, train and construct a multimodal audit model through an iterative fine-tuning process.
[0111] Furthermore, according to the multimodal large model base, select the basic model for the current audit task and the fine-tuning dataset;
[0112] Load the fine-tuning dataset into the training environment;
[0113] It should be noted that loading the selected fine-tuning dataset into the training environment realizes the seamless docking of data with the multimodal model. This process not only simplifies the transition from data preparation to training but also ensures the consistency and integrity of data transmission. The configuration of the training environment allows for flexible adjustment of parameter settings to adapt to datasets of different scales, achieving the effect of improving training efficiency, reducing time waste and technical problems caused by improper data processing, and laying a solid foundation for the subsequent fine-tuning training.
[0114] Use the loaded fine-tuning dataset to perform iterative fine-tuning on the multimodal audit model;
[0115] Specifically, use the loaded fine-tuning dataset to perform iterative fine-tuning on the multimodal audit model to gradually improve its performance. During each iteration, the multimodal audit model can learn more specific details about the audit task and continuously optimize its prediction ability. Iterative fine-tuning helps the multimodal audit model overcome the initial bias and gradually approach the optimal solution. Eventually, the beneficial effects of enhanced accuracy and robustness are achieved, enabling it to perform well in specific audit tasks and be able to more accurately handle complex and changing practical application scenarios.
[0116] Evaluate the generalization ability of the fine-tuned multimodal audit model using an independent validation set.
[0117] Among them, the role of the independent validation set is to provide a set of unseen data samples to test whether the multimodal audit model truly masters the core features of the audit task, rather than simply memorizing the training data. This step ensures that the multimodal audit model can not only perform well on the training data but also maintain stable and reliable performance on new and unknown data.
[0118] S6. Receive an audit request through the multimodal audit model, activate the context-aware intelligent engine, and dynamically generate audit rules in combination with the audit task for audit determination.
[0119] Furthermore, the multimodal audit model receives the audit request and parses the text, picture, audio, and video metadata of the audit request;
[0120] By receiving and parsing the text, picture, audio, and video metadata of the audit request through the multimodal audit model, after receiving the audit request, the multimodal audit model uses natural language processing NLP, computer vision CV, and audio processing technologies to parse the text, picture, audio, and video metadata therein, extract key information such as text content, image objects, speech-to-text, and video behaviors, and generate structured audit request metadata, providing comprehensive and detailed data support for subsequent context analysis and audit decision-making, thereby enhancing the accuracy and reliability of the audit process. Achieve a comprehensive understanding of complex audit requests. This step is not limited to processing a single type of data, but covers various information sources such as text, images, audio, and video, ensuring that all relevant information can be fully considered in the audit process. Ultimately, enhance the comprehensiveness of audit decisions and provide a solid foundation for more accurate subsequent context analysis.
[0121] According to the parsed audit request metadata, activate the context-aware intelligent engine to analyze the context characteristics of the audit request;
[0122] Specifically, to analyze the context characteristics of the audit request, the multimodal audit model is used to compare historical data or known patterns, identify the similarities and differences between the current audit request and previous cases, and based on the parsed information, evaluate the risk level of the audit request and attempt to infer the motivation or purpose behind the user's submission of the audit request, which helps to more accurately judge the true intention of the request.
[0123] Introduce a context adaptability scoring function , to guide the generation and adjustment of audit rules, and the expression is:
[0124] ;
[0125] Among them, is the context adaptability scoring function, is the time range, is the time decay coefficient, is the information filtering function, used to measure the context characteristics at time and the audit request The matching degree between them, and this function can be implemented by selecting an appropriate algorithm according to the actual situation, is the reference time point;
[0126] It should be noted that a context adaptability scoring function is introduced to quantitatively evaluate the degree of match between the review request and the current context features. The time decay coefficient in the formula and the information filtering function act together to enable the score to dynamically reflect the change in the matching degree over different time periods. This mechanism helps to maintain sensitivity and responsiveness in a constantly changing context, achieving the effect of optimizing the generation and adjustment of review rules, ensuring that the rules always meet the latest context requirements, and enhancing flexibility and adaptability.
[0127] According to the results of the context adaptability scoring function and the review task objectives, an adaptive parameterized review function is used to dynamically generate review rules;
[0128] Preferably, combining the results of the context adaptability scoring function with the review task objectives, an adaptive parameterized review function is used to dynamically generate review rules to achieve personalized customization of review rules. This method flexibly adjusts the review criteria according to specific context features, ensures the consistency of the rules, and enhances the pertinence. It improves the review efficiency and accuracy, enabling the review process to respond quickly and execute accurately, meeting diverse review requirements.
[0129] Apply the review rules to conduct a review and determination of the review request, and refine the application of the review rules by combining the context information provided by the context adaptability scoring function.
[0130] Among them, the context information provided by the context adaptability scoring function refines the rule application, realizing the refined management of review decisions. It not only relies on preset rules but also makes full use of the results of real-time context analysis to ensure that each review rule can be appropriately applied to specific situations. This improves the review quality, ensures the fairness and reasonableness of the review results, and at the same time improves user satisfaction and credibility.
[0131] Specifically, based on the adaptive parameterized review function to dynamically generate review rules, combining the context information provided by the context adaptability scoring function, the review request is matched and screened with the review rules. First, a preliminary screening is carried out according to the text content and video tags to identify keywords that may involve sensitive topics. For the review requests that pass the screening, the review rules are refined and adjusted according to the context information provided by the context adaptability scoring function. For normal situations, review conclusions are automatically generated according to the refined review rules. For complex situations, these review requests are marked and manual intervention is required for a second review, completing the review and determination efficiently and accurately. It makes full use of the results of real-time context analysis to ensure that each review rule can be appropriately applied to specific situations, improving the accuracy of the review.
[0132] This embodiment also provides a multimodal content review system, including: a data collection module, a model training module, a task fine-tuning module, a prompt generation module, an iterative optimization module, and an intelligent review module; The data collection module is used to collect text, pictures, audio, and video resources on the Internet in real time; The model training module is used to pre-train text, pictures, audio, and video resources through self-supervised learning to generate a multimodal large model base; The task fine-tuning module is used to adapt the multimodal large model base to the review task through fine-tuning training to obtain a multimodal model; The prompt generation module is used to use prompt engineering to guide the multimodal model to generate an understanding Q&A dataset as the fine-tuning dataset; The iterative optimization module is used to train and construct a multimodal review model based on the fine-tuning dataset through an iterative fine-tuning process; The intelligent review module is used to receive a review request through the multimodal review model, activate the context-aware intelligent engine, and dynamically generate review rules in combination with the review task for review determination.
[0133] This embodiment also provides a computer device applicable to the multimodal content review method, including: a memory and a processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the multimodal content review method proposed in the above embodiment.
[0134] This computer device can be a terminal. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of this computer device is used to provide computing and control capabilities. The memory of this computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of this computer device is used to communicate with an external terminal in a wired or wireless manner. The wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. The display screen of this computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of this computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad set on the computer device housing. It can also be an external keyboard, touchpad, or mouse, etc.
[0135] This embodiment also provides a storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the multi-modal content review method proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, abbreviated as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, abbreviated as EEPROM), erasable programmable read-only memory (Erasable Programmable Read Only Memory, abbreviated as EPROM), programmable read-only memory (Programmable Red-Only Memory, abbreviated as PROM), read-only memory (Read-Only Memory, abbreviated as ROM), magnetic memory, flash memory, magnetic disk or optical disc.
[0136] In summary, the present invention collects various types of data (text, pictures, audio, and video) in real time, uses self-supervised learning technology to generate a multi-modal large model base, integrates the pre-training process of multi-modal data, enables the multi-modal model to fully understand the correlation between different modalities, and applies it to the review task. Secondly, by dynamically generating review rules and context-adaptive scoring functions, the multi-modal review model can perform context-aware analysis on real-time review requests and generate more accurate review strategies. In addition, the introduction of iterative fine-tuning and feedback mechanisms enables the multi-modal review model to be continuously optimized in practical applications, further improving its performance. These series of innovative steps give the present invention significant technical advantages in the field of multi-modal content review and contribute to the realization of an efficient and secure network environment.
[0137] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A multimodal content review method, characterized by: include, Collect text, pictures, audio and video resources on the Internet in real time; Through self-supervised learning, pre-train text, image, audio and video resources to generate a multi-modal large model base; The multimodal large model base is fine-tuned to adapt to the audit task to obtain a multimodal model; Use prompt engineering to guide multimodal models to generate understanding question answering datasets as fine-tuning datasets; Initialize the multimodal large model base, using the CLIP model as a starting point; Based on the CLIP model and multimodal dataset, an adaptive parameterized audit function is created by combining exponential decay integral normalization and complex summation function. The expression is: ; in, is the feature vector of the content to be reviewed, is the corresponding expected output, is the total number of samples, For the The learning rate parameter of the sample, is the sample index variable, is the attenuation factor, To measure the The feature vector of the content to be reviewed of samples with expected output The distance metric function of the difference between Indicates time Integrate infinitesimal increments of Using the learning rate parameter and attenuation factor , calculate the adaptive parameterized audit function score for each sample in the multimodal dataset; Based on the adaptive parameterized audit function scores, the CLIP model is fine-tuned using multimodal datasets; In each fine-tuning, the adaptive parameterized audit function score is recalculated to evaluate the adaptability of the CLIP model to the audit task and to dynamically adjust the learning rate. and attenuation factor parameter; The CLIP model after multiple fine-tuning training and evaluation is used as a multimodal model; Based on the fine-tuning dataset, a multimodal audit model is constructed through iterative fine-tuning training. Receive audit requests through the multimodal audit model, activate the context-aware intelligent engine and dynamically generate audit rules based on the audit tasks to make audit decisions; Select social media platforms and news websites as data sources; Based on the data source, Kafka is used to coordinate the distributed crawler nodes of the Scrapy framework to perform text, image, audio and video resource crawling; Preprocess the captured text, images, audio and video resources through OpenCV and FFmpeg; The pre-processed text, image, audio and video resources are encrypted and stored in Ceph, forming a data stream containing text, image, audio and video.
2. The multimodal content review method according to claim 1, characterized in that: The self-supervised learning is used to pre-train text, pictures, audio and video resources to generate a multi-modal large model base. The specific steps are as follows: Integrate text, image, audio and video data streams, and annotate them to generate multimodal datasets with annotated information; Based on multimodal datasets, the CLIP model is used as the basis, self-supervised learning strategy is used, and contrastive learning method is used for pre-training; After pre-training, adjust the hyperparameters to optimize performance, and evaluate the generalization ability of the CLIP model through the validation set; Save the evaluated and optimized CLIP model as a multimodal large model base.
3. The multimodal content review method according to claim 2, characterized in that: The method of using prompt engineering to guide the multimodal model to generate an understanding question-answering dataset as a fine-tuning dataset has the following specific steps: Create prompt templates in the form of task instructions, contextual information, and expected responses through enhancement engineering; Apply the prompt template to the multimodal model to generate question-answer pairs; Apply graph neural network (GNN) to capture the relationship between different modal data in multimodal datasets and construct graph structures; Optimize question-answer pairs through graph structure and generate question-answer pair dataset; An evaluation mechanism is introduced to evaluate the question-answer pair dataset using a validation set. Adjust the evaluated question-answer pair dataset into a high-quality question-answer pair dataset through feedback; Integrate a high-quality question-answer pair dataset as a fine-tuning dataset.
4. The multimodal content review method according to claim 3 is characterized in that: Based on the fine-tuning dataset, the multimodal audit model is constructed through iterative fine-tuning training. The specific steps are as follows: Based on the multimodal large model base, select the basic model and fine-tuning dataset for the current audit task; Load the fine-tuning dataset into the training environment; Iteratively fine-tune the multimodal audit model using the loaded fine-tuning dataset; The generalization ability of the fine-tuned multimodal audit model is evaluated using an independent validation set.
5. The multimodal content review method according to claim 4, characterized in that: The multimodal audit model receives the audit request, activates the context-aware intelligent engine, and dynamically generates the audit rules in combination with the audit task to make the audit decision. The specific steps are as follows: The multimodal review model receives the review request and parses the text, image, audio, and video metadata of the review request; Based on the parsed audit request metadata, activate the context-aware intelligent engine to analyze the contextual characteristics of the audit request; Introducing contextual adaptability scoring function , guiding the generation and adjustment of audit rules, the expression is: ; in, is the context adaptability scoring function, is the time range, is the time decay coefficient, is the information filtering function, is the reference time point; According to the results of the situational adaptability scoring function and the audit task objectives, the audit rules are dynamically generated using the adaptive parameterized audit function; Apply audit rules to make audit decisions on audit requests, and refine the application of audit rules based on context information provided by the situational adaptability scoring function.
6. A multimodal content review system, based on the multimodal content review method according to any one of claims 1 to 5, characterized in that: Including data collection module, model training module, task fine-tuning module, prompt generation module, iterative optimization module and intelligent review module; The data collection module is used to collect text, pictures, audio and video resources on the Internet in real time; The model training module is used to pre-train text, image, audio and video resources through self-supervised learning to generate a multi-modal large model base; The task fine-tuning module is used to fine-tune the multimodal large model base to adapt to the audit task and obtain a multimodal model; The prompt generation module is used to use prompt engineering to guide the multimodal model to generate an understanding question-answering dataset as a fine-tuning dataset; The iterative optimization module is used to construct a multimodal audit model through iterative fine-tuning process training based on the fine-tuning dataset; The intelligent audit module is used to receive audit requests through a multimodal audit model, activate the context-aware intelligent engine, dynamically generate audit rules in combination with the audit tasks, and make audit decisions.
7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the multimodal content review method described in any one of claims 1 to 5 are implemented.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the multimodal content review method described in any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Model training data generation method and device, equipment and storage medium
CN117972052A
Multi-mode network content security intelligent auditing system and method thereof
CN118312922A
Intelligent text auditing method and device fusing big data and small data model
CN118761409A