A method for detecting objectionable content based on generative artificial intelligence driving
By using a generative artificial intelligence-driven approach, we can automatically collect and analyze data on inappropriate content, generate diverse and highly concealed content, optimize the detection model, and solve the problems of inaccurate detection and reliance on manual labor in existing technologies. This achieves efficient and accurate detection of inappropriate content and reduces costs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING UNIV OF POSTS & TELECOMM
- Filing Date
- 2023-11-06
- Publication Date
- 2026-05-26
Smart Images

Figure CN117407725B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and more specifically to a method for detecting inappropriate content based on generative artificial intelligence. Background Technology
[0002] Some manufacturers have developed various content monitoring platforms to achieve intelligent analysis and review of big data, thereby detecting and controlling inappropriate content. However, the following problems exist:
[0003] Poor transferability: Currently, mainstream AI content moderation technologies mainly include blacklists, keyword filtering, rule matching, machine learning classification models, and multimodal feature fusion. These technologies primarily rely on rule matching or model training based on existing inappropriate content to detect content with similar characteristics. However, these methods learn a limited amount of inappropriate content and struggle to cope with the rapid spread, iteration, deep concealment, diversity, and sheer volume of such content. Therefore, carefully crafted, highly concealed, and diverse inappropriate content can bypass the detection of existing content security moderation platforms.
[0004] High cost: Existing technologies for content detection employ methods such as user profiling, high-frequency monitoring, and real-time clustering, which typically require significant resource and human resources. This includes establishing and maintaining user profile databases, continuously monitoring online content flows, and processing and analyzing large amounts of data in real time. These operations not only require substantial financial investment but also necessitate hiring professionals to perform them, increasing the cost burden of content moderation.
[0005] Over-reliance on manual review: Due to the limitations of existing technology, manual review is often necessary. This is not only inefficient, but may also harm the physical and mental health of reviewers. Violent, terrifying, and bloody content may cause psychological discomfort to reviewers, and the high intensity of review work may even lead to psychological problems such as depression and stress disorder.
[0006] Therefore, there is a need to research an efficient, accurate, low-cost, AI-driven method for detecting inappropriate content. Summary of the Invention
[0007] To address the problems of inaccurate detection of inappropriate content and over-reliance on manual methods in existing methods, this invention provides a method for detecting inappropriate content based on generative artificial intelligence.
[0008] The technical solution adopted by this invention to solve the technical problem is as follows:
[0009] A method for detecting inappropriate content based on generative artificial intelligence includes the following steps:
[0010] S1. Automatically collect inappropriate content based on keywords and / or key words to obtain the first inappropriate content dataset;
[0011] S2. Input the first malicious dataset into the content security detection platform for detection to obtain the first detection result, extract the sample labels of the data in the first malicious dataset, and obtain the malicious features and their combination rules that were not detected by the content security detection platform based on the first detection result and the sample labels.
[0012] S3. By manually analyzing the inappropriate features and their combination rules obtained in S2, highly concealed inappropriate features and novel combination methods of inappropriate content are obtained. Based on the inappropriate features and their combination rules obtained in S2, as well as the highly concealed inappropriate features and novel combination methods of inappropriate content, a prompting project based on AIGC-generated inappropriate content is constructed.
[0013] S4. According to the prompting process, use the AIGC model to generate several texts, images and videos of inappropriate content, input the texts, images and videos of inappropriate content into the content security review platform for detection to obtain a second detection result, obtain the inappropriate content generated by the AIGC model that has not been detected by the content security review platform based on the second detection result, and construct a multi-dimensional, highly concealed inappropriate content dataset accordingly.
[0014] S5. Using the first dataset of inappropriate content and the multivariate dataset of highly concealed inappropriate content, extract the multimodal features of the inappropriate data; based on the sample labels of the first dataset of inappropriate content and the multivariate dataset of highly concealed inappropriate content and the multimodal features, train a content security detection model; optimize the content security detection model until the content security detection model meets the preset requirements.
[0015] The beneficial effects of this invention are:
[0016] This invention, through automatic dataset collection and the integration of an AIGC model, yields a content security detection model with high efficiency and accuracy, capable of detecting highly concealed and diverse inappropriate content. Using this content security detection model reduces reliance on extensive manual labor, lowers costs, alleviates the burden of manual review, and protects the physical and mental health of reviewers. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of the process of the present invention. Detailed Implementation
[0018] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments.
[0019] This invention provides a method for detecting inappropriate content based on generative artificial intelligence, such as... Figure 1 ,include:
[0020] S1. Based on the keywords and / or key words of the inappropriate content, an inappropriate content dataset is automatically collected. The inappropriate content dataset obtained by automatically collecting inappropriate content includes several inappropriate samples. The inappropriate content dataset obtained by automatically collecting inappropriate content is called the first inappropriate content dataset.
[0021] S2. Input the dataset of inappropriate content obtained by automatic collection into the content security review platform for detection to obtain the first detection result. Extract the sample labels of the data in the dataset of inappropriate content obtained by automatic collection. Based on the first detection result and the sample labels, obtain the inappropriate features and their combination rules that were not detected by the content security review platform.
[0022] S3. By manually analyzing the undetected features and their combination patterns of the undesirable content that were not detected by the content security review platform, we obtain highly concealed undesirable features and novel combination methods of the undesirable content (the novel combination method is a combination method that is different from the combination method of undesirable content in the first undesirable content dataset). Based on the undetected features and their combination patterns of the undesirable content that were not detected by the content security review platform, and also based on the highly concealed undesirable features and novel combination methods of the undesirable content, we construct a prompting project for undesirable content generated by AIGC.
[0023] S4. Based on the AIGC-based prompting project for generating inappropriate content obtained in S3, the AIGC model is used to generate several texts, images, and videos of inappropriate content. All the texts, images, and videos of inappropriate content generated by the AIGC model are input into the content security review platform for detection to obtain a second detection result. Based on the second detection result, the inappropriate content generated by the AIGC model that has not been detected by the content security review platform is obtained, and a multi-dimensional, highly concealed inappropriate content dataset is constructed accordingly.
[0024] S5. Analyze the dataset of inappropriate content obtained by automatically collecting inappropriate content and the dataset of inappropriate content with high hidden diversity. Extract the multimodal features of the inappropriate content dataset. Input these multimodal features and the sample labels of the inappropriate content dataset into the content security detection model to train the content security detection model. Optimize the content security detection model until it meets the preset requirements.
[0025] S5 specifically involves: first, classifying and labeling all data in the automatically collected malicious content dataset and the multivariate high-concealment malicious content dataset; extracting multimodal features from the malicious content dataset; using these multimodal features and sample labels from the automatically collected malicious content dataset and the multivariate high-concealment malicious content dataset to train a content security detection model; obtaining the sensitivity of the content security detection model to different data types based on the classification and feature labeling results; and optimizing the content security detection model based on the sensitivity to different data types. Different data types include text data, image data, audio data, and video data.
[0026] The above detection methods will be described in further detail below.
[0027] To collect a wealth of inappropriate content, it is necessary to use keywords related to such content and employ various technologies such as natural language processing, image recognition, and speech recognition to automatically collect relevant text, images, videos, and audio from the internet (using web crawling technology). Data sources can include platforms such as Weibo, Douyin, WeChat official accounts, converged media, posts, comments, bullet comments, updates, news, user profiles, photo albums, books, songs, short videos, albums, and web pages.
[0028] Specifically, based on multiple classification categories, words are combined and derived using existing large-scale language models. The generated text is then used as keywords to systematically search short videos and blog posts to collect potential bad samples. These potential bad samples constitute an automatically collected dataset of bad content.
[0029] OCR is used to convert text images into machine-readable text formats, while tools such as Stable Diffusion and ChatGPT Vision can extract image information, and ChatGPT and Wenxin Yiyan large-scale language models can extract text summaries. Therefore, this embodiment uses tools such as OCR, Stable Diffusion, and ChatGPT to recognize the collected images and text, and then extracts the main content of the automatically collected inappropriate content dataset as sample labels (i.e., sample descriptions). The dataset of automatically collected inappropriate content is input into a mainstream content security review platform for detection and recognition. The detection results and sample labels of the content security review platform are comprehensively analyzed to extract inappropriate features that can bypass the content security review platform's detection, as well as the combination rules of features, such as text style and layout, image color and elements, etc. In a preferred embodiment, the automatically collected inappropriate content also includes audio and video. The aforementioned content security review platform includes an automatic detection module and a manual detection module.
[0030] AIGC can generate rich text, images, audio, and video, thus AIGC models can be used to generate diverse and highly concealed inappropriate content. When using AIGC models, users typically need to provide an initial text prompt, and prompt engineering is crucial for data generation. Therefore, this invention requires the design of efficient and reliable inappropriate content prompting engineering.
[0031] First, based on the undetected characteristics and combination patterns of inappropriate content, combined with human understanding and annotation, highly concealed features and potential novel combinations of inappropriate content are extracted. These highly concealed features include inappropriate content type, inappropriate content format, tool confidence level, and difficulty for human differentiation. Then, prompts for inappropriate content are designed and input into tools such as ChatGPT and Stable Diffusion, generating a large amount of text, images, and videos containing subtle, implicit inappropriate information. Finally, the generated inappropriate content is tested on a content security review platform. The platform's detection results are compared and analyzed, including the understanding of platform personnel, to filter out inappropriate content that can bypass the content security review platform, thus constructing a multi-dimensional, highly concealed dataset of inappropriate content. The specific process is as follows:
[0032] 1) AIGC Analysis Tools: This invention analyzes the generation strategies and network structures of mainstream AIGC tools (such as Stable Diffusion, MidJourney, DALL·E, etc.). For different types of data (text, images, audio, video, etc.), it delves into their generation logic and patterns, analyzes the correspondence between prompts and generation results, and conducts relevant experiments, especially for obscure, negative, and violent prompts, recording information such as tags and captions of inappropriate content.
[0033] 2) AIGC-generated content: Based on the characteristics of the malicious content dataset obtained by automatically collecting malicious content and the given malicious content settings, a prompting project for malicious content generated by AIGC is constructed. This includes the mutation and combination of malicious features, and through multiple iterations, a dataset of malicious content with extremely high concealment and diversity is constructed using AIGC tools.
[0034] 3) Attacking Mainstream Content Security Review Platforms: Test datasets of inappropriate content generated by AIGC models on mainstream content security review platforms and conduct manual reviews to test the capability of AIGC-generated inappropriate content to attack these platforms. Observe the judgment results and specific descriptions of the content security review platforms, and combine them with the results of manual analysis to extract key characteristics of successful attacks. By comprehensively comparing the output of the content security review platforms and the results of manual analysis, improve the attack capability, concealment, and accuracy of AIGC-generated inappropriate content.
[0035] The specific process of optimizing the content security detection model is as follows: the dataset of inappropriate content generated based on the AIGC model is detected on the content security review model; the data in the dataset of inappropriate content that has not been detected by the content security detection platform is analyzed, and multimodal features of highly concealed inappropriate content are extracted based on the analysis results.
[0036] Based on the multimodal characteristics and universal patterns of highly concealed harmful content, we optimize the harmful content detection model to address diverse and highly concealed harmful features.
[0037] In this phase, the invention will explore security review technology for multimodal content to better address attacks from diverse malicious content and ensure that the content security monitoring model can maintain efficient detection capabilities in diverse data scenarios.
[0038] 1) Multimodal data preparation: The data includes not only malicious datasets collected from the Internet, but also malicious content datasets generated by AIGC models, including text, images, audio, video, etc., to cover as much diverse content in the real world as possible. That is, the malicious content dataset obtained by automatically collecting malicious content and the multi-dimensional, highly concealed malicious content dataset together constitute the multimodal malicious content dataset.
[0039] 2) Multimodal Feature Fusion: The text, images, element positions, and styles from the multimodal inappropriate content dataset are classified and labeled with features. This dataset is then input into a content security detection model. By analyzing the model's response to the multimodal inappropriate content dataset, we can gain a deeper understanding of the model's sensitivity to different data types and types of inappropriate content. Furthermore, by comprehensively analyzing multiple features and combining them with the main rules, data labels, and descriptions used in the data generation phase's prompting process, we can design a multimodal feature fusion scheme for inappropriate content.
[0040] 3) Optimize the content security monitoring model: Based on sensitivity analysis of the content security detection model, optimization and processing are performed on the model's performance under multimodal data. Simultaneously, this invention further enhances the content security detection model using reinforcement learning methods based on the multimodal characteristics of inappropriate content. By comparing the descriptions of the output and input data, similarity assessment and inappropriateness detection are performed in conjunction with models such as NLP (Natural Language Processing). This continuous optimization process will help make the content security detection model more robust and reliable in responding to attacks from inappropriate content.
[0041] This invention discloses a method for detecting inappropriate content based on generative artificial intelligence. It automatically collects datasets and drives the generation of a large amount of diverse and highly concealed inappropriate content datasets based on AIGC (AI Generative Content Generation). The multimodal features of these datasets are then analyzed to train a content security detection model with high detection efficiency and accuracy, capable of detecting highly concealed and diverse inappropriate content. Using this content security detection model eliminates the need for extensive manual review, reducing costs, alleviating the burden on manual reviewers, and protecting the mental and physical health of reviewers.
[0042] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for detecting objectionable content based on generative artificial intelligence driving, characterized in that, Includes the following steps: S1. Automatically collect inappropriate content based on keywords and / or key words to obtain the first inappropriate content dataset; S2. Input the first dataset of inappropriate content into the content security review platform for detection to obtain the first detection result, extract the sample labels of the data in the first dataset of inappropriate content, and obtain the inappropriate features and their combination rules that were not detected by the content security review platform based on the first detection result and the sample labels. S3. By manually analyzing the inappropriate features and their combination rules obtained in S2, highly concealed inappropriate features and novel combination methods of inappropriate content are obtained. Based on the inappropriate features and their combination rules obtained in S2, as well as the highly concealed inappropriate features and novel combination methods of inappropriate content, a prompting project based on AIGC-generated inappropriate content is constructed. The novel combination method is a combination method that is different from the combination method of inappropriate content in the first inappropriate content dataset. S4. According to the prompting process, use the AIGC model to generate several texts, images and videos of inappropriate content, input the texts, images and videos of inappropriate content into the content security review platform for detection to obtain a second detection result, obtain the inappropriate content generated by the AIGC model that has not been detected by the content security review platform based on the second detection result, and construct a multi-dimensional, highly concealed inappropriate content dataset accordingly. S5. Using the first dataset of inappropriate content and the multivariate dataset of highly concealed inappropriate content, extract the multimodal features of the inappropriate data; Based on the sample labels and multimodal features of the first dataset of inappropriate content and the dataset of inappropriate content with high concealment, a content security detection model is trained; the content security detection model is optimized until it meets preset requirements.
2. The method for detecting inappropriate content based on generative artificial intelligence as described in claim 1, characterized in that, The specific process of training the content security detection model using the first dataset of inappropriate content and the multivariate dataset of highly concealed inappropriate content is as follows: First, classify and label all the data in the first dataset of inappropriate content and the multivariate dataset of highly concealed inappropriate content. Analyze the multimodal features and sample labels of the first dataset of inappropriate content and the multivariate dataset of highly concealed inappropriate content, and input them into the content security detection model to train the content security detection model.
3. The method for detecting inappropriate content based on generative artificial intelligence as described in claim 2, characterized in that, The optimization of the content security detection model specifically involves: obtaining the sensitivity of the content security detection model to different data types based on the classification and feature annotation results, and optimizing the content security detection model based on the sensitivity.
4. The method for detecting inappropriate content based on generative artificial intelligence as described in claim 1, characterized in that, The detection method also includes the steps of optimizing the AIGC model and generating a prompting process for inappropriate content based on the AIGC model.
5. The method for detecting inappropriate content based on generative artificial intelligence as described in claim 4, characterized in that, The specific process of optimizing the AIGC model is as follows: the dataset of inappropriate content generated by the AIGC model is detected on a content security review platform; the data in the dataset of inappropriate content that was not detected by the content security review platform is analyzed, and the AIGC model and the prompting process for inappropriate content generated based on the analysis results are optimized.
6. The method for detecting inappropriate content based on generative artificial intelligence as described in claim 1, characterized in that, The content security auditing platform includes an automatic detection module and a manual detection module.