Multi-mode private harmful content detection method and device, electronic equipment and storage medium

By constructing a multimodal implicit and harmful content sample set and training a security detection model, the problem that the multimodal security detection model is unable to identify implicit and harmful content is solved, and accurate detection and identification of implicit and harmful content is achieved.

CN120744566APending Publication Date: 2025-10-03TSINGHUA UNIVERSITY
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510636316.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-10-03

Smart Images

  • Figure CN120744566A_ABST
    Figure CN120744566A_ABST
Patent Text Reader

Abstract

The invention provides a multi-mode private and harmful content detection method and device, electronic equipment and a storage medium. The method comprises the steps that multi-mode content to be detected is acquired; based on a pre-trained security detection model, according to the multi-modal content, obtaining a detection result of the multi-modal obscure harmful content; wherein the security detection model is obtained by performing training optimization according to a multi-mode obscure content sample set, and the multi-mode obscure content sample set comprises a multi-mode obscure content data pair, a security label corresponding to the multi-mode obscure content data pair, a security risk category and risk reasoning data. According to the method, harmful content detection is carried out on the multi-modal content through the safety detection model, harmful information generated after association of all single-modal contents can be better found, the obscure harmful content is recognized, and accurate recognition of the multi-modal obscure harmful content is achieved; meanwhile, the innovatively constructed multi-mode private harmful content sample set provides high-quality data resources for harmful content detection and research.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of content security detection, and in particular to a method, device, electronic device and storage medium for detecting multimodal hidden and harmful content. Background Art

[0002] With the rapid development of the internet and digital communication technologies, multimodal content, combining images and text, is becoming increasingly abundant in cyberspace. Compared to single-modal content, the association of images and text gives multimodal content a stronger ability to express semantics, allowing the combination of images and text to convey more complex messages. However, this characteristic can also be exploited to conceal harmful content, a phenomenon known as "hidden harmful content"—images and text that meet safety standards when viewed individually may convey harmful meanings when combined.

[0003] Multimodal, implicitly harmful content can exist within multimodal statements expressing specific viewpoints or positions on social platforms. It can also be fed into large vision-language models (LVLMs) as multimodal prompts, inducing unsafe responses and fostering multimodal conversations with inherent risks. Multimodal statements, multimodal prompts, and multimodal conversations are three typical forms of multimodal implicitly harmful content. Because it's difficult to identify in a single modality, traditional detection mechanisms struggle to address it, easily leading to the uncontrolled spread of harmful content. Therefore, effectively identifying and preventing this type of implicitly harmful content has become a crucial task in maintaining a secure online environment.

[0004] Currently, security detection for multimodal content primarily relies on LVLM and other traditional detection mechanisms. For example, OpenAI's OpenAI Multimodal Moderation service, launched in September 2024, and Meta's LlamaGuard 3 Vision tool are dedicated to identifying explicit risks in multimodal statements and explicitly unsafe content in multimodal prompts and conversations, respectively. However, these existing security detection models are mostly designed to handle explicitly risky graphic and text content, and their ability to identify implicitly harmful content, where graphic and text appear safe individually but produce harmful semantics when combined, is limited.

[0005] Current mainstream multimodal security detection models have several significant shortcomings: First, existing models often rely on content features that directly present offensive, uncivilized, and other explicitly inappropriate information, and lack effective recognition of implicit expressions; second, existing models usually analyze images or text independently, failing to fully consider the new semantics that may arise from the combination of images and text; finally, because implicit harmful content takes various forms and changes rapidly, traditional detection methods based on rules or fixed patterns are difficult to adapt to new and evolving threats.

[0006] Therefore, how to solve the problem that existing multimodal security detection methods are unable to effectively identify and prevent multimodal implicit and harmful content is an important issue that needs to be urgently addressed in the field of content security detection. Summary of the Invention

[0007] The present invention provides a multimodal implicit and harmful content detection method, device, electronic device and storage medium, which are used to overcome the defects of existing multimodal security detection methods that cannot effectively identify and prevent multimodal implicit and harmful content, and realize accurate identification of multimodal implicit and harmful content.

[0008] On the one hand, the present invention provides a method for detecting multimodal implicit and harmful content, comprising: obtaining multimodal content to be detected; obtaining a multimodal implicit and harmful content detection result based on the multimodal content based on a pre-trained security detection model; wherein the security detection model is obtained by training and optimizing based on a multimodal implicit and harmful content sample set, and the multimodal implicit and harmful content sample set includes multimodal implicit and harmful content data pairs and their corresponding security labels, security risk categories, and risk inference data.

[0009] Furthermore, constructing the multimodal implicitly harmful content sample set specifically includes: determining a preset security risk category, and generating a risk instance based on the preset security risk category; performing cross-modal decomposition on the risk instance to obtain the multimodal implicitly harmful content data pair; generating the risk inference data based on the multimodal implicitly harmful content data pair and its corresponding security risk category; constructing the multimodal implicitly harmful content sample set based on the multimodal implicitly harmful content data pair and its corresponding security label, the security risk category and the risk inference data.

[0010] Furthermore, the risk instance is cross-modally decomposed to obtain the multimodal implicit harmful content data pair, which then includes: checking the unimodal security and multimodal implicit harmfulness of the multimodal implicit harmful content data pair to obtain an inspection result; if the inspection result is qualified, retaining the multimodal implicit harmful content data pair; if the inspection result is unqualified, modifying the multimodal implicit harmful content data pair and cross-validating it, and retaining the multimodal implicit harmful content data pair after cross-validation.

[0011] Furthermore, the risk reasoning data is generated according to the multimodal implicit harmful content data pairs and their corresponding security risk categories, including: based on the visual language big model, the risk reasoning data is generated according to the multimodal implicit harmful content data pairs and their corresponding security risk categories, as well as structured reasoning step prompts.

[0012] Furthermore, the structured reasoning step prompts include: independently checking the behavior and intention conveyed by each single modal content in the multimodal implicit and harmful content data pair; collaboratively reasoning about the consequences of the association between each single modal content in the multimodal implicit and harmful content data pair to obtain a cross-modal association reasoning result; and judging whether there is implicit and harmful content based on the cross-modal association reasoning result.

[0013] Furthermore, training the security detection model specifically includes: using the multimodal implicit harmful content data pair as model input, using the implicit harmful content detection result as model output, using the risk inference data, the security label and the security risk category as supervision signals, iteratively optimizing the security detection model to obtain a security detection model trained to convergence; wherein, the implicit harmful content detection result includes the risk inference result, the predicted security label and the predicted security risk category.

[0014] In a second aspect, the present invention also provides a multimodal implicit and harmful content detection device, comprising: a multimodal content acquisition module, for acquiring multimodal content to be detected; a multimodal implicit and harmful content detection module, for acquiring multimodal implicit and harmful content detection results based on the multimodal content based on a pre-trained security detection model; wherein the security detection model is obtained by training and optimizing based on a multimodal implicit and harmful content sample set, and the multimodal implicit and harmful content sample set includes multimodal implicit and harmful content data pairs and their corresponding security labels, security risk categories, and risk inference data.

[0015] In a third aspect, the present invention further provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the multimodal implicit and harmful content detection method as described above is implemented.

[0016] In a fourth aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the multimodal implicit and harmful content detection methods described above.

[0017] In a fifth aspect, the present invention further provides a computer program product, comprising a computer program, which, when executed by a processor, implements any of the multimodal implicit and harmful content detection methods described above.

[0018] The multimodal implicit harmful content detection method provided by the present invention obtains multimodal content to be detected and, based on a pre-trained security detection model, obtains multimodal implicit harmful content detection results based on the multimodal content. The security detection model is trained and optimized based on a multimodal implicit harmful content sample set, which includes multimodal implicit harmful content data pairs and their corresponding security labels, security risk categories, and risk inference data. This method uses the security detection model to perform harmful content detection on multimodal content, enabling better discovery of harmful information generated by the association of individual modal contents, identification of implicit harmful content, and accurate identification of multimodal implicit harmful content. Furthermore, the innovatively constructed multimodal implicit harmful content sample set provides a high-quality data resource for harmful content detection research. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0020] Figure 1 4 is a flowchart of a multimodal method for detecting implicit and harmful content provided by an embodiment of the present invention.

[0021] Figure 2 2 is a schematic diagram of the structure of a multimodal implicit and harmful content detection device provided by an embodiment of the present invention.

[0022] Figure 3 It is a schematic diagram of the physical structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0023] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0024] It's important to note that currently mainstream security detection models are mostly designed for graphic and textual content with explicit risks. For example, OpenAI Multimodal Moderation, a graphic and textual security detection service launched by OpenAI in September 2024, primarily identifies explicit risks in multimodal presentations. Meta's Llama Guard 3 Vision, designed for multimodal prompts and conversations, focuses on detecting explicitly unsafe content. However, these models generally struggle to identify implicitly harmful content—text and images that appear safe individually but produce harmful semantics when combined—and therefore are unable to effectively address this potential threat.

[0025] In view of this, the present invention proposes a multimodal implicit harmful content detection method, specifically, Figure 1 A schematic flow chart of a multimodal method for detecting implicit and harmful content provided by an embodiment of the present invention is shown.

[0026] like Figure 1 As shown, the method includes: S110, obtaining multimodal content to be detected; S120, obtaining a multimodal implicit harmful content detection result according to the multimodal content based on a pre-trained security detection model; wherein the security detection model is obtained by training and optimizing according to a multimodal implicit harmful content sample set, and the multimodal implicit harmful content sample set includes multimodal implicit harmful content data pairs and their corresponding security labels, security risk categories and risk inference data.

[0027] The following will describe steps S110 - S120 and related steps in detail.

[0028] S110: Acquire multimodal content to be detected.

[0029] It is easy to understand that the multimodal content to be detected in this step can exist in the form of multimodal statements, multimodal prompts, or multimodal dialogues. Among them, multimodal statements refer to posts or comments posted by users on social platforms or forums that express specific views or positions through multimodal combinations. Multimodal prompts refer to multimodal combination instructions or questions input to the visual language model. Multimodal dialogues refer to the dialogue content with the model obtained after inputting multimodal prompts into the visual language model. The types of "multimodal" in multimodal content include, but are not limited to, visual modalities, textual modalities, and audio modalities. Visual modalities include images and videos. Images include photographs and illustrations, while videos include dynamic image sequences. Textual modalities include textual content, such as articles, reviews, conversation logs, and any other form of written language. Audio modalities include speech, such as recordings of user speech; music, such as various types of musical works and their components; and ambient sound, such as background noise or other non-speech sounds.

[0030] It should be noted that while each unimodal content contained in the multimodal content to be detected should meet security standards individually, the multimodal content formed by combining the unimodal content may convey unsafe or harmful content. This embodiment of the present invention primarily detects potentially harmful content in the multimodal content formed by combining the unimodal content. The specific detection process is implemented in step S120.

[0031] S120, based on a pre-trained security detection model, obtain a multimodal implicit harmful content detection result according to the multimodal content; wherein, the security detection model is obtained by training and optimizing according to a multimodal implicit harmful content sample set, and the multimodal implicit harmful content sample set includes multimodal implicit harmful content data pairs and their corresponding security labels, security risk categories, and risk inference data.

[0032] It's easy to understand that, for the multimodal content to be detected, embodiments of the present invention pre-train a security detection model specifically designed to detect implicit and harmful content within that multimodal content. Specifically, during actual detection, the multimodal content to be detected is input into the pre-trained security detection model, which then outputs a multimodal implicit and harmful content detection result.

[0033] It is worth mentioning that the security detection model in this embodiment performs security detection on the input multimodal content through the logical thinking of "thinking and reasoning first, then safety judgment". Specifically, the security detection model will first perform cross-modal association reasoning on the multimodal content, respectively identify the intention of each single modal content in the multimodal content and the harmful consequences that may arise after the association of each single modal content, and then give the multimodal content security judgment label and the corresponding security risk category. In this process, the cross-modal association reasoning results can be used as an analysis and interpretation of the final multimodal implicit harmful content detection results, thereby improving the interpretability of the multimodal implicit harmful content detection results.

[0034] The multimodal implicit harmful content detection results include target reasoning analysis results, target security labels, and target security risk categories.

[0035] It should be noted that the security detection model in this embodiment is pre-trained. Specifically, a high-quality, multimodal sample set of implicit and harmful content is first constructed through innovative human-machine collaboration, providing high-quality data resources for harmful content detection research. This multimodal sample set of implicit and harmful content is then used to iteratively optimize the security detection model, resulting in a trained security detection model. The process of constructing the multimodal sample set of implicit and harmful content and training the security detection model will be detailed in the embodiments below.

[0036] Among them, the multimodal implicit harmful content sample set includes four parts, namely, multimodal implicit harmful content data pairs and their corresponding security labels, security risk categories and risk reasoning data.

[0037] A multimodal implicitly harmful content data pair refers to data where each modal content, when viewed individually, meets safety standards, but when combined, conveys unsafe information. A multimodal implicitly harmful content data pair can encompass two or more modal types, without specific limitations. For example, a multimodal implicitly harmful content data pair might include both a textual modality and a visual modality, such as "textual modality: video recording of a moment; visual modality: a movie theater showing a movie." While each modal content may appear safe when viewed individually, when combined, it may convey "illegal video recording inside a movie theater," conveying unsafe information.

[0038] Safety labels refer to the safety of multimodal, implicitly harmful content data pairs, such as safe or unsafe. Security risk categories include, but are not limited to, seven types of harmful content: offense and insult, prejudice and discrimination, physical harm, illegal behavior, immorality, privacy violation, and false information. Risk inference data includes the inference analysis process conducted on multimodal, implicitly harmful content data pairs.

[0039] In this embodiment, multimodal content to be detected is obtained, and based on the multimodal content, a multimodal implicit harmful content detection result is obtained based on the multimodal content. The security detection model is trained and optimized based on a multimodal implicit harmful content sample set, which includes multimodal implicit harmful content data pairs and their corresponding security labels, security risk categories, and risk inference data. This method uses the security detection model to perform harmful content detection on multimodal content, enabling better discovery of harmful information generated by the association of individual modal contents, identification of implicit harmful content, and accurate identification of multimodal implicit harmful content. Furthermore, the innovatively constructed multimodal implicit harmful content sample set provides a high-quality data resource for harmful content detection research.

[0040] On the basis of the above embodiments, the following will further describe in detail the process of constructing a multimodal implicit and harmful content sample set by taking the multimodality of text modality-image modality association as an example.

[0041] Constructing a multimodal implicit and harmful content sample set specifically includes: determining preset security risk categories, and generating risk instances based on the preset security risk categories; performing cross-modal decomposition on the risk instances to obtain multimodal implicit and harmful content data pairs; generating risk inference data based on the multimodal implicit and harmful content data pairs and their corresponding security risk categories; constructing a multimodal implicit and harmful content sample set based on the multimodal implicit and harmful content data pairs and their corresponding security labels, security risk categories and risk inference data.

[0042] As will be readily understood, to better train the security detection model, this embodiment innovatively designs a method for constructing a multimodal sample set of implicitly harmful content. Based on the concept of "cross-modal risk element decomposition," this embodiment utilizes a large visual language model to map the risk elements of specific harmful behaviors or scenarios to each single modality (e.g., text modality, image modality). Through manual review and screening, this method generates a set of multimodal statements and prompts that implicitly convey harmful information, namely, multimodal implicitly harmful content data pairs. Based on these multimodal implicitly harmful content data pairs, corresponding risk inference data is generated to obtain the final multimodal sample set of implicitly harmful content.

[0043] Specifically, first, the preset security risk categories are determined. In this embodiment, the security risk categories cover seven typical security risk categories and their corresponding 31 subcategories. The seven typical security risk categories include offense and insult, prejudice and discrimination, physical harm, illegal behavior, immorality, privacy invasion, and false information. The 31 subcategories corresponding to these seven typical security risk categories include personal attacks, malicious ridicule, disrespectful remarks, age discrimination, threats of violence, fraudulent information, copyright infringement, phone number disclosure, home address exposure, false news, misleading advertising, rumor spreading, etc., which will not be detailed here.

[0044] Then, risk instances are generated based on the preset security risk categories. In order to ensure data diversity, it is necessary to make the samples under each security risk category cover as many risk instances as possible. To this end, this embodiment further refines the above-mentioned 7 security risk categories and expands a group of typical harmful behaviors or scenarios under each category. During the expansion process, a structured prompt method can be adopted: first guide the visual language model to generate specific security risk subcategories under the security risk category, and then guide the visual language model to generate representative and diverse risk instances.

[0045] For example, under the security risk category “false information”, the corresponding content generated by the guided visual language model is as follows (1)-(3).

[0046] (1) Fake news. Risk example: A news article posted on a social media platform claimed that a well-known local restaurant used expired ingredients, but did not provide any evidence to support the claim. The report attracted widespread public attention and caused a significant drop in the restaurant's business.

[0047] (2) Spread of rumors. Risk example: During a regional infectious disease outbreak, a rumor began circulating online claiming that drinking a certain brand of mineral water could prevent the disease. This unverified rumor spread quickly, causing a rush to buy that brand of mineral water and a surge in its price.

[0048] (3) Misleading advertising. Risk example: A weight loss product claims on its official website that users can lose 10 kg in two weeks without changing their diet or increasing their exercise. In fact, the effectiveness of this product has not been scientifically verified, and long-term use may be harmful to health.

[0049] It should be noted that the large visual language model in this embodiment, such as GPT-4o, is not specifically limited here.

[0050] After generating risk instances, the risk instances are cross-modally decomposed to obtain multimodal implicit and harmful content data pairs. Specifically, based on the above risk instances, instructions can be designed to guide large-scale visual language models to decompose each risk instance into image descriptions and text content, dispersing potentially harmful semantics into different modalities, making the image descriptions and text content individually safe, but conveying harmful content when combined. For example, the risk instance "illegal video recording in a cinema" can be decomposed into "image: a cinema showing a movie, text: the video records wonderful moments", which may induce illegal behavior after being combined. Subsequently, a diffusion model can be used to generate corresponding images based on the image descriptions to construct image-text pairings to form preliminary multimodal implicit and harmful content data pairs.

[0051] After obtaining preliminary multimodal data pairs containing potentially harmful content, manual correction and screening are performed. Specifically, data annotators examine the single-modal security (text modality security and image modality security) and multimodal potential harmfulness of the multimodal data pairs, generating corresponding inspection results.

[0052] If the inspection result is satisfactory, meaning the multimodal implicit and harmful content data pair meets the requirements, then the multimodal implicit and harmful content data pair is retained. Conversely, if the inspection result is unsatisfactory, meaning the multimodal implicit and harmful content data pair does not meet the requirements, then manual modifications are made and cross-validated by different data annotators. The multimodal implicit and harmful content data pairs that meet the requirements after cross-validation are retained. This results in the multimodal implicit and harmful content data pairs that ultimately participate in the training of the security detection model.

[0053] To better guide the model's identification of implicit and harmful content in images and text, this embodiment also adds risk inference data to the security detection model's training sample set. Specifically, the aforementioned multimodal implicit and harmful content data pairs and their corresponding security risk categories are used as input to the visual language model, which generates corresponding risk inference data based on structured reasoning steps.

[0054] Among them, the structured reasoning steps include: (1) independently checking the behaviors and intentions conveyed by each single modal content in the multimodal implicit harmful content data pair; (2) collaboratively reasoning about the consequences of the association between each single modal content in the multimodal implicit harmful content data pair to obtain the cross-modal association reasoning result; (3) judging whether there is implicit harmful content based on the cross-modal association reasoning result.

[0055] Based on the above, a multimodal implicit harmful content sample set can be constructed. Each multimodal implicit harmful content sample includes four parts, namely, a multimodal implicit harmful content data pair and its corresponding security label, security risk category, and risk reasoning data.

[0056] In this embodiment, by determining preset security risk categories and generating risk instances based on the preset security risk categories, and then performing cross-modal decomposition on the risk instances, multimodal implicit harmful content data pairs are obtained. Risk inference data is then generated based on the multimodal implicit harmful content data pairs and their corresponding security risk categories. Thus, a multimodal implicit harmful content sample set is constructed based on the multimodal implicit harmful content data pairs and their corresponding security labels, security risk categories, and risk inference data. This method trains and optimizes a security detection model using the constructed multimodal implicit harmful content sample set, enabling the security detection model to effectively identify implicit harmful content in multimodal content. Furthermore, the innovatively constructed multimodal implicit harmful content sample set provides a high-quality data resource for harmful content detection research.

[0057] On the basis of the above embodiments, the training optimization process of the security detection model will be further described in detail below.

[0058] Training a security detection model specifically includes: using multimodal implicit and harmful content data pairs as model input, implicit and harmful content detection results as model output, risk reasoning data, security labels, and security risk categories as supervision signals, iteratively optimizing the security detection model, and obtaining a security detection model that is trained to convergence; wherein the implicit and harmful content detection results include risk reasoning results, predicted security labels, and predicted security risk categories.

[0059] It's easy to understand that each sample in the constructed multimodal implicit and harmful content sample set includes a multimodal implicit and harmful content data pair and its corresponding security label, security risk category, and risk inference data. Training and optimization of the security detection model involves iteratively optimizing the security detection model using multiple samples from the multimodal implicit and harmful content sample set.

[0060] Specifically, each time the security detection model is trained and optimized, the multimodal implicit and harmful content data is input into the initialized security detection model. The security detection model generates risk reasoning results, predicted safety labels, and predicted security risk categories in sequence according to the logical idea of ​​"thinking and reasoning first, then safety judgment". The generated risk reasoning results, predicted safety labels, and predicted security risk categories use the risk reasoning data, safety labels, and security risk categories in the multimodal implicit and harmful content sample set as supervision signals to guide the security detection model to learn how to understand and analyze the implicit correlations between each single modal content, thereby detecting implicit and harmful content.

[0061] For a trained security detection model, the multimodal content to be detected is directly input into the security detection model, which will automatically perform the reasoning and analysis process and generate explainable multimodal implicit and harmful content detection results.

[0062] In this embodiment, by using multimodal implicit and harmful content data pairs as model input, implicit and harmful content detection results as model output, and risk reasoning data, security labels, and security risk categories as supervision signals, the security detection model is iteratively optimized, so that the security detection model can better discover the harmful information generated after the association of each single modal content, identify implicit and harmful content, and achieve accurate identification of multimodal implicit and harmful content; at the same time, the innovatively constructed multimodal implicit and harmful content sample set used to train the security detection model provides high-quality data resources for harmful content detection research.

[0063] In some embodiments, the visual language large model described in the above embodiments is built based on GPT-4o, and the security detection model is built based on qwenVL2-7b.

[0064] GPT-4o is a large multimodal language model developed by OpenAI (supporting multimodal inputs such as text, images, and files). Its core capabilities lie in multimodal collaborative understanding, logical reasoning, and complex task automation. It demonstrates the potential for cross-modal reasoning through joint training of text and images. qwenVL2-7b is a large multimodal model based on the Transformer architecture. It supports multimodal inputs such as text, images, and videos, and has multimodal joint reasoning and generation capabilities. Its parameter size reaches 7 billion (7B).

[0065] Of course, the visual language model and the security detection model can also be other qualified architectures, and are not specifically limited here.

[0066] In some other embodiments, the overall process of the multimodal implicit and harmful content detection method provided by the embodiments of the present invention will be described in detail below.

[0067] First, a multimodal implicit harmful content sample set is constructed: the preset security risk categories are determined, and risk instances are generated based on the preset security risk categories; the risk instances are cross-modally decomposed to obtain multimodal implicit harmful content data pairs; the unimodal security and multimodal implicit harmfulness of the multimodal implicit harmful content data pairs are inspected to obtain inspection results; if the inspection results are qualified, the multimodal implicit harmful content data pairs are retained; if the inspection results are unqualified, the multimodal implicit harmful content data pairs are modified and cross-validated, and the cross-validated multimodal implicit harmful content data pairs are retained; risk inference data are generated based on the multimodal implicit harmful content data pairs and their corresponding security risk categories; a multimodal implicit harmful content sample set is constructed based on the multimodal implicit harmful content data pairs and their corresponding security labels, security risk categories, and risk inference data.

[0068] Then, the security detection model is trained and optimized: multimodal implicit and harmful content data pairs are used as model input, implicit and harmful content detection results are used as model output, and risk reasoning data, security labels, and security risk categories are used as supervision signals. The security detection model is iteratively optimized to obtain a security detection model that is trained to convergence; among them, the implicit and harmful content detection results include risk reasoning results, predicted security labels, and predicted security risk categories.

[0069] Finally, the security detection model is used to detect implicit and harmful content in multimodal content: the multimodal content to be detected is obtained; the multimodal content to be detected is input into the pre-trained security detection model to obtain the output multimodal implicit and harmful content detection results. The multimodal implicit and harmful content detection results include the target reasoning analysis results, the target security label and the target security risk category in sequence.

[0070] Corresponding to the multimodal implicit and harmful content detection method described in the above embodiments, the present invention also provides a multimodal implicit and harmful content detection device.

[0071] Specifically, Figure 2 FIG2 shows a schematic structural diagram of a multimodal implicit and harmful content detection device provided by an embodiment of the present invention.

[0072] like Figure 2 As shown, the device includes: a multimodal content acquisition module 210, which is used to obtain multimodal content to be detected; a multimodal implicit harmful content detection module 220, which is used to obtain multimodal implicit harmful content detection results according to the multimodal content based on a pre-trained security detection model; wherein the security detection model is obtained by training and optimizing according to a multimodal implicit harmful content sample set, and the multimodal implicit harmful content sample set includes multimodal implicit harmful content data pairs and their corresponding security labels, security risk categories and risk inference data.

[0073] In this embodiment, the multimodal content acquisition module 210 acquires multimodal content to be detected. The multimodal implicit harmful content detection module 220 obtains multimodal implicit harmful content detection results based on the multimodal content, based on a pre-trained security detection model. The security detection model is trained and optimized using a multimodal implicit harmful content sample set, which includes pairs of multimodal implicit harmful content data and their corresponding security labels, security risk categories, and risk inference data. This device uses the security detection model to perform harmful content detection on multimodal content, enabling better discovery of harmful information generated by the association of individual modal contents and identification of implicit harmful content, thus achieving accurate identification of multimodal implicit harmful content. Furthermore, the innovatively constructed multimodal implicit harmful content sample set provides a high-quality data resource for harmful content detection research.

[0074] It should be noted that the multimodal implicit harmful content detection apparatus provided in the embodiment of the present invention and the multimodal implicit harmful content detection method described in the above embodiments can be referred to in correspondence with each other, and will not be described in detail here.

[0075] Figure 3 An example of a physical structure diagram of an electronic device is shown below. Figure 3As shown, the electronic device may include: a processor 310, a communications interface 320, a memory 330, and a communications bus 340. The processor 310, the communications interface 320, and the memory 330 communicate with each other via the communications bus 340. The processor 310 may invoke logic instructions in the memory 330 to execute a method for detecting multimodal implicit and harmful content. The method includes: obtaining multimodal content to be detected; and obtaining a multimodal implicit and harmful content detection result based on the multimodal content based on a pre-trained security detection model. The security detection model is obtained by training and optimizing a multimodal implicit and harmful content sample set, wherein the multimodal implicit and harmful content sample set includes multimodal implicit and harmful content data pairs and their corresponding security labels, security risk categories, and risk inference data.

[0076] Furthermore, the logic instructions in the aforementioned memory 330 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0077] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the multimodal implicit harmful content detection method provided by the above-mentioned methods, which includes: obtaining multimodal content to be detected; based on a pre-trained security detection model, obtaining a multimodal implicit harmful content detection result according to the multimodal content; wherein, the security detection model is obtained by training and optimizing according to a multimodal implicit harmful content sample set, and the multimodal implicit harmful content sample set includes multimodal implicit harmful content data pairs and their corresponding security labels, security risk categories and risk inference data.

[0078] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the multimodal implicit harmful content detection method provided by the above-mentioned methods, the method comprising: obtaining multimodal content to be detected; obtaining a multimodal implicit harmful content detection result according to the multimodal content based on a pre-trained security detection model; wherein the security detection model is obtained by training and optimizing based on a multimodal implicit harmful content sample set, and the multimodal implicit harmful content sample set includes multimodal implicit harmful content data pairs and their corresponding security labels, security risk categories, and risk inference data.

[0079] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0080] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0081] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A multimodal method for detecting implicit and harmful content, characterized in that: include: Obtain multimodal content to be detected; Obtaining a multimodal implicit and harmful content detection result based on the multimodal content based on a pre-trained security detection model; The security detection model is obtained by training and optimizing based on a multimodal implicit harmful content sample set, wherein the multimodal implicit harmful content sample set includes multimodal implicit harmful content data pairs and their corresponding security labels, security risk categories, and risk inference data.

2. The multimodal implicit harmful content detection method according to claim 1, characterized in that: Constructing the multimodal implicit and harmful content sample set specifically includes: Determining a preset security risk category and generating a risk instance based on the preset security risk category; Performing cross-modal decomposition on the risk instance to obtain the multi-modal implicit and harmful content data pair; generating the risk inference data according to the multimodal implicit harmful content data pairs and their corresponding security risk categories; The multimodal implicit harmful content sample set is constructed based on the multimodal implicit harmful content data pairs and their corresponding security labels, the security risk categories, and the risk inference data.

3. The multimodal implicit harmful content detection method according to claim 2, characterized in that: The cross-modal decomposition of the risk instance to obtain the multi-modal implicit harmful content data pair then includes: Checking the unimodal security and multimodal implicit harmfulness of the multimodal implicit harmful content data pair to obtain a check result; If the inspection result is qualified, retaining the multimodal implicit and harmful content data pair; If the inspection result is unqualified, the multimodal implicit and harmful content data pair is modified and cross-validated, and the cross-validated multimodal implicit and harmful content data pair is retained.

4. The multimodal implicit harmful content detection method according to claim 2, characterized in that: Generating the risk inference data according to the multimodal implicit harmful content data pair and its corresponding security risk category includes: Based on the visual language big model, the risk reasoning data is generated according to the multimodal implicit harmful content data pairs and their corresponding security risk categories, as well as structured reasoning step prompts.

5. The multimodal implicit harmful content detection method according to claim 4, characterized in that: The structured reasoning steps include: Independently examine the actions and intentions conveyed by each unimodal content in the multimodal potentially harmful content data pair; Collaboratively reasoning about the consequences of associations between the single-modal content in the multimodal implicit and harmful content data pair to obtain a cross-modal association reasoning result; Determine whether there is obscure and harmful content based on the cross-modal association reasoning result.

6. The multimodal implicit and harmful content detection method according to any one of claims 1 to 5, characterized in that: Training the security detection model specifically includes: Iteratively optimizing the security detection model using the multimodal implicit and harmful content data pair as a model input, implicit and harmful content detection results as a model output, and the risk inference data, the security label, and the security risk category as a supervisory signal to obtain a security detection model trained to convergence; The implicit harmful content detection result includes a risk reasoning result, a predicted safety label, and a predicted safety risk category.

7. A multimodal implicit and harmful content detection device, characterized in that: include: A multimodal content acquisition module, used to acquire multimodal content to be detected; a multimodal implicit and harmful content detection module, configured to obtain a multimodal implicit and harmful content detection result based on the multimodal content based on a pre-trained security detection model; The security detection model is obtained by training and optimizing based on a multimodal implicit harmful content sample set, wherein the multimodal implicit harmful content sample set includes multimodal implicit harmful content data pairs and their corresponding security labels, security risk categories, and risk inference data.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the multimodal implicit and harmful content detection method according to any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the multimodal implicit and harmful content detection method according to any one of claims 1 to 6 is implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the multimodal implicit and harmful content detection method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Implicit toxic text generation method and device based on reinforcement learning

    CN117764037A

  • Modal characterization decomposition method for multi-modal metaphor detection

    CN118673401A

  • Multi-mode error information detection method and system of mixed source

    CN119598382A

  • Social big data cross-modal meta-learning early rumor detection method based on small samples

    CN119739930A

  • Illegal data detection method based on reading understanding

    CN119884368A