Harmful information detection method and device, electronic equipment, storage medium and product

CN122594950APending Publication Date: 2026-08-18TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610252063.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-03
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

[0005]本发明提供一种有害信息检测方法、装置、电子设备、存储介质及产品,以解决相关技术效率与精准度失衡、复杂语义易误判、检测结果无推理过程且不可解释的问题,提升有害信息检测的精准度与检测性能,提高有害信息检测的适用性

Benefits of technology

[0011]The harmful information detection method provided by the present invention first acquires the multimodal data to be detected, performs explicit harmful information detection on it to obtain the explicit harmful information clue results and determines the gate control discrimination result, then conducts discrimination debate on the multimodal data based on the gate control discrimination result to obtain the discrimination debate history, and finally obtains the harmful information detection result of the multimodal data based on the discrimination debate history. This method solves the problems of imbalance between efficiency and accuracy of related technologies, easy misjudgment of complex semantics, and lack of reasoning process and interpretation of detection results, thereby improving the accuracy and performance of harmful information detection and enhancing the applicability of harmful information detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122594950A_ABST
    Figure CN122594950A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of artificial intelligence, in particular to a harmful information detection method and device, electronic equipment, storage medium and product, the method comprising: obtaining multi-modal data to be detected; performing explicit harmful information detection on the multi-modal data to be detected to obtain harmful information explicit clue results of the multi-modal data to be detected, and discriminating the harmful information explicit clue results to obtain a gating discrimination result; based on the gating discrimination result, discriminating the multi-modal data to be detected, obtaining a discrimination argument history of the multi-modal data to be detected, and obtaining harmful information detection results of the multi-modal sample to be detected according to the discrimination argument history of the multi-modal data to be detected, solving the problems of imbalance between efficiency and accuracy, easy misjudgment of complex semantics, no reasoning process and uninterpretable detection results in related technologies, improving the accuracy and detection performance of harmful information detection, and improving the applicability of harmful information detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, electronic device, storage medium, and product for detecting harmful information. Background Technology

[0002] With the development of social media, harmful speech has evolved from simple text to multimodal forms combining text and images. Existing automated detection systems have made some progress in handling overtly harmful speech (such as explicit abusive language or violent images), but they still have significant shortcomings when dealing with subtle multimodal harmful speech.

[0003] In related technologies, the common approach is direct feature fusion, which is single-channel direct detection. It relies on zero-sample multimodal large models to directly perform feature matching or shallow semantic analysis on multimodal data combining images and text.

[0004] However, the relevant technologies have the following problems. First, single-channel processing cannot efficiently handle explicit and simple cases, nor can it deeply analyze implicit and complex content such as metaphors and cultural codes, resulting in an imbalance between efficiency and accuracy. Second, the semantic misjudgment rate is high, and the lack of adversarial reasoning makes it easy to be misled by surface content when dealing with complex semantics such as irony and accusation. Finally, the relevant technologies only output classification labels without reasoning processes, and the results are uninterpretable, resulting in insufficient credibility and traceability of the detection results, which urgently need to be addressed. Summary of the Invention

[0005] This invention provides a method, apparatus, electronic device, storage medium, and product for detecting harmful information, in order to solve the problems of imbalance between efficiency and accuracy in related technologies, easy misjudgment of complex semantics, and detection results without reasoning process and uninterpretable, thereby improving the accuracy and performance of harmful information detection and enhancing the applicability of harmful information detection.

[0006] To achieve the above objectives, a first aspect of the present invention provides a method for detecting harmful information, comprising the following steps: acquiring multimodal data to be detected; performing explicit harmful information detection on the multimodal data to be detected to obtain explicit harmful information clue results of the multimodal data to be detected, and judging the explicit harmful information clue results to obtain a gating judgment result; based on the gating judgment result, performing discrimination debate on the multimodal data to be detected to obtain a discrimination debate history of the multimodal data to be detected, and obtaining the harmful information detection result of the multimodal data to be detected based on the discrimination debate history of the multimodal data to be detected.

[0007] Furthermore, in some embodiments, the step of judging the explicit clue result of the harmful information to obtain the gating judgment result includes: identifying whether the explicit clue result of the harmful information is an explicit clue; if the explicit clue result of the harmful information is an explicit clue, then the gating judgment result is a trial based on the first channel; otherwise, the gating judgment result is a debate based on the second channel; wherein, the processing rate of the first channel is higher than the processing rate of the second channel.

[0008] Furthermore, in some embodiments, the gating judgment result is based on the trial conducted through the first channel. The step of performing discrimination and debate on the multimodal data to be detected based on the gating judgment result to obtain the discrimination and debate history of the multimodal data to be detected includes: detecting the multimodal data to be detected according to a preset prosecutor intelligent agent and generating an indictment corresponding to the multimodal data to be detected; refuting the indictment according to a preset defense intelligent agent and obtaining rebuttal opinions; and generating the discrimination and debate history based on the indictment and the rebuttal opinions.

[0009] Furthermore, in some embodiments, the gating discrimination result is based on the second channel for debate. The process of discriminating and debating the multimodal data to be detected based on the gating discrimination result to obtain the discriminating debate history of the multimodal data to be detected includes: performing clue detection on the multimodal data to be detected according to the preset prosecutor's intelligent agent to obtain a clue detection result; if the clue detection result is empty, a harmless debate history is generated; otherwise, an adversarial debate is conducted through the prosecutor's intelligent agent and the defense intelligent agent to generate a harmful debate history; and the discriminating debate history is determined based on the harmless debate history and the harmful debate history.

[0010] Furthermore, in some embodiments, obtaining the harmful information detection result of the multimodal data to be detected based on the discriminant debate history of the multimodal data to be detected includes: evaluating the multimodal data to be detected and the discriminant debate history of the multimodal data to be detected based on a preset judge agent to obtain harmful information classification labels of the multimodal data to be detected; performing source interpretation on the harmful information classification labels to generate natural language interpretation results corresponding to the harmful information classification labels; and obtaining the harmful information detection result based on the harmful information classification labels and the natural language interpretation results corresponding to the harmful information classification labels.

[0011] The harmful information detection method provided by the present invention first acquires the multimodal data to be detected, performs explicit harmful information detection on it to obtain the explicit harmful information clue results and determines the gate control discrimination result, then conducts discrimination debate on the multimodal data based on the gate control discrimination result to obtain the discrimination debate history, and finally obtains the harmful information detection result of the multimodal data based on the discrimination debate history. This method solves the problems of imbalance between efficiency and accuracy of related technologies, easy misjudgment of complex semantics, and lack of reasoning process and interpretation of detection results, thereby improving the accuracy and performance of harmful information detection and enhancing the applicability of harmful information detection.

[0012] To achieve the above objectives, a second aspect of the present invention provides a harmful information detection device, comprising: an acquisition module for acquiring multimodal data to be detected; a discrimination module for performing explicit harmful information detection on the multimodal data to be detected to obtain explicit harmful information clue results of the multimodal data to be detected, and discriminating the explicit harmful information clue results to obtain a gating discrimination result; and a detection module for performing discrimination debate on the multimodal data to be detected based on the gating discrimination result to obtain a discrimination debate history of the multimodal data to be detected, and obtaining a harmful information detection result of the multimodal data to be detected based on the discrimination debate history of the multimodal data to be detected.

[0013] Furthermore, in some embodiments, the discrimination module is specifically used to: identify whether the result of the harmful information explicit clue is an explicit clue; if the result of the harmful information explicit clue is an explicit clue, then the gating discrimination result is a trial based on the first channel; otherwise, the gating discrimination result is a debate based on the second channel; wherein, the processing rate of the first channel is higher than the processing rate of the second channel.

[0014] Furthermore, in some embodiments, the detection module is specifically used for: detecting the multimodal data to be detected according to a preset prosecutor intelligent agent, and generating an indictment corresponding to the multimodal data to be detected; refuting the indictment according to a preset defense intelligent agent, and obtaining rebuttal opinions; and generating the discrimination debate history based on the indictment and the rebuttal opinions.

[0015] Furthermore, in some embodiments, the detection module is also used to: perform clue detection on the multimodal data to be detected according to the preset prosecutor intelligent agent, and obtain clue detection results; if the clue detection results are empty, generate a harmless debate history; otherwise, conduct adversarial debate through the prosecutor intelligent agent and the defense intelligent agent to generate a harmful debate history; and determine the discriminative debate history based on the harmless debate history and the harmful debate history.

[0016] Furthermore, in some embodiments, the detection module is also used to: evaluate the multimodal data to be detected and the discrimination debate history of the multimodal data to be detected based on a preset judge agent, and obtain a harmful information classification label for the multimodal data to be detected; perform source interpretation on the harmful information classification label and generate a natural language interpretation result corresponding to the harmful information classification label; and obtain the harmful information detection result based on the harmful information classification label and the natural language interpretation result corresponding to the harmful information classification label.

[0017] The harmful information detection device provided in this embodiment of the invention first acquires multimodal data to be detected, performs explicit harmful information detection on it to obtain explicit harmful information clue results and determines the gate control discrimination result, then conducts discrimination debate on the multimodal data based on the gate control discrimination result to obtain the discrimination debate history, and finally obtains the harmful information detection result of the multimodal data based on this discrimination debate history. This solves the problems of imbalance between efficiency and accuracy of related technologies, easy misjudgment of complex semantics, and detection results without reasoning process and uninterpretable, thereby improving the accuracy and detection performance of harmful information detection and enhancing the applicability of harmful information detection.

[0018] To achieve the above objectives, a third aspect of the present invention provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the harmful information detection method as described in the above embodiments.

[0019] To achieve the above objectives, a fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which is executed by a processor to implement the harmful information detection method as described in the above embodiments.

[0020] A fifth aspect of the present invention provides a computer program product, including a computer program that is executed to implement the harmful information detection method as described in the above embodiments.

[0021] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0022] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 A flowchart of a harmful information detection method provided according to an embodiment of the present invention; Figure 2This is a schematic flowchart of a harmful information detection method according to a specific embodiment of the present invention; Figure 3 A schematic diagram illustrating a fast trial process according to a specific embodiment of the present invention; Figure 4 A schematic diagram of an in-depth debate provided according to a specific embodiment of the present invention; Figure 5 This is a block diagram of a harmful information detection device provided according to an embodiment of the present invention; Figure 6 This is a schematic diagram of the structure of an electronic device provided according to an embodiment of the present invention. Detailed Implementation

[0023] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.

[0024] The following describes the method, apparatus, electronic device, storage medium, and product for detecting harmful information according to embodiments of the present invention with reference to the accompanying drawings. First, the method for detecting harmful information according to embodiments of the present invention will be described with reference to the accompanying drawings.

[0025] Figure 1 A flowchart illustrating a harmful information detection method provided according to an embodiment of the present invention.

[0026] like Figure 1 As shown, the method for detecting harmful information includes the following steps: In step S101, the multimodal data to be detected is acquired.

[0027] Specifically, the multimodal data to be detected refers to social media content samples that include both text and images (such as emojis with text, tweets with images, and comments combining text and images). For example, a multimodal sample... , containing text and images The text and images interact semantically through metaphor, irony, or cultural context. They may be harmless individually but form a harmful intent when combined, or appear harmful on the surface but are actually benign.

[0028] In step S102, explicit harmful information detection is performed on the multimodal data to be detected to obtain explicit harmful information clues of the multimodal data to be detected, and the explicit harmful information clues are judged to obtain the gating judgment result.

[0029] Among them, the explicit harmful information cue result refers to the judgment record obtained after the explicit harmful information detection of multimodal samples, and the gating discrimination result refers to the binarized instruction output after the judgment is performed by the gating function.

[0030] Furthermore, in some embodiments, the results of explicit harmful information clues are judged to obtain a gating judgment result, including: identifying whether the results of explicit harmful information clues are explicit clues; if the results of explicit harmful information clues are explicit clues, the gating judgment result is a trial based on the first channel, otherwise, the gating judgment result is a debate based on the second channel; wherein, the processing rate of the first channel is higher than the processing rate of the second channel.

[0031] Specifically, the system first performs a rapid scan of the multimodal data sample to be tested, input by the prosecutor's intelligent agent team, to detect whether the multimodal data sample contains explicit harmful symbols or abusive words. For example, if explicit clues are detected, a gating function is then activated. Output 1 indicates that the gating decision result is to enter the fast trial channel (i.e., the trial is conducted based on the first channel). If no explicit clues are detected, the gating function... Output 0, the gating result is to enter the deep debate channel (i.e. debate is based on the second channel).

[0032] In step S103, based on the gating discrimination result, discrimination and debate are performed on the multimodal data to be detected to obtain the discrimination and debate history of the multimodal data to be detected, and the harmful information detection result of the multimodal data to be detected is obtained based on the discrimination and debate history of the multimodal data to be detected.

[0033] Furthermore, in some embodiments, the gating judgment result is based on the first channel for trial. Based on the gating judgment result, the multimodal data to be detected is used for discrimination and debate to obtain the discrimination and debate history of the multimodal data to be detected, including: detecting the multimodal data to be detected according to the preset prosecutor intelligent agent and generating the indictment corresponding to the multimodal data to be detected; refuting the indictment according to the preset defense intelligent agent and obtaining the rebuttal opinion; and generating the discrimination and debate history based on the indictment and the rebuttal opinion.

[0034] The indictment refers to data generated by a pre-set prosecutor AI that clearly points out specific explicit harmful clues in the data (such as abusive words, violent images and symbols), the corresponding attack targets, and the specific reasons why the content constitutes harmful information. The rebuttal refers to data generated by a pre-set defense AI that makes targeted defenses against the indictment issued by the prosecutor AI from the perspectives of context, purpose of use, etc.

[0035] Specifically, the prosecutor's intelligent agent directly generates an indictment based on the sample. The defense should identify specific harmful keywords or symbols, their targets, and the reasons for the attacks. Immediately provide contextual rebuttal (e.g., point out that this is quoting harmful language and criticize it), generating rebuttal opinions. Finally, according to the indictment and rebuttal opinions Generate debate history .

[0036] Furthermore, in some embodiments, the gating judgment result is based on the second channel for debate. Based on the gating judgment result, the multimodal data to be detected is subjected to a discriminative debate to obtain the discriminative debate history of the multimodal data to be detected, including: according to the preset prosecutor's intelligent agent, the multimodal data to be detected is subjected to clue detection to obtain clue detection results; if the clue detection results are empty, a harmless debate history is generated; otherwise, the prosecutor's intelligent agent and the defense agent conduct adversarial debate to generate a harmful debate history; and the discriminative debate history is determined based on the harmless debate history and the harmful debate history.

[0037] Among them, the clue detection result refers to the Boolean value result obtained by retrieving the multimodal data to be detected without explicit harmful clues, the harmless debate history refers to the debate history when the clue detection result is empty, and the harmful debate history refers to the debate history when the clue detection result is not empty.

[0038] Specifically, the prosecutor's intelligent agent conducts a comprehensive and in-depth semantic and content scan of the multimodal data to be detected that lacks explicit harmful clues. It focuses on uncovering various implicit harmful clues hidden in the data, such as metaphorical mappings, cultural stereotypes, and historical allusions, and marks the discovered clues as... If, after a complete deep scan, the prosecutor's AI agent does not find any hidden harmful clues that meet the judgment criteria, it will directly reject the multimodal data to be detected, resulting in a harmless debate history. If the prosecutor's AI agent successfully uncovers hidden harmful clues such as metaphors and stereotypes during the deep scan, it will immediately initiate a multi-round iterative adversarial debate process. This process consists of K rounds of debate, where K is a flexibly adjustable parameter; in this embodiment, K is set to 3. In the k-th round of the debate process, the prosecutor's AI agent will comprehensively refer to its own accusations and the defense arguments presented in the previous round, combined with the core content of the multimodal data to be detected, to generate more targeted new accusations. The defense agent, based on its previous defense arguments, will launch targeted rebuttals against the new charges brought by the prosecutor in this round, generating corresponding defense statements by combining key information such as the context and usage scenarios of the data. The state transition formula for iterative debate is: ; ; Finally, based on the history of harmless debate and the history of harmful debate, the discriminant debate history is determined. .

[0039] Furthermore, in some embodiments, obtaining the harmful information detection result of the multimodal data to be detected based on the discriminant debate history of the multimodal data to be detected includes: evaluating the multimodal data to be detected and the discriminant debate history of the multimodal data to be detected based on a preset judge agent to obtain harmful information classification labels of the multimodal data to be detected; performing source interpretation on the harmful information classification labels to generate natural language interpretation results corresponding to the harmful information classification labels; and obtaining the harmful information detection result based on the harmful information classification labels and the natural language interpretation results corresponding to the harmful information classification labels.

[0040] Among them, the harmful information classification label refers to the classification result that includes whether the information is harmful and the specific category of the harmful information, while the natural language interpretation result refers to the transformation of the source interpretation into easy-to-understand and well-organized natural language text content.

[0041] Specifically, the judge's intelligent agent Receive raw input and the complete history of the debate (include or The judge does not participate in the debate, but rather assesses the logical strength and evidentiary sufficiency of both sides' arguments, thus obtaining a classification label. This involves determining whether the content is harmful information and its specific category (such as racial discrimination, gender discrimination, etc.) and interpreting natural language. That is, to explain the reasons for the judgment (e.g., "Although the text contains offensive words, in combination with the context of the citation pointed out by the defense, the content is actually a criticism of harmful speech, and therefore is deemed harmless").

[0042] It should be noted that the prosecutor's intelligent agent, defense lawyer's intelligent agent, and judge's intelligent agent in the embodiments of the present invention all use existing multimodal models, including closed-source commercial models (such as GPT, Gemini series) and open-source models (such as Qwen series). They do not require training and can be implemented through local deployment or API calls to realize the harmful information detection method of the embodiments of the present invention.

[0043] To enable those skilled in the art to better understand the harmful information detection method of the present invention, the following explanation will be provided in conjunction with specific embodiments.

[0044] For example, Figure 2 This is a schematic flowchart of a harmful information detection method according to a specific embodiment of the present invention, as shown below. Figure 2As shown, the system first uses multimodal data combining text and images as input. The prosecutor's AI agent then performs a rapid scan for explicit harmful information. If explicit harmful clues are detected, the system enters the fast-track trial channel (Track I, targeting explicit content), where the prosecutor's AI agent initiates charges, the defense AI agent refutes them, and a debate history is generated. The data is then submitted to the judge's AI agent. If no explicit harmful clues are found, the data enters the deep debate channel (Track II, for implicit content). The prosecutor's AI agent first performs a deep scan of the implicit harmful information. If no clues are found, the data is directly rejected and marked as harmless. If clues are found, K rounds of adversarial debate are conducted, generating a debate history. The data is then submitted to the judge's AI agent; ultimately, the judge's AI agent, combining the corresponding debate history, makes the final ruling and outputs a classification label containing harmful information. and natural language interpretation The complete ruling.

[0045] further, Figure 3 A schematic diagram of a fast trial provided according to a specific embodiment of the present invention, such as Figure 3 As shown, the samples to be detected include images of people with bloodstains and text information. Their true label is non-hate speech. The zero-sample multimodal large model misjudged them as transgender hate speech. In this embodiment of the invention, the images are first detected to contain threatening language, thereby triggering a fast trial channel. The prosecution agent (i.e., the prosecutor's agent) accuses the content of equating transgender women with predatory men and creating dangerous rhetoric (i.e., the indictment). The defense agent then refutes this by arguing that the tweet actually quotes content to condemn, defining it as violence hidden in hate speech (i.e., the rebuttal). Finally, the judge agent, based on the views of both sides (i.e., the debate history), rules that the sample is non-hate speech because the tweet explicitly condemns hate and violence in the emoji.

[0046] further, Figure 4 A schematic diagram of an in-depth debate provided according to a specific embodiment of the present invention, such as Figure 4As shown, the sample to be tested includes text and image information of a specific group of people performing mathematical calculations. The true label is negative stereotype. The zero-sample multimodal large model misjudged that it emphasizes the importance of mathematical ability and did not find any harmful or negative expressions. In this embodiment of the invention, the lack of detection of explicit insults triggered in-depth debate. The prosecution agent (i.e., the prosecutor's agent) first pointed out the stereotype, that is, the text and image together point to "a specific group of people are good at math", which is reinforced by the image of a child. The defense agent (i.e., the defense attorney's agent) countered that the text only praised mathematical ability and did not mention a specific group of people. It was just a normal learning scenario (i.e., the defense statement). The prosecution agent then further pointed out that it ignored the relationship between the text and image and deliberately chose specific features to deepen the stereotype. The defense used excellent math as a positive evaluation and it was difficult to determine that there was any malicious response. Finally, the judge agent combined the debate between the two sides and ruled that the sample was a negative stereotype statement. The reason is that the text and the image of a specific group jointly reinforced the stereotype. Although there was no explicit attack, it was a subtle stereotype based on ethnic attributes and conveyed a negative stereotype.

[0047] The harmful information detection method provided by the present invention first acquires the multimodal data to be detected, performs explicit harmful information detection on it to obtain the explicit harmful information clue results and determines the gate control discrimination result, then conducts discrimination debate on the multimodal data based on the gate control discrimination result to obtain the discrimination debate history, and finally obtains the harmful information detection result of the multimodal data based on the discrimination debate history. This method solves the problems of imbalance between efficiency and accuracy of related technologies, easy misjudgment of complex semantics, and lack of reasoning process and interpretation of detection results, thereby improving the accuracy and performance of harmful information detection and enhancing the applicability of harmful information detection.

[0048] Next, the harmful information detection device according to an embodiment of the present invention will be described with reference to the accompanying drawings.

[0049] Figure 5 This is a block diagram of a harmful information detection device provided according to an embodiment of the present invention.

[0050] like Figure 5 As shown, the harmful information detection device includes: an acquisition module 100, a discrimination module 200, and a detection module 300.

[0051] The acquisition module 100 is used to acquire the multimodal data to be detected; the discrimination module 200 is used to detect explicit harmful information in the multimodal data to be detected to obtain the explicit harmful information clue results of the multimodal data to be detected, and to discriminate the explicit harmful information clue results to obtain the gating discrimination result; the detection module 300 is used to perform discrimination and debate on the multimodal data to be detected based on the gating discrimination result to obtain the discrimination and debate history of the multimodal data to be detected, and to obtain the harmful information detection result of the multimodal data to be detected based on the discrimination and debate history of the multimodal data to be detected.

[0052] Furthermore, in some embodiments, the discrimination module 200 is specifically used to: identify whether the result of the harmful information explicit clue is an explicit clue; if the result of the harmful information explicit clue is an explicit clue, the gating discrimination result is to conduct a trial based on the first channel; otherwise, the gating discrimination result is to conduct a debate based on the second channel; wherein, the processing rate of the first channel is higher than the processing rate of the second channel.

[0053] Furthermore, in some embodiments, the detection module 300 is specifically used for: detecting the multimodal data to be detected according to a preset prosecutor intelligent agent, generating an indictment corresponding to the multimodal data to be detected; refuting the indictment according to a preset defense intelligent agent, obtaining rebuttal opinions; and generating a discrimination debate history based on the indictment and the rebuttal opinions.

[0054] Furthermore, in some embodiments, the detection module 300 is also used to: perform clue detection on the multimodal data to be detected according to the preset prosecutor intelligent agent, and obtain the clue detection result; if the clue detection result is empty, generate a harmless debate history; otherwise, conduct adversarial debate through the prosecutor intelligent agent and the defense intelligent agent to generate a harmful debate history; and determine the discriminant debate history based on the harmless debate history and the harmful debate history.

[0055] Furthermore, in some embodiments, the detection module 300 is also used to: evaluate the discrimination and debate history of the multimodal data to be detected and the multimodal data to be detected based on a preset judge intelligent agent, and obtain harmful information classification labels of the multimodal data to be detected; perform source interpretation on the harmful information classification labels and generate natural language interpretation results corresponding to the harmful information classification labels; and obtain harmful information detection results based on the harmful information classification labels and the natural language interpretation results corresponding to the harmful information classification labels.

[0056] The harmful information detection device provided in this embodiment of the invention first acquires multimodal data to be detected, performs explicit harmful information detection on it to obtain explicit harmful information clue results and determines the gate control discrimination result, then conducts discrimination debate on the multimodal data based on the gate control discrimination result to obtain the discrimination debate history, and finally obtains the harmful information detection result of the multimodal data based on this discrimination debate history. This solves the problems of imbalance between efficiency and accuracy of related technologies, easy misjudgment of complex semantics, and detection results without reasoning process and uninterpretable, thereby improving the accuracy and detection performance of harmful information detection and enhancing the applicability of harmful information detection.

[0057] It should be noted that the foregoing explanation of the harmful information detection method embodiment also applies to the harmful information detection device of this embodiment, and will not be repeated here.

[0058] Figure 6This is a schematic diagram of an electronic device provided according to an embodiment of the present invention. The electronic device may include: The memory 601, the processor 602, and the computer program stored on the memory 601 and capable of running on the processor 602.

[0059] When the processor 602 executes the program, it implements the harmful information detection method provided in the above embodiments.

[0060] Furthermore, electronic devices also include: Communication interface 603 is used for communication between memory 601 and processor 602.

[0061] The memory 601 is used to store computer programs that can run on the processor 602.

[0062] The memory 601 may include high-speed RAM (Random Access Memory) memory, and may also include non-volatile memory, such as at least one disk storage.

[0063] If the memory 601, processor 602, and communication interface 603 are implemented independently, then the communication interface 603, memory 601, and processor 602 can be interconnected via a bus to complete communication between them. The bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 6 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0064] Optionally, in a specific implementation, if the memory 601, processor 602, and communication interface 603 are integrated on a single chip, then the memory 601, processor 602, and communication interface 603 can communicate with each other through an internal interface.

[0065] Processor 602 may be a CPU (Central Processing Unit), an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement embodiments of the present invention.

[0066] In addition, embodiments of the present invention also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method for detecting harmful information.

[0067] In addition, embodiments of the present invention also provide a computer program product, including a computer program, which is executed to implement the above-described method for detecting harmful information.

[0068] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0069] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0070] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A method for detecting harmful information, characterized in that, Includes the following steps: Acquire the multimodal data to be detected; The explicit harmful information detection is performed on the multimodal data to be detected to obtain the explicit harmful information clue results of the multimodal data to be detected, and the explicit harmful information clue results are judged to obtain the gating judgment result; Based on the gating discrimination result, the multimodal data to be detected is discriminated and debated to obtain the discrimination and debate history of the multimodal data to be detected, and the harmful information detection result of the multimodal data to be detected is obtained according to the discrimination and debate history of the multimodal data to be detected.

2. The method according to claim 1, characterized in that, The process of judging the explicit clues of harmful information to obtain the gating judgment result includes: Whether the result of identifying the explicit clues of harmful information is an explicit clue; If the result of the explicit clue of harmful information is the explicit clue, then the gated judgment result is to conduct a trial based on the first channel; otherwise, the gated judgment result is to conduct a debate based on the second channel. The processing rate of the first channel is higher than that of the second channel.

3. The method for detecting harmful information according to claim 2, characterized in that, The gating discrimination result is based on the first channel for adjudication. Based on the gating discrimination result, the multimodal data to be detected is subjected to discrimination and debate to obtain the discrimination and debate history of the multimodal data to be detected, including: The multimodal data to be detected is detected by a pre-set prosecutor's intelligent agent, and an indictment corresponding to the multimodal data to be detected is generated. The defense agent, based on a pre-set intelligent system, refutes the indictment and obtains rebuttal opinions. Based on the indictment and the rebuttals, the discriminative debate history is generated.

4. The method for detecting harmful information according to claim 2, characterized in that, The gated discrimination result is based on the second channel for discrimination. Based on the gated discrimination result, the multimodal data to be detected is discriminated and discriminated to obtain the discrimination and discrimination history of the multimodal data to be detected, including: Based on the preset prosecutor's intelligent agent, the multimodal data to be detected is used to detect clues, and the clue detection results are obtained; If the clue detection result is empty, a harmless debate history is generated; otherwise, a harmful debate history is generated through adversarial debate between the prosecutor's intelligent agent and the defense agent's intelligent agent. The discriminative debate history is determined based on the harmless debate history and the harmful debate history.

5. The method for detecting harmful information according to claim 1, characterized in that, The step of obtaining the harmful information detection result of the multimodal data to be detected based on the discrimination and debate history of the multimodal data to be detected includes: Based on a pre-defined judge's intelligent agent, the multimodal data to be detected and the discrimination debate history of the multimodal data to be detected are evaluated to obtain a harmful information classification label for the multimodal data to be detected. The harmful information classification tags are traced and interpreted to generate natural language interpretation results corresponding to the harmful information classification tags; The harmful information detection result is obtained based on the harmful information classification label and the natural language interpretation result corresponding to the harmful information classification label.

6. A device for detecting harmful information, characterized in that, include: The acquisition module is used to acquire the multimodal data to be detected; The discrimination module is used to detect explicit harmful information in the multimodal data to be detected to obtain explicit harmful information clues in the multimodal data to be detected, and to discriminate the explicit harmful information clues to obtain a gating discrimination result. The detection module is used to perform discrimination and debate on the multimodal data to be detected based on the gating discrimination result, obtain the discrimination and debate history of the multimodal data to be detected, and obtain the harmful information detection result of the multimodal data to be detected based on the discrimination and debate history of the multimodal data to be detected.

7. The apparatus according to claim 6, characterized in that, The discrimination module is specifically used for: Whether the result of identifying the explicit clues of harmful information is an explicit clue; If the result of the explicit clue of harmful information is the explicit clue, then the gated judgment result is to conduct a trial based on the first channel; otherwise, the gated judgment result is to conduct a debate based on the second channel. The processing rate of the first channel is higher than that of the second channel.

8. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and capable of running on the processor, the processor executing the program to implement the harmful information detection method as described in any one of claims 1-5.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the harmful information detection method as described in any one of claims 1-5.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the harmful information detection method as described in any one of claims 1-5.