A Reliable Method and System for Determining Multimodal Endogenous Harmful Knowledge Based on Trigger Link Resolution and Consistency Verification
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-25
- Publication Date
- 2026-08-14
AI Technical Summary
[0003]由于现有技术主要侧重最终输出分类、拒答过滤或中间状态风险检测,尚未将触发条件、跨模态交互、一致性冲突与推理过程校验整合为统一可信测定框架
(1)本发明将可信测定对象由最终输出扩展至触发条件、跨模态交互及推理过程,能够提高测定结果的可解释性与可追溯性。
Smart Images

Figure CN122571324A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence security governance and multimodal content understanding technology, specifically involving a method and system for reliable determination of multimodal endogenous harmful knowledge based on trigger link parsing and consistency verification. Background Technology
[0002] With the widespread application of multimodal large models in scenarios such as image understanding, visual question answering, image-text reasoning, and human-computer interaction, the problem of harmful knowledge implicit in the model being activated and generating non-compliant output under the combined effect of images and text is becoming increasingly prominent. Compared with traditional text security issues, multimodal endogenous harmful knowledge has stronger concealment and combinability. The same text may be suppressed under plain text conditions, but after introducing specific image cues, semantic deviation, dangerous detail completion, or role shift may occur, thereby generating high-risk content.
[0003] Since existing technologies mainly focus on final output classification, rejection filtering, or intermediate state risk detection, they have not yet integrated triggering conditions, cross-modal interactions, consistency conflicts, and inference process verification into a unified credibility measurement framework. Therefore, this paper proposes a multimodal endogenous harmful knowledge credibility measurement method that can simultaneously conduct systematic analysis of triggering conditions, cross-modal interactions, and inference processes, thereby providing more robust technical support for risk suppression, decision control, and model correction. Summary of the Invention
[0004] This invention aims to address the shortcomings of existing technologies and provides the following solutions: A reliable method for determining multimodal endogenous harmful knowledge based on trigger link resolution and consistency verification includes the following steps: Acquire multimodal input samples to be tested, including image input and text input, and generate candidate responses using a large multimodal model; A harmful trigger criterion is constructed, and a joint trigger search is performed on the multimodal input samples. Under the condition of maintaining semantic readability and controllable perturbation, the minimum trigger pair that can induce a risk response is obtained, and the triggering factors corresponding to the minimum trigger pair are decomposed into visual factors, linguistic factors and cross-modal link factors. Perform cross-modal consistency verification on the multimodal input samples, construct a statement subgraph, an evidence subgraph, and semantic paths between entities, and calculate the support, conflict, and background credibility between image evidence and text statements; The response process of the multimodal large model is divided into a candidate response development stage and a security review stage. The security review stage reviews and judges the candidate responses based on evidence fit, link coherence, and risk spillover. The obtained multi-source evidence is aggregated and combined with the priority constraint of counter-evidence and confidence calibration to output credible measurement results.
[0005] Preferably, the harmful triggering criterion is constructed using an external security evaluator; The external security evaluator outputs a risk assessment value for the candidate response and forms a trigger status flag accordingly. The risk assessment value is estimated using expected risk; The expected risk is obtained by averaging the evaluation results of the same input under different sampling temperatures, different decoding strategies, or multiple sampling conditions.
[0006] Preferably, the joint trigger search includes: at least one image perturbation operation, at least one text perturbation operation, and attribution explanation of the execution link of the minimum trigger pair, and the triggering factors are structured and stored as a triggering factor library; The image perturbation operations include region occlusion, region replacement, region rearrangement, and significant region enhancement; The text perturbation operations include synonym rewriting, insertion, deletion, word order transformation, and role cue rewriting, and utilize semantic similarity constraints and task preservation constraints to filter out non-minimum trigger pairs; The trigger factor library is used to record factor type, location index, semantic label, risk contribution, recurrence frequency, and corresponding risk category.
[0007] Preferably, the cross-modal consistency verification is achieved by uniformly constructing maps of image entities, text entities, and external retrieval evidence; The constructed cross-modal graph includes the statement subgraph and the evidence subgraph, and retrieves the external related evidence chains around the risk statement object in the text input, and forms a statement support representation accordingly; Supporting evidence representations and conflicting evidence representations are formed by positive attention aggregation and negative attention aggregation, respectively.
[0008] Preferably, the security review stage includes: The key judgment segments in the candidate response chain are aligned and verified with image evidence, text evidence, and external evidence to obtain the evidence fit. The link coherence is obtained by reviewing the preceding and following judgments and the overall link closure of the candidate response links. The risk spillover is obtained by detecting sensitive intent completion, dangerous detail expansion, and violation inference tendency in candidate response links. Based on the evidence fit, the link coherence, and the risk spillover, a decision action is output from four actions: pass, clarify, reject, and re-reasoning, to complete the process verification.
[0009] Preferably, the multi-source evidence is integrated into a comprehensive adjudication result after unified risk assessment, and the credibility of the comprehensive adjudication result is adjusted.
[0010] This invention also provides a multimodal endogenous harmful knowledge credibility determination system based on trigger link parsing and consistency verification. The system applies the above-mentioned method and includes: a multimodal sample access module, a trigger link extraction module, a graphic evidence verification module, a link review module, and a credibility adjudication module. The multimodal sample access module is used to acquire the multimodal input samples to be tested, including image input and text input, and to generate candidate responses by a multimodal large model; The trigger link extraction module is used to construct harmful trigger criteria, perform joint trigger search on the multimodal input samples, and obtain the minimum trigger pair that can induce risk response while maintaining semantic readability and controllable perturbation. The triggering factors corresponding to the minimum trigger pair are decomposed into visual factors, linguistic factors and cross-modal link factors. The image and text evidence verification module is used to perform cross-modal consistency verification on the multimodal input samples, construct a statement subgraph, an evidence subgraph, and semantic paths between entities, and calculate the support, conflict, and background credibility between image evidence and text statements. The link verification module is used to divide the response process of the multimodal large model into a candidate response unfolding stage and a security verification stage. The security verification stage performs verification and judgment based on the evidence fit, link coherence and risk spillover of the candidate response. The credible determination module is used to aggregate the obtained multi-source evidence and, in conjunction with the priority constraint of counter-evidence and confidence calibration, output credible determination results.
[0011] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) The present invention extends the credible measurement object from the final output to the triggering conditions, cross-modal interactions and reasoning processes, which can improve the interpretability and traceability of the measurement results.
[0012] (2) This invention constructs a search and trigger factor library through minimal triggering, which can provide a structured representation of visual factors, language factors and cross-modal link factors, making it convenient for risk reproduction, auditing and strategy updating.
[0013] (3) This invention improves the robustness of judgment of hidden scenarios, mismatch scenarios and background conflict scenarios through cross-modal consistency verification, process verification, priority aggregation of counter-evidence and confidence calibration, while taking into account the protection capability of normal samples.
[0014] (4) In a set of prototype system embodiments, the number of high-risk sample related counts decreased from 89 to 5, and the number of normal sample misjudgments decreased from 32 to 2, indicating that the present invention has good risk identification support capability and engineering application value. Attached Figure Description
[0015] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a schematic diagram of the method flow according to an embodiment of the present invention; Figure 2 This is a flowchart of a method according to an embodiment of the present invention. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0019] Example 1 In this embodiment, as Figure 1 , Figure 2 As shown, the multimodal endogenous harmful knowledge credibility determination method based on trigger link resolution and consistency verification includes the following steps: S1. Obtain the multimodal input samples to be tested, including image input and text input, and generate candidate responses by the multimodal large model.
[0020] S2. Construct harmful trigger criteria, perform joint trigger search on multimodal input samples, and obtain the minimum trigger pair that can induce risk response while maintaining semantic readability and controllable perturbation. Decompose the triggering factors corresponding to the minimum trigger pair into visual factors, linguistic factors and cross-modal link factors.
[0021] The harmful trigger criterion is constructed through an external security evaluator; the external security evaluator outputs a risk assessment value for candidate responses and forms a trigger status mark accordingly; the risk assessment value is estimated using the expected risk; the expected risk is obtained by averaging the evaluation results of the same input under different sampling temperatures, different decoding strategies, or multiple sampling conditions.
[0022] The joint trigger search includes: at least one image perturbation operation, at least one text perturbation operation, and a link attribution interpretation for the minimum trigger pair, and the triggering factors are structured and stored as a trigger factor library; the image perturbation operations include region occlusion, region replacement, region rearrangement, and salient region enhancement; the text perturbation operations include synonym rewriting, insertion, deletion, word order transformation, and role cue rewriting, and non-minimum trigger pairs are screened using semantic similarity constraints and task preservation constraints; the trigger factor library is used to record factor type, location index, semantic label, risk contribution, recurrence frequency, and corresponding risk category.
[0023] In this embodiment, given a test sample containing image input and text instructions, the target multimodal large model first generates one or more candidate responses, and then an external security evaluator outputs a risk assessment value for the candidate responses. An external security evaluator is used instead of relying solely on the model's self-scoring because the external evaluation result is more suitable as a unified triggering criterion. Furthermore, to reduce fluctuations caused by the randomness of a single decoding, this invention uses a multiple sampling averaging method to calculate the expected harmful risk, and uses this expected risk as the basic input for subsequent trigger search and consistency verification.
[0024] The expected harmful risk can be expressed as: in, This represents the expected harmful risk under the current input conditions. Indicates the number of samples. Indicates the first Index of the next sample Indicates the first Candidate responses obtained from the second sampling. This represents the risk assessment function of an external security assessor. This indicates that the external security evaluator is for the first Subsampling response The given risk assessment value.
[0025] During the trigger chain verification phase, this invention performs a minimum trigger pair search for samples with a trigger risk tendency. Specifically, for image input, region occlusion, region replacement, region rearrangement, or salient region enhancement are performed; for text input, synonym rewriting, insertion, deletion, word order transformation, or role cue transformation are performed. For each combination of image and text perturbations, the expected harmful risk is recalculated. Under the conditions of satisfying semantic similarity and task preservation thresholds, the combination with the smallest perturbation amplitude and the risk change reaching a preset condition is selected as the minimum trigger pair. If image perturbation alone can cause a significant change in risk, it is denoted as a visual factor; if text perturbation alone can cause a significant change in risk, it is denoted as a linguistic factor; if both image and text perturbations are required to cause a significant change in risk, it is denoted as a cross-modal link factor.
[0026] Furthermore, various triggering factors are written into a trigger factor library. Each entry in the library includes at least a factor number, factor type, location index, risk category, recurrence frequency, risk contribution value, and a sample input fragment. If a new sample appears that is highly similar to an existing entry, the corresponding entry can be directly used as supplementary verification. The trigger factor library can also be used for subsequent boundary sample construction, risk classification assessment, and strategy updates.
[0027] S3. Perform cross-modal consistency verification on multimodal input samples, construct a claim subgraph, an evidence subgraph, and semantic paths between entities, and calculate the support, conflict, and background credibility between image evidence and text claims.
[0028] Cross-modal consistency verification is achieved by uniformly constructing maps of image entities, text entities, and externally retrieved evidence. The constructed cross-modal map includes a claim subgraph and an evidence subgraph, and retrieves external related evidence chains around the risk claim objects in the image and text input, thereby forming a claim support representation. Supporting evidence representation and conflicting evidence representation are formed through positive attention aggregation and negative attention aggregation, respectively.
[0029] In this embodiment, during the cross-modal consistency constraint stage, image entities are first extracted from the image input, and text entities are extracted from the text input. A declaration subgraph is then constructed based on these two. Simultaneously, supplementary evidence is obtained through external retrieval, and an evidence subgraph is constructed. For entity pairs in the declaration subgraph, their semantically relevant paths in an external knowledge base are retrieved. When the path is short and the relationship is consistent, the corresponding information is aggregated into a consistent representation; when the path is long, the relationship is contradictory, or there is subject substitution, the corresponding information is aggregated into a conflict representation. Subsequently, combined with the query vector of the current image and text sample, support, conflict degree, and background credibility are calculated respectively as inputs for subsequent credibility determination.
[0030] Cross-modal consistency score can be expressed as: in, Indicates the cross-modal consistency score. This represents a consistent representation. Indicates conflict. This represents the query vector for the current image and text sample. This represents the consistency scoring function.
[0031] A high support level indicates that the image, text, and background knowledge are generally consistent. A high conflict level indicates that the sample contains obvious image-text mismatch, subject substitution, contextual misleading, or knowledge conflict.
[0032] S4. The response process of the multimodal large model is divided into a candidate response development stage and a security review stage. The security review stage reviews and judges the candidate response based on the evidence fit, link coherence and risk spillover.
[0033] The security review phase includes: aligning and verifying key judgment segments in the candidate response chain with image evidence, text evidence, and external evidence to obtain evidence fit; reviewing the sequential judgment relationships and overall chain closure of the candidate response chain to obtain chain coherence; detecting sensitive intent completion, dangerous detail expansion, and tendency to infer violations in the candidate response chain to obtain risk spillover; and outputting decision actions from four options—pass, clarify, reject, and re-reason—based on evidence fit, chain coherence, and risk spillover, thus completing process verification.
[0034] In this embodiment, during the credibility determination stage of the reasoning process, candidate answers and candidate response chains are first generated by reasoning, and then the evidence fit, chain coherence, and risk spillover are calculated in the process verification stage. Evidence fit measures the correspondence between candidate reasoning steps and image evidence, text evidence, and external evidence. Chain coherence measures the coherence of the preceding and following judgments and the overall chain closure. Risk spillover identifies sensitive intent completion, dangerous detail expansion, and tendency for illegal inference. Based on the combination of the three risk assessment values, the following actions are output: pass, clarify, reject, or re-reasoning: pass when evidence fit is high and process risk is low; clarify when evidence is insufficient but risk is not high; reject when process risk is significant; and re-reasoning when the logical chain is broken or self-consistency is insufficient.
[0035] The process verification module derives a credibility score and a risk score based on the three scores mentioned above, which can be expressed as follows: in, Indicates the index of the candidate response link. Indicates the first The trust score corresponding to each candidate response link Indicates the first Risk scores corresponding to each candidate response link Indicates the degree of relevance of the evidence. Indicates the link continuity. Indicates the degree of risk spillover. , , and All are non-negative weight parameters.
[0036] For the candidate set, candidates with high credibility scores and low risk scores are prioritized for approval. If none of the candidates meet the requirements, clarification, rejection, or re-inference actions are triggered.
[0037] S5. Aggregate the obtained multi-source evidence, and combine it with the priority constraint of counter-evidence and confidence calibration to output the credible measurement results.
[0038] After unified risk integration, multi-source evidence forms a comprehensive adjudication result, and the credibility of the comprehensive adjudication result is adjusted.
[0039] In this embodiment, during the evidence aggregation and confidence calibration stage, the present invention uniformly converts the trigger chain verification results, cross-modal consistency constraint results, and process verification results into risk probabilities, and further maps them to the logarithmic probability space for linear fusion. This approach is adopted because the original score scales output by different verification modules are not consistent; direct addition or averaging can easily lead to one module abnormally dominating the overall result. Considering that conflicting evidence has a higher decision priority in security scenarios, when any module provides high-confidence counter-evidence, the overall risk probability is increased to form an aggregation result prioritizing counter-evidence.
[0040] The overall prediction probability can be expressed as: in, Indicates the overall prediction probability. This represents the index of the verification module. Indicates the first The risk probability output by each verification module Indicates the first The weight parameters corresponding to each verification module Represents the logarithmic function. Let represent the Sigmoid function. Considering that proof by contradiction usually has stronger decision-making constraints in multimodal security scenarios, this embodiment further introduces a proof by contradiction priority term. The overall risk probability after introducing the proof by contradiction priority term can be expressed as: in, This represents the overall risk probability after introducing the priority term for proof by contradiction. This represents the strength coefficient of the priority constraint for proof by contradiction. An indicator function that indicates the existence of high-confidence conflicts or rebuttal evidence. Indicates a conflict or rebuttal of evidence. Represents a mapping function. Indicates the confidence level of disproving evidence. This represents a mapping function that monotonically increases with the confidence level of proof by contradiction.
[0041] Subsequently, the aggregated risk probabilities are calibrated using temperature scaling or monotonic mapping to obtain the final confidence level. The final confidence level can be expressed as: in, Indicates the final credibility. Represents the calibration function. This indicates the calibration temperature parameter.
[0042] The final output includes risk labels, calibration confidence levels, key triggering factors, consistency or conflict summaries, reasoning process verification conclusions, and audit record numbers. These can serve as preliminary risk screening results, as well as input for subsequent interception, clarification, and manual review.
[0043] Example 2 In this embodiment, to verify the effectiveness of the method of the present invention, the publicly available results of methods in the same field on the SafeEraser dataset are selected for comparison. Considering that multimodal harmful knowledge removal not only requires high removal capabilities but also requires minimizing false positives on normal samples, Table 1 uses the effective removal rate as the main comparison indicator. For the publicly available methods, the removal rate is first calculated based on the decrease in Generality ASR relative to the Vanilla model, and its expression can be expressed as: in, Indicates the original clearance rate. express The proportion of harmful responses to the model on the corresponding dataset. This represents the proportion of harmful responses to the corresponding method on the corresponding dataset. Further, combining the original clearance rate with the normal sample protection rate yields the effective clearance rate, which can be expressed as: in, Indicates the effective clearance rate. The percentage of normal sample protection is represented by the percentage decrease in the number of high-risk samples from 89 to 5. The percentage of normal samples retained after the false positive count decreased from 32 to 2 is used as the normal sample protection rate. The effective clearance rate is then calculated. The comparison results are shown in Table 1.
[0044] Table 1 In Table 1, the Generality ASR, Specificity, and SARR data for the disclosed methods are taken from the results published in the SafeEraser paper under LLaVA-v1.5-13B settings. GA+PD, GD+PD, KL+PD, and PO+PD represent the results of introducing Prompt Decouple Loss on the basis of the corresponding forgetting methods. The "Normal Sample Protection Related Indicators" in the table use Specificity for the disclosed methods and normal sample protection rate for the method of this invention. The "Over-forgetting or False Alarm Related Indicators" in the table use SARR for the disclosed methods and residual false alarm rate for normal samples for the method of this invention.
[0045] As shown in Table 1, if only the original removal rate is considered, the publicly available methods generally achieve high values, but their ability to protect normal samples is relatively limited, and the indicators related to excessive forgetting or false positives are relatively high. In contrast, the method of this invention achieves an original removal rate of 94.38%, a normal sample protection indicator of 93.75%, and reduces the excessive forgetting or false positives indicator to 6.25%, thus achieving an effective removal rate of 88.48%, which is the best among the methods listed in Table 1. This indicates that the method of this invention achieves a better overall balance between harmful knowledge removal and normal sample protection, and is more suitable as a technical support for the reliable determination of multimodal large models and subsequent risk control.
[0046] Example 3 In this embodiment, the multimodal endogenous harmful knowledge credibility determination system based on trigger link parsing and consistency verification includes: a multimodal sample access module, a trigger link extraction module, a graphic evidence verification module, a link review module, and a credibility adjudication module; The multimodal sample access module is used to acquire the multimodal input samples to be tested, including image input and text input, and to generate candidate responses by the multimodal large model; The trigger link extraction module is used to construct harmful trigger criteria, perform joint trigger search on multimodal input samples, and obtain the minimum trigger pair that can induce risk response while maintaining semantic readability and controllable perturbation. The triggering factors corresponding to the minimum trigger pair are decomposed into visual factors, linguistic factors and cross-modal link factors. The image and text evidence verification module is used to perform cross-modal consistency verification on multimodal input samples, construct a statement subgraph, an evidence subgraph, and semantic paths between entities, and calculate the support, conflict, and background credibility between image evidence and text statements. The link verification module is used to divide the response process of the multimodal large model into a candidate response unfolding stage and a security verification stage. The security verification stage reviews and judges the candidate response based on the evidence fit, link coherence and risk spillover. The credible determination module is used to aggregate the obtained multi-source evidence and combine it with the priority constraint of counter-evidence and confidence calibration to output credible determination results.
[0047] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.
Claims
1. A method for determining the credibility of multimodal endogenous harmful knowledge based on trigger link resolution and consistency verification, characterized in that, Includes the following steps: Acquire multimodal input samples to be tested, including image input and text input, and generate candidate responses using a large multimodal model; A harmful trigger criterion is constructed, and a joint trigger search is performed on the multimodal input samples. Under the condition of maintaining semantic readability and controllable perturbation, the minimum trigger pair that can induce a risk response is obtained, and the triggering factors corresponding to the minimum trigger pair are decomposed into visual factors, linguistic factors and cross-modal link factors. Perform cross-modal consistency verification on the multimodal input samples, construct a statement subgraph, an evidence subgraph, and semantic paths between entities, and calculate the support, conflict, and background credibility between image evidence and text statements; The response process of the multimodal large model is divided into a candidate response development stage and a security review stage. The security review stage reviews and judges the candidate responses based on evidence fit, link coherence, and risk spillover. The obtained multi-source evidence is aggregated and combined with the priority constraint of counter-evidence and confidence calibration to output credible measurement results.
2. The method for determining the credibility of multimodal endogenous harmful knowledge based on trigger link resolution and consistency verification according to claim 1, characterized in that, The harmful triggering criteria are constructed using an external security evaluator; The external security evaluator outputs a risk assessment value for the candidate response and forms a trigger status flag accordingly. The risk assessment value is estimated using expected risk; The expected risk is obtained by averaging the evaluation results of the same input under different sampling temperatures, different decoding strategies, or multiple sampling conditions.
3. The method for determining the credibility of multimodal endogenous harmful knowledge based on trigger link resolution and consistency verification according to claim 1, characterized in that, The joint trigger search includes: at least one image perturbation operation, at least one text perturbation operation, and attribution explanation of the execution link of the minimum trigger pair, and the triggering factors are structured and stored as a triggering factor library; The image perturbation operations include region occlusion, region replacement, region rearrangement, and significant region enhancement; The text perturbation operations include synonym rewriting, insertion, deletion, word order transformation, and role cue rewriting, and utilize semantic similarity constraints and task preservation constraints to filter out non-minimum trigger pairs; The trigger factor library is used to record factor type, location index, semantic label, risk contribution, recurrence frequency, and corresponding risk category.
4. The method for determining the credibility of multimodal endogenous harmful knowledge based on trigger link resolution and consistency verification according to claim 1, characterized in that, The cross-modal consistency verification is achieved by uniformly constructing a map of image entities, text entities, and external retrieval evidence. The constructed cross-modal graph includes the statement subgraph and the evidence subgraph, and retrieves the external related evidence chains around the risk statement object in the text input, and forms a statement support representation accordingly; Supporting evidence representations and conflicting evidence representations are formed by positive attention aggregation and negative attention aggregation, respectively.
5. The method for determining the credibility of multimodal endogenous harmful knowledge based on trigger link resolution and consistency verification according to claim 1, characterized in that, The security review phase includes: The key judgment segments in the candidate response chain are aligned and verified with image evidence, text evidence, and external evidence to obtain the evidence fit. The link coherence is obtained by reviewing the preceding and following judgments and the overall link closure of the candidate response links. The risk spillover is obtained by detecting sensitive intent completion, dangerous detail expansion, and violation inference tendency in candidate response links. Based on the evidence fit, the link coherence, and the risk spillover, a decision action is output from four actions: pass, clarify, reject, and re-reasoning, to complete the process verification.
6. The method for determining the credibility of multimodal endogenous harmful knowledge based on trigger link resolution and consistency verification according to claim 1, characterized in that, The multi-source evidence is integrated into a unified risk assessment to form a comprehensive adjudication result, and the credibility of the comprehensive adjudication result is adjusted.
7. A multimodal endogenous harmful knowledge credibility determination system based on trigger link resolution and consistency verification, wherein the system applies the method described in any one of claims 1-6, characterized in that, include: The module includes a multimodal sample access module, a trigger link extraction module, a graphic evidence verification module, a link review module, and a trustworthy adjudication module. The multimodal sample access module is used to acquire the multimodal input samples to be tested, including image input and text input, and to generate candidate responses by a multimodal large model; The trigger link extraction module is used to construct harmful trigger criteria, perform joint trigger search on the multimodal input samples, and obtain the minimum trigger pair that can induce risk response while maintaining semantic readability and controllable perturbation. The triggering factors corresponding to the minimum trigger pair are decomposed into visual factors, linguistic factors and cross-modal link factors. The image and text evidence verification module is used to perform cross-modal consistency verification on the multimodal input samples, construct a statement subgraph, an evidence subgraph, and semantic paths between entities, and calculate the support, conflict, and background credibility between image evidence and text statements. The link verification module is used to divide the response process of the multimodal large model into a candidate response unfolding stage and a security verification stage. The security verification stage performs verification and judgment based on the evidence fit, link coherence and risk spillover of the candidate response. The credible determination module is used to aggregate the obtained multi-source evidence and, in conjunction with the priority constraint of counter-evidence and confidence calibration, output credible determination results.