基于多轮交互意图累积图谱的内容安全评估方法

By constructing an intent accumulation graph and using dynamic threshold adjustment, the problem of identifying progressive inducement attacks in multi-turn dialogues is solved, enabling accurate detection and defense of dialogue risks and significantly improving the accuracy and adaptability of content security assessment.

CN122087832BActive Publication Date: 2026-07-17ZHEJIANG CHUANGLIN TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHEJIANG CHUANGLIN TECH CO LTD
Filing Date
2026-04-21
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing multi-turn dialogue risk detection technologies suffer from problems such as long-distance forgetting, lack of quantitative accumulation, and fixed thresholds when facing progressive inducement attacks. They are difficult to effectively identify cross-turn intent associations and lack interpretability due to their reliance on large-scale labeled data and black-box characteristics.

Method used

The method adopts a multi-turn interaction intent accumulation graph approach. By constructing an intent accumulation graph, it calculates a comprehensive risk value using a predefined logical dependency matrix and time decay factor, and dynamically adjusts the interception threshold. By combining intent path accumulation and time decay, it achieves accurate detection and defense of dialogue risks.

Benefits of technology

It achieves accurate detection of progressively induced attacks, provides interpretability through explicit logical dependency matrix injection, reduces dependence on large-scale labeled data, and improves the system's sensitivity and adaptability through dynamic threshold adjustment, making it suitable for high-frequency interaction scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122087832B_ABST
    Figure CN122087832B_ABST
Patent Text Reader

Abstract

本发明公开了一种基于多轮交互意图累积图谱的内容安全评估方法,包含:接收多轮对话数据,对每轮对话进行意图识别得到意图类型和基础风险分值;构建意图累积图谱,将每轮对话作为节点,基于预定义的逻辑依赖矩阵确定节点间边的权重;基于所述意图累积图谱,采用融合时间衰减与意图路径累积的非线性累积公式计算当前轮的综合风险值;将所述综合风险值与拦截阈值比较,根据比较结果触发相应防御动作。本申请的基于多轮交互意图累积图谱的内容安全评估方法,通过意图图谱化建模、非线性风险累积的协同机制,实现了多轮对话中渐进式诱导攻击的精准检测,显著提升了内容安全评估的准确性、实时性和适应性。
Need to check novelty before this filing date? Find Prior Art