一种基于后门水印与感知加密的问答数据集版权保护方法及装置

By introducing natural language adverb watermarking and semantic vector encryption mechanisms into the question-and-answer dataset, and combining watermark-triggered detection, the usability and traceability problems of copyright protection in existing technologies are solved, and efficient data copyright protection and secure distribution are achieved.

CN121765697BActive Publication Date: 2026-07-17INST OF GEOGRAPHICAL SCI & NATURAL RESOURCE RES CAS

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INST OF GEOGRAPHICAL SCI & NATURAL RESOURCE RES CAS
Filing Date
2025-12-23
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing copyright protection schemes for question-and-answer datasets struggle to strike a balance between usability, imperceptibility, and traceability. Explicit watermarks are easily identified and removed, and traditional encryption methods cannot effectively track and assign responsibility after data decryption.

Method used

A backdoor watermarking and perceptual encryption method is adopted. Natural language adverb watermarks are generated through a large language model and high-dimensional vector encryption is performed by combining a semantic vectorization model. Copyright verification is performed by combining watermark triggering and discrimination mechanisms.

Benefits of technology

It achieves hidden copyright identification of question-and-answer datasets without changing the original semantics, and has machine-detectable and black-box forensic capabilities, improving the security and controllability of data use. It is suitable for closed-source APIs and managed inference environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121765697B_ABST
    Figure CN121765697B_ABST
Patent Text Reader

Abstract

本发明为一种基于后门水印与感知加密的问答数据集版权保护方法及装置,包括:划分水印嵌入子集与加密保护子集;构建水印副词库及自然语言改写指令,利用大语言模型生成水印响应样本,并在问题侧嵌入触发词,形成水印问答子集。采用文本语义向量化模型将问答样本映射为高维向量,并以用户特定的正交矩阵密钥实施向量域加密,生成加密问答子集;授权场景下借助密钥管理执行可逆解密恢复可用样本;检测阶段,输入预设触发问题,利用水印判别机制识别回答中的隐式水印响应,实现对涉嫌侵权模型的版权核验。本方法不仅保证了语义质量,而且结合感知加密机制确保了数据的可控使用与防泄漏安全,对于AI问答大模型训练数据的知识产权保护具有重要意义。
Need to check novelty before this filing date? Find Prior Art