An adversarial prompt copyright verification method and device based on multi-model joint gradient optimization

By generating adversarial hints through multi-model joint gradient optimization, the robustness and applicability issues of copyright verification for large language models are solved, and stable copyright verification across models is achieved, which is applicable to the copyright protection of multiple homologous models.

CN122333429APending Publication Date: 2026-07-03HANGZHOU JUNTONG FUTURE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU JUNTONG FUTURE TECHNOLOGY CO LTD
Filing Date
2025-04-16
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing copyright verification techniques for large language models lack robustness, especially single-model fingerprint watermark generation, which is weak in adversarial capabilities and has poor transferability, making it difficult to effectively cover diverse variations of the same series of models. Furthermore, traditional methods may impair model performance or be easily tampered with.

Method used

We employ a multi-model joint gradient optimization approach, which constructs a multi-model ensemble through adversarial suffix generation and optimization. This generates consistent and effective adversarial hints across multiple homologous models, avoiding reliance on model fine-tuning and enhancing the cross-model adaptability and robustness of fingerprints.

Benefits of technology

It significantly improves the robustness and applicability of copyright verification, ensures the stability and accuracy of verification results, avoids model performance degradation, and is suitable for copyright protection of multiple large-scale language model series with the same origin.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122333429A_ABST
    Figure CN122333429A_ABST
Patent Text Reader

Abstract

This invention discloses an adversarial hint copyright verification method and device based on multi-model joint gradient optimization. Addressing the weaknesses of single-model fingerprint generation in terms of weak adversarial capabilities, poor transferability, and insufficient concealment, this invention constructs multiple downstream models to simulate technical variations of the original model series. It employs multi-model joint gradient optimization to generate adversarial hints with cross-model recognition capabilities as model fingerprint features. Copyright verification is achieved by detecting the specific response patterns of the test model to the adversarial hints. This invention relies on multiple downstream models to capture the most essential features of the original model. Adversarial hints optimized using these features possess stronger specificity and effectiveness, effectively addressing intellectual property risks such as model leakage and unauthorized distribution.
Need to check novelty before this filing date? Find Prior Art