基于欧几里得流匹配的大语言模型祛毒方法及装置

By using the Euclidean flow matching feature transformation operator to perform smooth mapping in the feature space of the sparse autoencoder, the contradiction between detoxification effect and generation quality in the existing technology is resolved, achieving efficient and stable detoxification effect while maintaining generation fluency and semantic integrity.

CN122286767BActive Publication Date: 2026-07-17CIVIL AVIATION FLIGHT UNIV OF CHINA

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CIVIL AVIATION FLIGHT UNIV OF CHINA
Filing Date
2026-05-25
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing detoxification schemes for large language models struggle to effectively reduce the probability of outputting harmful content while maintaining generation quality. Furthermore, existing sparse autoencoder methods suffer from loss of generation fluency and semantic corruption due to indiscriminate intervention.

Method used

The Euclidean flow matching feature transformation operator is used to perform continuous smooth mapping in the feature space of the sparse autoencoder. By smoothly transforming toxic features into non-toxic features, and combining minimization, maximization and adaptive intervention strategies, the feature transformation process is precisely controlled to achieve a balance between detoxification effect and generation quality.

Benefits of technology

It significantly reduces the probability of toxic output while only slightly reducing the quality of generated content, improving the stability and interpretability of anti-virus, adapting to different LLM architectures and maintaining downstream task capabilities, and is suitable for a variety of security control needs and scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122286767B_ABST
    Figure CN122286767B_ABST
Patent Text Reader

Abstract

本发明公开了一种基于欧几里得流匹配的大语言模型祛毒方法及装置,属于人工智能领域。为更好地平衡祛毒效果‑生成质量,本发明通过稀疏自编码器提取LLM输出的特征,并基于欧几里得流匹配进行祛毒变换,最后替换原始输入激活,引导LLM生成非毒性内容。本发明应用于大语言模型领域,可对LLM高效祛毒且不会显著牺牲流畅度。
Need to check novelty before this filing date? Find Prior Art