基于欧几里得流匹配的大语言模型祛毒方法及装置
By using the Euclidean flow matching feature transformation operator to perform smooth mapping in the feature space of the sparse autoencoder, the contradiction between detoxification effect and generation quality in the existing technology is resolved, achieving efficient and stable detoxification effect while maintaining generation fluency and semantic integrity.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CIVIL AVIATION FLIGHT UNIV OF CHINA
- Filing Date
- 2026-05-25
- Publication Date
- 2026-07-17
AI Technical Summary
Existing detoxification schemes for large language models struggle to effectively reduce the probability of outputting harmful content while maintaining generation quality. Furthermore, existing sparse autoencoder methods suffer from loss of generation fluency and semantic corruption due to indiscriminate intervention.
The Euclidean flow matching feature transformation operator is used to perform continuous smooth mapping in the feature space of the sparse autoencoder. By smoothly transforming toxic features into non-toxic features, and combining minimization, maximization and adaptive intervention strategies, the feature transformation process is precisely controlled to achieve a balance between detoxification effect and generation quality.
It significantly reduces the probability of toxic output while only slightly reducing the quality of generated content, improving the stability and interpretability of anti-virus, adapting to different LLM architectures and maintaining downstream task capabilities, and is suitable for a variety of security control needs and scenarios.
Smart Images

Figure CN122286767B_ABST