Large model security protection method and device, equipment, storage medium and computer program product

By splitting attack requests into text fragments and generating distributed instruction image sequences, attacks are launched against multimodal large models. This solves the problem of security defense failure in multimodal large language models under multi-image sequence joint reasoning scenarios and improves security protection capabilities.

CN122413441APending Publication Date: 2026-07-17BEIJING QIHOOD TECHNOLOGY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING QIHOOD TECHNOLOGY CO LTD
Filing Date
2026-04-22
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing multimodal large language models lack effective security mechanisms in multi-image sequence joint reasoning scenarios, resulting in defense failure and an inability to effectively detect and defend against attacks in which malicious intent is dispersed across multiple images.

Method used

The key text instructions in the attack request are split into multiple text fragments, and a background image corresponding to each text fragment is generated. A distributed instruction image sequence is generated based on each text fragment and the background image. The multimodal large model is attacked through the distributed instruction image sequence, and security protection is implemented based on the attack results.

Benefits of technology

By disassembling harmful instructions into text fragments without independent malicious semantics and distributing them into multiple independent images, security vulnerabilities of multimodal large models in multi-image sequence joint reasoning scenarios can be detected, thereby improving the security protection capabilities of multimodal large models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122413441A_ABST
    Figure CN122413441A_ABST
Patent Text Reader

Abstract

本申请涉及人工智能安全技术领域,公开了一种大模型安全防护方法、装置、设备、存储介质及计算机程序产品,包括:响应于攻击请求,将攻击请求中的关键文本指令拆分为多个文本片段,生成各文本片段对应的背景图像,并根据各文本片段和各文本片段对应的背景图像生成分布式指令图像序列,基于分布式指令图像序列对多模态大模型进行攻击,并根据攻击结果对多模态大模型进行安全防护;本申请通过将有害指令拆解为无独立恶意语义的文本片段并编码至多张独立图像中对多模态大模型进行攻击,从而能够检测多模态大模型在多图像序列联合推理场景下的安全漏洞,并且根据攻击结果优化多模态大模型的安全防护机制,进而能够提高多模态大模型的安全防护能力。
Need to check novelty before this filing date? Find Prior Art