Model prompt word injection attack behavior detection method and device

CN122365485APending Publication Date: 2026-07-10STATE GRID BEIJING ELECTRIC POWER CO +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
STATE GRID BEIJING ELECTRIC POWER CO
Filing Date
2026-04-10
Publication Date
2026-07-10

Smart Images

  • Figure CN122365485A_ABST
    Figure CN122365485A_ABST
Patent Text Reader

Abstract

The application discloses a method and device for detecting injection attack behavior of a model prompt word, and relates to the fields of artificial intelligence and information security. The method comprises the following steps: in the case that a large language model receives a model prompt word, extracting attention behavior features of each word element of the model prompt word by a word element detector, and performing injection risk scoring on each word element in the model prompt word; performing smoothing processing on the injection risk scores of each word element in the model prompt word by using a sliding window to obtain a posterior injection probability sequence; detecting the maximum length of a high-risk word element sequence according to the posterior injection probability sequence; and determining a behavior detection label of the model prompt word according to the detected maximum length of the high-risk word element sequence. The application solves the technical problem of inaccurate detection of indirect prompt injection attacks by a large language model in the prior art.
Need to check novelty before this filing date? Find Prior Art