Model prompt word injection attack behavior detection method and device
CN122365485APending Publication Date: 2026-07-10STATE GRID BEIJING ELECTRIC POWER CO +2
View PDF 0 Cites 0 Cited by
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- STATE GRID BEIJING ELECTRIC POWER CO
- Filing Date
- 2026-04-10
- Publication Date
- 2026-07-10
Smart Images

Figure CN122365485A_ABST
Abstract
The application discloses a method and device for detecting injection attack behavior of a model prompt word, and relates to the fields of artificial intelligence and information security. The method comprises the following steps: in the case that a large language model receives a model prompt word, extracting attention behavior features of each word element of the model prompt word by a word element detector, and performing injection risk scoring on each word element in the model prompt word; performing smoothing processing on the injection risk scores of each word element in the model prompt word by using a sliding window to obtain a posterior injection probability sequence; detecting the maximum length of a high-risk word element sequence according to the posterior injection probability sequence; and determining a behavior detection label of the model prompt word according to the detected maximum length of the high-risk word element sequence. The application solves the technical problem of inaccurate detection of indirect prompt injection attacks by a large language model in the prior art.
Need to check novelty before this filing date? Find Prior Art