一种基于分层序列标注的高隐蔽性音频伪造定位方法及系统
By constructing a hierarchical sequence labeling method for audio forgery localization, and utilizing EAT, Conformer, and Bi-LSTM networks combined with CRF layers, the problem of high-concealment audio forgery localization driven by a large language model is solved, achieving accurate detection and localization of local forgeries and improving the accuracy and robustness of detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- XIANGTAN UNIV
- Filing Date
- 2026-05-25
- Publication Date
- 2026-07-17
AI Technical Summary
Existing technologies lack effective localization capabilities when detecting and locating local deep forgery attacks driven by large language models. In particular, they have weak perception capabilities at splicing boundaries, and the prediction results are easily fragmented, making it difficult to capture the subtle acoustic boundaries and semantic temporal logical breaks of highly covert audio forgeries.
A hierarchical sequence labeling-based approach is adopted. By constructing a locally tampered forged dataset, acoustic features are extracted using a pre-trained EAT model. Temporal modeling is performed by combining Conformer and Bi-LSTM networks, and structured prediction is performed using a boundary-aware hybrid loss function and CRF layer to accurately locate the subtle acoustic inconsistencies introduced by content tampering in audio.
It achieves precise location of highly covert audio forgeries, improves location accuracy and boundary clarity, effectively identifies extreme phrase semantic tampering segments and suppresses random noise, and has strong detection capabilities and cross-speaker generalization capabilities.
Smart Images

Figure CN122245350B_ABST