一种基于分层序列标注的高隐蔽性音频伪造定位方法及系统

By constructing a hierarchical sequence labeling method for audio forgery localization, and utilizing EAT, Conformer, and Bi-LSTM networks combined with CRF layers, the problem of high-concealment audio forgery localization driven by a large language model is solved, achieving accurate detection and localization of local forgeries and improving the accuracy and robustness of detection.

CN122245350BActive Publication Date: 2026-07-17XIANGTAN UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
XIANGTAN UNIV
Filing Date
2026-05-25
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing technologies lack effective localization capabilities when detecting and locating local deep forgery attacks driven by large language models. In particular, they have weak perception capabilities at splicing boundaries, and the prediction results are easily fragmented, making it difficult to capture the subtle acoustic boundaries and semantic temporal logical breaks of highly covert audio forgeries.

Method used

A hierarchical sequence labeling-based approach is adopted. By constructing a locally tampered forged dataset, acoustic features are extracted using a pre-trained EAT model. Temporal modeling is performed by combining Conformer and Bi-LSTM networks, and structured prediction is performed using a boundary-aware hybrid loss function and CRF layer to accurately locate the subtle acoustic inconsistencies introduced by content tampering in audio.

Benefits of technology

It achieves precise location of highly covert audio forgeries, improves location accuracy and boundary clarity, effectively identifies extreme phrase semantic tampering segments and suppresses random noise, and has strong detection capabilities and cross-speaker generalization capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122245350B_ABST
    Figure CN122245350B_ABST
Patent Text Reader

Abstract

本发明属于人工智能与网络空间安全技术领域,具体公开一种基于分层序列标注的高隐蔽性音频伪造定位方法及系统。该方法包括:利用预训练的EAT提取对非语音伪影敏感的高维声学特征;通过堆叠的Conformer与双向LSTM网络进行多尺度时序建模,捕捉局部声学相关性与长程韵律依赖;最后利用CRF层结合发射与转移分数,利用维特比算法解码得到结构一致的伪造标签序列;且在模型训练过程中引入边界感知混合损失函数,强化对拼接边界微弱痕迹的识别。本发明能有效解决传统模型预测碎片化的问题,精准定位由高隐蔽性篡改引发的局部声学不一致性,具有高鲁棒性。
Need to check novelty before this filing date? Find Prior Art