Method and system for spoofed speech attribution based on front-end time-frequency attention

CN119785829BActive Publication Date: 2026-05-26ARMY ENG UNIV OF PLA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ARMY ENG UNIV OF PLA
Filing Date
2024-12-31
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing methods for attributing spoofed speech perform poorly on encoded and decoded speech signals and are unable to effectively identify spoofed speech.

Method used

A forged speech attribution method based on front-end time-frequency attention is adopted. By constructing a front-end time-frequency attention module, a feature extraction module, and a feature classification module, the attribution performance of forged speech methods is improved by using time-frequency features for weighted and classified calculations. This improves the attribution performance of forged speech that has been converted by encoding and decoding as well as speech that has not been converted by encoding and decoding.

Benefits of technology

It significantly improves the performance of forgery recognition methods for both encoded and unencoded speech, and enhances the robustness and accuracy of attributing forged speech.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119785829B_ABST
    Figure CN119785829B_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for attributing forged speech based on front-end time-frequency attention. The method includes: inputting the time-frequency features of a speech sample into a trained forged speech attribution model, and outputting the probability that the time-frequency features of the speech sample belong to various speech forgery methods, as the recognition result of the forged speech attribution. The forged speech attribution model includes: a front-end time-frequency attention module, used to weight the time-frequency features of the speech sample using front-end time-frequency attention; a feature extraction module, used to extract the deep features of the speech sample from the time-frequency features weighted by the time-frequency attention; and a feature classification module, used to classify the deep features of the speech sample. This invention, when performing forged speech attribution, can simultaneously attribute forged methods to both encoded and unencoded speech, and is beneficial to improving the performance of forged method attribution for both encoded and unencoded speech.
Need to check novelty before this filing date? Find Prior Art