Countercheck method for automatically identifying speaker aiming to voice deception

A speaker recognition and speaker technology, applied in speech analysis, instruments, etc., can solve problems such as fragile confrontation capabilities

CN105139857AActive Publication Date: 2015-12-09SUN YAT SEN UNIV +1
7 Cites 39 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Publication Date
2015-12-09

Smart Images

  • Figure 1
    Figure 1
  • Figure 2
    Figure 2
  • Figure 3
    Figure 3
Patent Text Reader

Abstract

The invention provides a countercheck method for automatically identifying a speaker aiming to voice deception, which is a voice anti-spoofing technology based on a method combining various features and a plurality of sub-systems. According to the invention, the serial features of the posterior probability of a phoneme in the phonological level and the MFCC features of voice level or MFDCC features of phase level are combined, thus the performance of the system is significantly enhanced. By combining the provided i-vector sub-system and OpenSMILE (open Speech and Music Interpretation by Large Space Extraction criterion containing voice and rhythmic information, the final presentation of the system is further enhanced. To a back-end model, the development datum are used; and under the situation of knowing deceptive attacks, a two-level support vector machine has better performance compared with one-level cosine similarity or PLDA evaluations, while the one-level evaluation approach has better robustness under the situation without seeing the test datum and knowing the deceptive conditions.
Need to check novelty before this filing date? Find Prior Art

Description

technical field

[0001] The invention relates to the field of automatic speaker recognition, and more particularly, relates to a countermeasure against speech deception in automatic speaker recognition. Background technique

[0002] The purpose of speaker recognition is to automatically confirm the identity of a known speaker through a piece of speech. In the past decade, speaker recognition has attracted the attention of many researchers, and also achieved very remarkable results. However, it has been recently reported that many existing speaker recognition systems are vulnerable to different spoofing attacks, such as speaker-adaptive speech synthesis, voice conversion, and voice playback.

[0003] Since the spoken content is restricted or pre-defined, text-based speaker recognition is more robust to voice playback spoofing attacks than text-independent speaker recognition. Speaker-adaptive voice synthesis and voice transformation are the most commonly used methods of dece...

Examples

Embodiment Construction

[0057] The drawings are for illustrative purposes only, and should not be construed as limitations on this patent; in order to better illustrate this embodiment, some parts in the drawings will be omitted, enlarged or reduced, and do not represent the size of the actual product;

[0058] For those skilled in the art, it is understandable that some well-known structures and descriptions thereof may be omitted in the drawings. The technical solutions of the present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0059] ⅣExperimental results

[0060] Table 1 shows the experimental results of the 4 subsystems on the development data. It can be observed that fusing PPP features at the feature level improves the performance. Compared with the MFCCi-vector subsystem (EER=6.63%), the error rate of MFCC-PPPi-vector is reduced by 1.06%. On the other hand, the results of the OpenSmile feature are better than those of the MFCCi...