Voice spoofing detection method and system based on multistage mixed feature fusion

Through the speech spoofing detection method of multi-level hybrid feature fusion, the problem of performance degradation after mute segment removal is solved. Multi-level feature extraction and residual network fusion technology are used to improve the robustness and accuracy of speech spoofing detection and adapt to unknown attacks.

CN120299480APending Publication Date: 2025-07-11ANHUI UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510605445.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The performance of existing voice spoof detection methods has greatly decreased after removing the mute segment, making it difficult to adapt to unknown or new attacks, especially after voice activity detection, and the detection performance has been significantly reduced.

Method used

The speech spoof detection method using multi-level mixed feature fusion, including a multi-level mixed feature extraction unit, a mixed feature residual network unit and a feature classifier, is used to extract context-related features and fundamental frequency information features from the speech signal, and combine a self-supervised speech processing model and feature residual network to fusion information to reduce mute segment interference.

Benefits of technology

The excellent detection performance is maintained after removing the silent segment, which is significantly better than the baseline method, improving the feature extraction ability of non-silent segments, and improving the robustness and accuracy of speech spoof detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299480A_ABST
    Figure CN120299480A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of speech recognition, in particular to a speech deception detection method and system based on multistage mixed feature fusion. The invention provides a voice deception detection model with a new architecture, and the method comprises the steps: firstly extracting a mixed feature expression X1 mixed with a context-related feature XSSL and a fundamental frequency information feature XF0 through a multi-stage mixed feature extraction part, thereby enabling the model to be more focused on a non-silence segment while recognizing a deep voice feature; the X1 is subjected to information fusion in a channel, a time domain and a frequency domain through a mixed feature residual network part to obtain a fusion feature expression X2, so that the model can be guided to recognize forged clues in the voice more accurately; and finally, X2 is classified through a feature classifier to obtain a discrimination result. Through experimental comparison, even if a mute section is removed, the method still keeps excellent detection performance, and is obviously superior to a baseline method and other frontier models.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech recognition, and more specifically, to: 1. A speech spoofing detection method based on multi-level hybrid feature fusion; 2. A speech spoofing detection system using the speech spoofing detection method. Background Art

[0002] Current speech spoofing detection methods generally consist of two main stages: front-end feature extraction and back-end classification.

[0003] In early research, detection methods mainly relied on handcrafted features, such as traditional acoustic features like Mel Frequency Cepstral Coefficients (MFCC), and used machine learning algorithms like Gaussian Mixture Models (GMM) for classification. However, the effectiveness of such methods is often limited by the prior knowledge of specific attack types, and their generalization ability to unknown or new attacks is weak, making it difficult to adapt to evolving spoofing techniques.

[0004] With the rapid development of deep learning technology, spoofing detection methods based on deep neural networks (DNN) have gradually become a research hotspot. These methods can not only automatically learn more discriminative features but also improve the robustness of the system by optimizing the classification network. For example, architectures such as Convolutional Neural Networks (CNN) and Recurrent Neural Networks (RNN) are used to enhance the system's ability to model the spatio-temporal dependence of speech data.

[0005] However, the combination of traditional handcrafted features and deep learning still has certain limitations, especially in adapting to different types of spoofing attacks. To break through this bottleneck, recent research has begun to explore end-to-end speech spoofing detection methods, which directly extract forgery traces from the original waveform, avoiding the dependence on specific features.

[0006] Although neural network-based speech anti-spoofing methods perform well in detecting known attacks, they still suffer from insufficient robustness when dealing with unknown forgery methods. Specifically, after Voice Activity Detection (VAD) removes the silent segments, the performance of these methods will drop significantly, limiting their effectiveness in practical applications. Summary of the Invention

[0007] Based on this, in view of the problem that the performance of existing methods drops significantly after removing the language silent segments, a speech spoofing detection method and system based on multi-level hybrid feature fusion are provided.

[0008] The present invention is implemented by the following technical solutions:

[0009] In a first aspect, the present invention discloses a speech spoofing detection method based on multi-level hybrid feature fusion, including the following steps:

[0010] Obtain the current voice signal Voice0;

[0011] Perform voice activity detection processing on Voice0 to remove silent segments, and obtain the current input voice Voice1;

[0012] Use the trained voice spoofing detection model to process Voice1 and output the discrimination result.

[0013] Among them, the sample data set used by the voice spoofing detection model during training is the sample voice with silent segments removed.

[0014] The voice spoofing detection model includes: a multi-level hybrid feature extraction unit, a hybrid feature residual network unit, and a feature classifier. The multi-level hybrid feature extraction unit is used to: extract the hybrid feature representation X1 that mixes the context-related feature X SSL , fundamental frequency information feature X F0 from Voice1. The hybrid feature residual network unit is used to: perform information fusion on X1 in the channel, time domain, and frequency domain to obtain the fused feature representation X2. The feature classifier is used to: classify X2 to obtain the discrimination result.

[0015] The voice spoofing detection method based on multi-level hybrid feature fusion implements the method or process according to the embodiments of the present disclosure.

[0016] In a second aspect, the present invention discloses a voice spoofing detection system based on multi-level hybrid feature fusion, which uses the voice spoofing detection method based on multi-level hybrid feature fusion disclosed in the first aspect.

[0017] The voice spoofing detection system based on multi-level hybrid feature fusion includes: an audio acquisition module, an audio preprocessing module, and an audio detection module.

[0018] The audio acquisition module is used to: obtain the current voice signal Voice0. The audio preprocessing module is used to: perform voice activity detection processing on Voice0 to remove silent segments, and obtain the current input voice Voice1. The audio detection module is built with a trained voice spoofing detection model and is used to: use the trained voice spoofing detection model to process Voice1 and output the discrimination result.

[0019] The voice spoofing detection system based on multi-level hybrid feature fusion implements the method or process according to the embodiments of the present disclosure.

[0020] Compared with the prior art, the present invention has the following beneficial effects:

[0021] 1. In view of the problem that the performance of existing methods drops significantly after removing the silent segments of language, the present invention provides a voice spoofing detection model with a new architecture. First, a multi-level hybrid feature extraction unit extracts a hybrid feature representation X1 that mixes context-related features X SSL and fundamental frequency information features X F0 , so that the model focuses more on non-silent segments while identifying deep voice features; then, a hybrid feature residual network unit performs information fusion on X1 in the channel, time domain, and frequency domain to obtain a fused feature representation X2, to guide the model to more accurately identify forged clues in the voice; finally, a feature classifier classifies X2 to obtain a discrimination result. Through experimental comparison, even after removing the silent segments, the present invention still maintains excellent detection performance, significantly superior to the baseline method and other cutting-edge models.

[0022] 2. The sample data set used by the voice spoofing detection model of the present invention during training is the sample voice with silent segments removed, which can reduce the interference caused by silent information. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0024] Figure 1 It is a flowchart of the voice spoofing detection method based on multi-level hybrid feature fusion proposed in Embodiment 1 of the present invention;

[0025] Figure 2 For Figure 1 the structure diagrams of the pre-trained self-supervised speech processing model and the F0 sub-band feature extractor in

[0026] Figure 3 For Figure 1 the structure diagram of the post-processing feature unit in

[0027] Figure 4 For Figure 1 the structure diagram of the hybrid feature residual network unit in DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.

[0029] It should be noted that when a component is referred to as "installed on" another component, it can be directly on the other component or there may be an intermediate component. When a component is considered to be "set on" another component, it can be directly set on the other component or there may be an intermediate component at the same time. When a component is considered to be "fixed to" another component, it can be directly fixed to the other component or there may be an intermediate component at the same time.

[0030] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this invention belongs. The terms used in the description of the present invention herein are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The term "or / and" used herein includes any and all combinations of one or more of the related listed items.

[0031] Embodiment 1

[0032] Please refer to Figure 1 , Figure 1 , which is a flowchart of the voice spoofing detection method based on multi-level hybrid feature fusion in Embodiment 1.

[0033] It should be noted that the present invention provides a voice spoofing detection model with a new architecture - in order to highlight the advantages of the model, the actual input voice is the voice with the silent segments removed.

[0034] Generally speaking, the voice spoofing detection method based on multi-level hybrid feature fusion includes the following steps:

[0035] Obtain the current voice signal Voice0;

[0036] Perform voice activity detection (VAD) processing on Voice0 to remove the silent segments and obtain the current input voice Voice1;

[0037] Process Voice1 using the trained voice spoofing detection model and output the discrimination result.

[0038] The core of the present invention lies in providing a voice spoofing detection model with a new architecture - refer to Figure 1 , and this voice spoofing detection model can be divided into: a multi-level hybrid feature extraction part, a hybrid feature residual network part, and a feature classifier.

[0039] The following is a detailed introduction to each part:

[0040] I. The multi-level hybrid feature extraction part is used to: extract the mixed context-related feature X from Voice1SSL 1. The fundamental frequency (F0) information feature X F0 is combined to form the hybrid feature representation X1.

[0041] The multi-level hybrid feature extraction unit aims to make the model focus more on non-silent segments while identifying deep speech features. This is because the differences in silent segments (especially at the beginning and end of speech) are often used to distinguish real speech from spoof speech, but this dependence may cause the model to learn silent features related to the target label rather than the spoof features of the speech itself. Therefore, when designing the model in the present invention, it is made to focus on non-silent segments to extract more comprehensive anti-spoofing features.

[0042] As Figure 2 shown, the multi-level hybrid feature extraction unit includes: a pre-trained self-supervised speech processing model, an F0 sub-band feature extractor, and a post-processing unit for features.

[0043] 101. The pre-trained self-supervised speech processing model is used to: extract X from Voice1 SSL .

[0044] The pre-trained self-supervised speech processing model preferably adopts the Wav2vec 2.0 self-supervised learning framework to learn powerful representations from speech audio. In this Embodiment 1, the pre-trained self-supervised speech processing model adopts the XLS-R model - which is an advanced model pre-trained on a large-scale cross-lingual corpus and has powerful feature representation capabilities.

[0045] As Figure 2 shown, the XLS-R model adopts a two-stage feature extraction mechanism: First, Voice1 is converted from the time-domain raw waveform signal to a hidden feature sequence with rich time-frequency information through the CNN layer; then the Transformer layer is used to deeply process the hidden feature sequence to generate an output sequence with context awareness ability and use it as X SSL . That is to say, the XLS-R model can effectively capture the local features and global dependencies in the speech signal, making X SSL have rich context-related features - including the semantic information and temporal dependencies of the speech.

[0046] 102. The F0 sub-band feature extractor is used to: extract X from Voice1 F0 ;

[0047] Referring to the above, since the silent segment usually lacks F0 information, and the F0 feature contains key discriminative information for speech spoofing detection. Therefore, the present invention considers extracting the F0 feature and then combining it with X SSLPerform fusion to direct the focus of the model to the non - silent segments. However, since the F0 feature is represented as a single dimension and cannot be directly used as an input feature, the F0 feature is regarded as sub - bands, and the F0 sub - band feature extractor is used for extraction to obtain X F0 。

[0048] As Figure 2 shown, the F0 sub - band feature extractor is designed to include: a domain transformation layer, an LPS calculation layer, a one - dimensional convolutional layer, and a fully - connected layer.

[0049] Ⅰ. The domain transformation layer is used to: convert Voice1 to the time - frequency domain through the short - time Fourier transform.

[0050] Among them, Voice1 can be represented as the original waveform signal X[k] in the time domain, and k represents the time index of the waveform in the time domain. Then, the processing process of the domain transformation layer can be expressed by the formula:

[0051] STFT(X[k]) = X r [t,f]+jX i [t,f];

[0052] In the formula, STFT(.) represents the short - time Fourier transform; X r [t,f] represents the real part in the time - frequency domain; X i [t,f] represents the imaginary part in the time - frequency domain; t represents the time index in the time - frequency domain; f represents the frequency index in the time - frequency domain.

[0053] Ⅱ. The LPS calculation layer is used to: first calculate the log - power spectrum LPS of the full - band based on the output of the domain transformation layer full , and then extract the log - power spectrum LPS full corresponding to the F0 sub - band from LPS f0 。

[0054] Among them, the calculation formula of LPS full is:

[0055]

[0056] Since the frequency corresponding to the F0 feature is only a part of the full - band, then this frequency range is regarded as the F0 sub - band, and LPS full is extracted from LPS f0 . For example, if the F0 feature is concentrated between 0 - 400 Hz, then 0 - 400 Hz is the F0 sub - band, and there is:

[0057] Ⅲ. The one - dimensional convolutional layer is used to: perform one - dimensional convolutional processing on LPS f0 。

[0058] The fully connected layer is used to: adjust the dimension of the output of the convolutional layer to obtain X F0 .

[0059] The processing procedures of the above two layers can be expressed by the formula:

[0060] X F0 = linear(conv1d(LPS f0 ));

[0061] In the formula, linear(.) represents the fully connected layer; conv1d(.) represents the one-dimensional convolutional layer.

[0062] It should be noted that the X F0 obtained after the processing of the above two layers has the same specification as X SSL .

[0063] 103. The post-processing feature unit is used to: fuse X SSL and X F0 into X1.

[0064] The post-processing feature unit aims to combine X SSL and X F0 : X SSL has rich context-related features, and X F0 reflects the pitch contour and prosody features of the speech and provides important pitch-related clues; the combination of the two not only makes full use of the knowledge representation ability obtained by pre-training the self-supervised speech processing model on a large-scale cross-lingual corpus, but also enhances the modeling ability of the pitch-related information in the speech signal by introducing the F0 feature. This multi-modal feature fusion strategy enables the entire model to more accurately distinguish speech segments from non-speech segments, thus significantly improving the feature extraction effect of non-silent segments.

[0065] As Figure 3 shown, the post-processing feature unit is designed to include: a stacking layer, a transposed layer, a channel adjustment layer, a max pooling layer, a batch normalization layer, and an activation function layer.

[0066] The stacking layer is used to: stack X SSL and X F0 . The transposed layer is used to: perform transpose processing on the output of the stacking layer. The channel adjustment layer is used to: add a channel dimension to the output of the transposed layer. The max pooling layer is used to: perform max pooling processing on the output of the channel adjustment layer. The batch normalization layer is used to: perform batch normalization processing on the output of the max pooling layer. The activation function layer is used to: process the output of the batch normalization layer through the SELU activation function to obtain X1.

[0067] The processing procedure of the above post-processing feature unit can be expressed by the formula:

[0068] X1 = SELU(BN(MaxPool(Add((X SSL + X F0 )))); T ))));

[0069] In the formula, SELU(.) represents the SELU activation function; BN(.) represents the batch normalization layer; MaxPool(.) represents the max pooling layer; Add(.) represents the channel adjustment layer.

[0070] Second, the hybrid feature residual network part is used to: perform information fusion on X1 in the channel, time domain, and frequency domain to obtain the fused feature representation X2.

[0071] Since the pre-trained self-supervised speech processing model learns general speech features from a large-scale speech dataset, and these features do not always effectively enhance the spoofing detection performance - because previous work cannot explore whether the features extracted by the self-supervised model are applicable to the speech spoofing detection task. Then, if X SSL is fed into subsequent tasks, redundant information is introduced - which actually needs to be removed.

[0072] The hybrid feature residual network part is designed to solve this problem. Refer to Figure 4 , the hybrid feature residual network part is designed to include: a channel feature fusion layer (Channel Feature Fusion, CFF), a time feature fusion layer (Spectral Feature Fusion, SFF), and a spectral feature fusion layer (Temporal Feature Fusion, TFF).

[0073] Ⅰ. The channel feature fusion layer is used to perform information fusion on X1 in the channel dimension.

[0074] This is because the mutual connection between the time patterns within each channel is crucial for effective speech spoofing detection. Therefore, information fusion needs to be performed first in the channel dimension. In this Embodiment 1, the channel feature fusion layer is preferably SENet (Squeeze-and-Excitation Network), which can model the correlation between feature channels to solve the redundancy of information within the channel.

[0075] Ⅱ. The time feature fusion layer is used to perform information fusion on the output of the channel feature fusion layer in the time domain dimension. The spectral feature fusion layer is used to perform information fusion on the output of the channel feature fusion layer in the frequency domain dimension.

[0076] It should be noted that the output W1 of the time feature fusion layer and the output W2 of the spectral feature fusion layer together constitute X2.

[0077] Since the pre-trained self-supervised speech processing model generally models the audio signal with a time window, signals at different time intervals may not carry important independent information, which may also lead to potential redundancy in the time-frequency domain features. Therefore, the time feature fusion layer and the spectral feature fusion layer adopt a design such as Figure 4 follows:

[0078] ①. The time feature fusion layer includes: 2 two-dimensional convolutional layers in the time domain, 2 activation function layers, 1 batch normalization layer, and 1 residual layer.

[0079] In the time feature fusion layer:

[0080] The first two-dimensional convolutional layer in the time domain is used to perform two-dimensional convolutional processing on the output of the channel feature fusion layer in the time domain; the first activation function layer is used to process the output of the first two-dimensional convolutional layer in the time domain through the SELU activation function; the batch normalization layer is used to perform batch normalization processing on the output of the activation function layer; the second two-dimensional convolutional layer in the time domain is used to perform two-dimensional convolutional processing on the output of the batch normalization layer in the time domain; the second activation function layer is used to process the output of the second two-dimensional convolutional layer in the time domain through the Softmax activation function; the residual layer is used to perform product processing on the output of the channel feature fusion layer and the output of the second activation function layer to obtain W1.

[0081] The processing process of the above time feature fusion layer can be expressed by the formula:

[0082]

[0083] In the formula, CFF(.) represents the time feature fusion layer; Softmax(.) represents the Softmax activation function; T_conv2d(.) represents the two-dimensional convolutional layer in the time domain; SELU(.) represents the SELU activation function; BN(.) represents the batch normalization layer.

[0084] ②. The spectral feature fusion layer includes: 2 two-dimensional convolutional layers in the frequency domain, 2 activation function layers, 1 batch normalization layer, and 1 residual layer.

[0085] In the spectral feature fusion layer:

[0086] The first two-dimensional frequency-domain convolutional layer is used for: performing two-dimensional convolutional processing in the frequency domain on the output of the channel feature fusion layer; the activation function layer is used for: processing the output of the first two-dimensional frequency-domain convolutional layer through the SELU activation function; the batch normalization layer is used for: performing batch normalization processing on the output of the activation function layer; the second two-dimensional frequency-domain convolutional layer is used for: performing two-dimensional convolutional processing in the frequency domain on the output of the batch normalization layer; the second activation function layer is used for: processing the output of the second two-dimensional frequency-domain convolutional layer through the Softmax activation function; the residual layer is used for: multiplying the output of the channel feature fusion layer and the output of the second activation function layer to obtain W2.

[0087] The processing process of the above spectral feature fusion layer can be expressed by the formula:

[0088]

[0089] In the formula, CFF(.) represents the time feature fusion layer; Softmax(.) represents the Softmax activation function; F_conv2d(.) represents the two-dimensional frequency-domain convolutional layer; SELU(.) represents the SELU activation function; BN(.) represents the batch normalization layer.

[0090] Thirdly, the feature classifier is used for: classifying X2 to obtain the discrimination result.

[0091] The innovation of this method is concentrated in the multi-level hybrid feature extraction part and the hybrid feature residual network part; while the feature classifier can adopt the existing classifier design.

[0092] In this Embodiment 1, the feature classifier of this method refers to the design in the Chinese invention patent with the application number 2023112084492 (a method and system for voice spoofing detection based on an indirect heterogeneous graph attention model) - including: a homogeneous attention network part, a heterogeneous attention network part, and an information fusion network.

[0093] Regarding X2 as the current channel feature map, its specification is S×T×C. S represents the height, T represents the width, and C represents the number of channels. Then, for the c-th channel, the current channel feature has S current frequency-domain features along the height direction and T current time-domain features along the width direction; c ∈ [1, C].

[0094] Generally speaking: the homogeneous attention network part is used for aggregating X2 respectively through the frequency-domain graph attention network and the time-domain graph attention network; the heterogeneous attention network part is used for aggregating the output of the homogeneous attention network part respectively through the direct aggregation type heterogeneous attention network and the indirect aggregation type heterogeneous attention network; the information fusion network is used for fusing the output of the heterogeneous attention network part and obtaining the discrimination result.

[0095] For the specific process of the above-mentioned feature classifier, please refer to the description of the Chinese invention patent with the application number 2023112084492, which will not be elaborated herein.

[0096] Based on the above-designed voice spoofing detection model, after training it, it can be applied to this method to achieve voice spoofing detection.

[0097] Example 2

[0098] The purpose of this Example 2 is to verify and illustrate the effect of the method in Example 1 (abbreviated as Ours) through experiments.

[0099] This Example 2 adopts three different training and detection methods:

[0100] Method 1: Train and evaluate the model on unprocessed speech;

[0101] Method 2: Train the model on unprocessed speech, but evaluate the model on VAD-processed speech;

[0102] Method 3: Train and evaluate the model on VAD-processed speech.

[0103] 1. Five existing models (FFT, CQT, CONV Layers, AASIST, W2V2) are introduced for comparison. Using the ASVspoof2019 LA dataset, training and detection are carried out through the above three methods, and the equal error rate (EER, the smaller the value, the better the performance) of each model is examined. The results are shown in Table 1.

[0104] Table 1 Comparison of EER results

[0105] Model Method 1 Method 2 Method 3 FFT 1.21 27.38 20.40 CQT 1.82 27.07 19.56 CONV Layers 1.80 26.63 23.04 AASIST 1.10 23.93 19.88 W2V2 1.12 14.34 9.28 Ours 0.30 6.95 5.82

[0106] As can be seen from Table 1, the performance of this method is the best among the 6 models in all three methods. Although there is a certain performance decline in this method under different methods, the decline range is much smaller than that of the other 6 models and is acceptable.

[0107] 2. The existing baseline MiaGATs model is introduced for comparison, and two evaluation metrics are adopted: equal error rate (EER), minimum normalized cumulative detection cost function (min t-DCF, the smaller the value, the better the performance).

[0108] First, training and detection are carried out through the above three methods on the ASVspoof2019 LA dataset, and the performance of the 2 models is examined. The results are shown in Table 2.

[0109] Table 2 Performance comparison

[0110]

[0111] Then, training and detection were carried out on the ASVspoof2021 LA dataset through the above three methods, and the performances of 2 models were investigated. The results are shown in Table 3.

[0112] Table 3 Performance Comparison

[0113]

[0114] As can be seen from Table 2 and Table 3, the two evaluation indicators of this method under the three methods are better than those of MiaGATs, showing stronger robustness and accuracy.

[0115] In addition, the above experimental verification also shows that it is more recommended to use Method 3 to obtain the trained voice spoofing detection model required in Example 1. That is to say, the voice spoofing detection model used in Example 1 uses the sample speech dataset with the silent segments removed during training, so as to reduce the interference caused by silent information.

[0116] Example 3

[0117] This Example 3 discloses a voice spoofing detection system based on multi-level hybrid feature fusion, which uses the voice spoofing detection method based on multi-level hybrid feature fusion disclosed in Example 1.

[0118] The voice spoofing detection system based on multi-level hybrid feature fusion includes: an audio acquisition module, an audio preprocessing module, and an audio detection module.

[0119] The audio acquisition module is used to: acquire the current voice signal Voice0. The audio preprocessing module is used to: perform voice activity detection processing on Voice0 to remove the silent segments and obtain the current input voice Voice1. The audio detection module is built-in with a trained voice spoofing detection model and is used to: process Voice1 using the trained voice spoofing detection model and output a discrimination result.

[0120] Since the voice spoofing detection system based on multi-level hybrid feature fusion disclosed in this Example 3 uses the voice spoofing detection method based on multi-level hybrid feature fusion in Example 1, it also has the advantages of Example 1 and will not be repeated here.

[0121] Example 4

[0122] This Example 4 discloses a computer device, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps of the voice spoofing detection method based on multi-level hybrid feature fusion disclosed in Example 1 are implemented.

[0123] Among them, the computer device can be: a mobile terminal or a fixed terminal. The former is, for example: a mobile phone, a notebook computer, a digital broadcast receiver, a PDA (Personal Digital Assistant), a PAD (Portable Application Description: tablet computer), a PMP (Portable Media Player), a vehicle-mounted terminal (such as a vehicle-mounted navigation terminal), etc.; the latter is, for example: a digital TV, a desktop computer, etc.

[0124] Embodiment 4 of the present disclosure also discloses a readable storage medium storing computer program instructions, which, when read and executed by a processor, perform the steps of the voice spoofing detection method based on multi-level hybrid feature fusion disclosed in Embodiment 1.

[0125] Among them, the readable storage medium may include, but is not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM) or a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0126] Embodiment 4 of the present disclosure also discloses a computer program product including a computer program. When the computer program is executed by a processor, the steps of the voice spoofing detection method based on multi-level hybrid feature fusion disclosed in Embodiment 1 are implemented.

[0127] It should be noted that the above computer program can be written in one or more programming languages or a combination thereof. Among them, the programming languages include object-oriented programming languages - such as Java, Smalltalk, C++, and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The above computer program can be executed entirely on the user's computer, partially on the user's computer, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN).

[0128] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0129] The above-described embodiments only represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.

Claims

1. A voice spoofing detection method based on multi-level hybrid feature fusion, characterized in that, It includes the following steps: Obtain the current voice signal Voice0; Perform voice activity detection processing on Voice0 to remove the silent segments, and obtain the current input voice Voice1; Use the trained voice spoofing detection model to process Voice1 and output the discrimination result; Among them, the voice spoofing detection model includes: A multi-level hybrid feature extraction unit, which is used to: extract a hybrid feature representation X1 that mixes context-related feature X SSL , fundamental frequency information feature X F0 from Voice1; A hybrid feature residual network part, which is used to: perform information fusion on X1 in the channel, time domain, and frequency domain to obtain a fused feature representation X2; and A feature classifier, which is used to: classify X2 to obtain the discrimination result.

2. The voice spoofing detection method based on multi-level hybrid feature fusion according to claim 1, wherein The multi-level hybrid feature extraction part includes: A pre-trained self-supervised speech processing model for extracting X from Voice1 SSL ; F0 sub-band feature extractor, which is used to: extract X from Voice1 F0 ; And Post - feature processing unit, which is used to: combine X SSL and X F0 into X1.

3. The method for voice spoofing detection based on multi-level hybrid feature fusion according to claim 2, wherein The pre-trained self-supervised speech processing model adopts the XLS-R model.

4. The method for voice spoofing detection based on multi-level hybrid feature fusion according to claim 2, wherein The F0 sub-band feature extractor includes: A domain transformation layer, which is used to: convert Voice1 to the time-frequency domain through short-time Fourier transform; LPS calculation layer, which is used to: first calculate the log power spectrum LPS of the full frequency band based on the output of the domain transformation layer full , and then from LPS full extract the log power spectrum LPS corresponding to the F0 sub-band f0 ; A one-dimensional convolutional layer for performing one-dimensional convolutional processing on LPS f0 and Fully connected layer, which is used to: adjust the dimension of the output of the convolutional layer to obtain X F0 .

5. The method for voice spoofing detection based on multi-level hybrid feature fusion according to claim 2, wherein The post-processing feature part includes: Overlay layer, which is used for: superimposing X SSL , X F0 for superimposition; A transpose layer, which is used to: perform transpose processing on the output of the stacking layer; A channel adjustment layer, which is used to: add a channel dimension to the output of the transpose layer; A max pooling layer, which is used to: perform max pooling processing on the output of the channel adjustment layer; A batch normalization layer, which is used to: perform batch normalization processing on the output of the max pooling layer; and An activation function layer, which is used to: process the output of the batch normalization layer through the SELU activation function to obtain X1.

6. The voice spoofing detection method based on multi-level hybrid feature fusion according to claim 1, wherein The hybrid feature residual network part includes: A channel feature fusion layer, which is used to perform information fusion on X1 in the channel dimension; A time feature fusion layer, which is used to perform information fusion on the output of the channel feature fusion layer in the time domain dimension; and A spectrum feature fusion layer, which is used to perform information fusion on the output of the channel feature fusion layer in the frequency domain dimension; Among them, the output W1 of the time feature fusion layer and the output W2 of the spectrum feature fusion layer jointly form X2.

7. The voice spoofing detection method based on multi-level hybrid feature fusion according to claim 6, wherein The channel feature fusion layer adopts SENet.

8. The voice spoofing detection method based on multi-level hybrid feature fusion according to claim 6, characterized in that The time feature fusion layer includes: 2 two-dimensional convolutional layers in the time domain, 2 activation function layers, 1 batch normalization layer, 1 residual layer; In the time feature fusion layer: The first two-dimensional convolutional layer in the time domain is used to: perform two-dimensional convolutional processing in the time domain on the output of the channel feature fusion layer; the first activation function layer is used to: process the output of the first two-dimensional convolutional layer in the time domain through the SELU activation function; the batch normalization layer is used to: perform batch normalization processing on the output of the activation function layer; the second two-dimensional convolutional layer in the time domain is used to: perform two-dimensional convolutional processing in the time domain on the output of the batch normalization layer; the second activation function layer is used to: process the output of the second two-dimensional convolutional layer in the time domain through the Softmax activation function; the residual layer is used to: perform product processing on the output of the channel feature fusion layer and the output of the second activation function layer to obtain W1; The spectrum feature fusion layer includes: 2 two-dimensional convolutional layers in the frequency domain, 2 activation function layers, 1 batch normalization layer, 1 residual layer; In the spectrum feature fusion layer: The first two-dimensional convolutional layer in the frequency domain is used for: performing two-dimensional convolutional processing in the frequency domain on the output of the channel feature fusion layer; the activation function layer is used for: processing the output of the first two-dimensional convolutional layer in the frequency domain through the SELU activation function; the batch normalization layer is used for: performing batch normalization processing on the output of the activation function layer; the second two-dimensional convolutional layer in the frequency domain is used for: performing two-dimensional convolutional processing in the frequency domain on the output of the batch normalization layer; the second activation function layer is used for: processing the output of the second two-dimensional convolutional layer in the frequency domain through the Softmax activation function; the residual layer is used for: performing a product process on the output of the channel feature fusion layer and the output of the second activation function layer to obtain W2.

9. The voice spoofing detection method based on multi-level hybrid feature fusion according to claim 1, characterized in that The feature classifier includes: The homogeneous attention network part, which is used to aggregate X2 through the frequency-domain graph attention network and the time-domain graph attention network respectively; The heterogeneous attention network part, which is used to aggregate the output of the homogeneous attention network part through the direct aggregation type heterogeneous attention network and the indirect aggregation type heterogeneous attention network respectively; and The information fusion network, which is used to fuse the output of the heterogeneous attention network part and obtain the discrimination result.

10. A voice spoofing detection system based on multi-level hybrid feature fusion, characterized in that, It uses the voice spoofing detection method based on multi-level hybrid feature fusion described in any one of claims 1-9; The voice spoofing detection system based on multi-level hybrid feature fusion includes: An audio acquisition module, which is used for: acquiring the current voice signal Voice0; An audio preprocessing module, which is used for: performing voice activity detection processing on Voice0 to remove the silent segments and obtain the current input voice Voice1; and An audio detection module, which is built with a trained voice spoofing detection model and is used for: processing Voice1 using the trained voice spoofing detection model and outputting the discrimination result.

Citation Information

Cited By

  • Voice anti-spoofing detection method and device

    CN121354594A