False Audio Detection Method and System Based on Complex Spectral Subband Fusion

By using the complex spectrum subband fusion method in the audio forgery detection technology, the low-frequency and high-frequency subbands are modeled and fused by the complex spectrum characteristics and logarithmic power spectrum characteristics, the problem of high false audio detection error rate in the prior art is solved, and higher detection accuracy and system performance are achieved.

CN114238849BActive Publication Date: 2025-06-27ANHUI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111481834.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-06
Publication Date
2025-06-27
Estimated Expiration
2041-12-06

AI Technical Summary

Technical Problem

Existing audio forgery detection technology is difficult to effectively utilize phase information and frequency-dividing band processing, resulting in a high error rate of false audio detection.

Method used

The detection method based on complex spectral subband fusion is adopted, and the complex spectral characteristics and logarithmic power spectral characteristics of speech waveforms are extracted, and the low-frequency and high-frequency subbands are modeled respectively, and the results are fused through the first- and second-level fusion algorithms.

Benefits of technology

It significantly reduces the error rate of false audio detection, improves the performance of anti-spoofing systems, and improves the accuracy of audio forgery detection technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114238849B_ABST
    Figure CN114238849B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for detecting fake audio based on complex spectrum sub-band fusion, comprising the following steps: S1: Extract the complex spectrum features and log power spectrum features of the original speech waveform as input features; S2: Divide the complex spectrum features into two sub-band frequency bands, namely low frequency and high frequency, model the high and low frequency bands of the complex spectrum respectively, and obtain corresponding prediction results; S3: Model the low frequency band features of the log power spectrum and obtain corresponding prediction results; S4: Fuse the prediction results of the high and low sub-band frequency bands of the complex spectrum through a primary fusion algorithm to obtain a primary fusion result; S5: Fuse the prediction result obtained from the low frequency band features of the log power spectrum and the primary fusion result through a secondary fusion algorithm to obtain a final result. A system for detecting fake audio based on complex spectrum sub-band fusion is also disclosed. The present invention can significantly improve the accuracy of audio forgery detection technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio forgery detection, and in particular to a method and system for detecting fake audio based on complex spectrum sub-band fusion. Background Art

[0002] An Automatic Speaker Verification (ASV) system is a typical biometric system, mainly used in fields such as access control, telephone banking, forensic evidence collection, and military reconnaissance. The system can use specific algorithms to perform pattern recognition and matching on the input speech to determine whether the speech of the speaker to be verified is the voice of a legitimate user. With the development of speech technology, the current ASV system faces various problems of fake audio attacks. Common fake audio can be divided into 4 forms: voice imitation, recording replay, speech synthesis, and voice conversion. Therefore, researchers have developed effective anti-spoofing systems to protect the ASV system from spoofing attacks by forged speech.

[0003] Audio forgery detection technology can effectively improve the performance of anti-spoofing systems. Current work mainly focuses on two aspects: 1) improving the acoustic features of audio; 2) designing new classification models. Compared with designing new models, extracting more representative features is particularly crucial. The magnitude spectrum and phase spectrum are two basic acoustic features obtained based on Fourier transform, reflecting different characteristics of audio. In early research, mainly the magnitude spectrum was processed while ignoring the phase information. Phase information is very important for the audio forgery detection task. Compared with the magnitude spectrum, the distribution of the phase spectrum is more irregular. Therefore, how to effectively utilize the phase information is a challenging problem. Currently, fake audio usually uses full-band information as features for modeling during the detection process. In fact, fake audio has different performances in the low-frequency band (0 - 4KHz) and high-frequency band (4 KHz - 8KHz). Sub-band processing can significantly reduce the error rate of fake audio detection and improve the performance of the anti-spoofing system.

[0004] Therefore, there is an urgent need to provide a new method for detecting fake audio to solve the above problems. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a method and system for detecting fake audio based on complex spectrum sub-band fusion, which can significantly improve the accuracy of audio forgery detection technology.

[0006] To solve the above technical problem, a technical solution adopted by the present invention is: to provide a method for detecting fake audio based on complex spectrum sub-band fusion, including the following steps:

[0007] S1: Extract the complex spectrum feature and logarithmic power spectrum feature of the original speech waveform as input features;

[0008] S2: Divide the complex spectral features into two sub - band frequency ranges, namely low - frequency and high - frequency. Model the high - frequency and low - frequency bands of the complex spectrum respectively, and obtain the corresponding prediction results;

[0009] S3: Model the features of the low - frequency band of the log - power spectrum and obtain the corresponding prediction results;

[0010] S4: Fuse the prediction results of the high - frequency and low - frequency sub - bands of the complex spectrum through a first - level fusion algorithm to obtain the first - level fusion result;

[0011] S5: Fuse the prediction result obtained from the low - frequency band features of the log - power spectrum and the first - level fusion result through a second - level fusion algorithm to obtain the final result.

[0012] In a preferred embodiment of the present invention, in step S1, extracting the complex spectral features of the original speech waveform includes the following steps:

[0013] S101: Use the Short - Time Fourier Transform (STFT) to convert the time - domain speech into the T - F domain:

[0014] X r [t, f]+i*X i [t, f]=STFT(X[k]) (1)

[0015] where x[k] represents the original speech waveform in the time domain, k is the time index of the speech signal, and are the corresponding real and imaginary parts of the STFT, t is the index of the number of time frames, and f is the index of the frequency unit;

[0016] S102: Concatenate the real - number part and the imaginary - number part of the STFT together to obtain the required complex spectral features, expressed as:

[0017] X complex =stack(X r , X i )∈R 2×F×T (2)

[0018] where stack represents the concatenation operation, F and T are the frequency and the number of time frames respectively.

[0019] In a preferred embodiment of the present invention, in step S1, extracting the log - power spectral features of the original speech waveform includes the following steps:

[0020] S111: Convert the original speech waveform into a complex spectrum through the Short - Time Fourier Transform (STFT):

[0021] D = STFT(X[k]) (3)

[0022] Among them, x[k] represents the original speech waveform in the time domain, k is the time index of the speech signal, and D represents the converted complex spectrum.

[0023] S112: Perform the absolute value and logarithm operations on the complex spectrum in sequence to obtain the logarithmic power spectrum feature:

[0024] LPS = log(abs(D)) (4)

[0025] Among them, abs and log respectively represent the absolute value and logarithm operations, and LPS is the required logarithmic power spectrum feature.

[0026] In a preferred embodiment of the present invention, the frequency range of the low-frequency sub-band is 0 - 4KHz, and the frequency range of the high-frequency sub-band is 4 - 8KHz.

[0027] In a preferred embodiment of the present invention, in step S4, the low-frequency sub-band of the complex spectrum feature is defined as The high-frequency sub-band of the complex spectrum feature is The first-level fusion formula is as follows:

[0028]

[0029] Among them, Distribution represents And The results obtained through the training and evaluation of the deep neural network classifier, S complex Is the first-level fusion result, and α is the weight coefficient of the first-level fusion algorithm.

[0030] In a preferred embodiment of the present invention, in step S5, the formula of the second-level fusion algorithm is as follows:

[0031]

[0032] Among them, Is the prediction result obtained by the deep neural network classifier for the low-frequency band feature of the logarithmic power spectrum, β is the weight coefficient of the second-level fusion algorithm, S complex Is the first-level fusion result, and S is the final result.

[0033] To solve the above technical problems, another technical solution adopted by the present invention is: to provide a false audio detection system for complex spectrum sub-band fusion, including:

[0034] A speech feature input module, configured to extract the complex spectrum feature and the logarithmic power spectrum feature of the original speech waveform as input features;

[0035] The complex spectrum feature processing module is used to divide the complex spectrum features into two sub-band frequency bands, namely low frequency and high frequency, model the high and low frequency bands of the complex spectrum respectively, and obtain the corresponding prediction results;

[0036] The logarithmic power spectrum processing module is used to model the low frequency band features of the logarithmic power spectrum and obtain the corresponding prediction results;

[0037] The primary fusion module is used to fuse the prediction results obtained by the complex spectrum feature processing module through the primary fusion algorithm to obtain the primary fusion result;

[0038] The secondary fusion module is used to fuse the prediction results obtained by the logarithmic power spectrum processing module and the results obtained by the primary fusion module through the secondary fusion algorithm to obtain the final result.

[0039] The beneficial effects of the present invention are as follows: By processing the false audio in frequency bands, the present invention can greatly reduce the error rate of false audio detection and improve the performance of the anti-spoofing system. Compared with the full-band system, fusing the low and high frequency sub-bands of the complex spectrum can improve the performance of the false audio detection system and the accuracy of the audio forgery detection technology. Brief Description of the Drawings

[0040] Figure 1 is a flowchart of the false audio detection method based on complex spectrum sub-band fusion of the present invention;

[0041] Figure 2 is a structural block diagram of the false audio detection system based on complex spectrum sub-band fusion. Detailed Embodiment

[0042] The following elaborates on the preferred embodiments of the present invention in conjunction with the drawings, so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby making a clearer and more definite definition of the protection scope of the present invention.

[0043] Please refer to Figure 1 and Figure 2 , the embodiments of the present invention include:

[0044] A false audio detection method based on complex spectrum sub-band fusion includes the following steps:

[0045] S1: Extract the complex spectrum features and logarithmic power spectrum (LPS) features of the original speech waveform as input features;

[0046] The steps for extracting the complex spectrum features of the original speech waveform include the following:

[0047] S101: Use the short-time Fourier transform STFT to convert the time-domain speech into the T-F domain:

[0048] Xr [t, f] + i * X i [t, f] = STFT(X[k]) (1)

[0049] Wherein, x[k] represents the original speech waveform in the time domain, k is the time index of the speech signal, and are the corresponding real and imaginary parts of the STFT, t is the index of the number of time frames, and f is the index of the frequency unit; Preferably, a Blackman Window with a window length of 1728 and a frame shift of 130 is used during STFT.

[0050] S102: Concatenate the real and imaginary parts of the STFT to obtain the required complex spectral features, expressed as:

[0051] X complex = stack(X r , X i ) ∈ R 2×F×T (2)

[0052] Wherein, stack represents the concatenation operation, F and T are the frequency and the number of time frames respectively.

[0053] Extracting the log power spectral features of the original speech waveform includes the following steps:

[0054] S111: Convert the original speech waveform into a complex spectrum through the short-time Fourier transform STFT:

[0055] D = STFT(X[k]) (3)

[0056] Wherein, x[k] represents the original speech waveform in the time domain, k is the time index of the speech signal, and D represents the converted complex spectrum.

[0057] S112: Perform the absolute value and logarithm operations on the complex spectrum in sequence to obtain the log power spectral features:

[0058] LPS = log(abs(D)) (4)

[0059] Wherein, abs and log represent the absolute value and logarithm operations respectively, and LPS is the required log power spectral feature.

[0060] S2: Divide the complex spectral features into two sub-band frequency ranges, low-frequency and high-frequency, model the high- and low-frequency bands of the complex spectrum respectively, and obtain the corresponding prediction results;

[0061] Preferably, the frequency range of the low-frequency sub-band is 0 - 4 KHz, and the frequency range of the high-frequency sub-band is 4 - 8 KHz. In this example, the low-frequency sub-band of the complex spectral features is defined as The high-frequency sub-band frequency range of the complex spectrum features is

[0062] It should be noted that in step S1, a complex spectrum with a dimension of 865 can be obtained. The first 0-433 dimensions are taken as the low-frequency sub-band, and the latter 433-865 are the high-frequency sub-band. Therefore, the sizes of the low-frequency and high-frequency sub-bands of the input complex spectrum features are 433×600 and 432×600 respectively. Then, the high- and low-frequency features of the complex spectrum are respectively used as the inputs of the deep neural network classifier, and a certain number of training rounds are set for the deep neural network classifier for training. Finally, the best model during training is selected for testing, and the obtained test results are used as the corresponding prediction results.

[0063] S3: Model the low-frequency band features of the logarithmic power spectrum and obtain the corresponding prediction results;

[0064] It should be noted that in step S3, first, by extracting the features of the logarithmic power spectrum from the original speech waveform, a logarithmic power spectrum with a dimension of 865 can be obtained. The first 0-433 dimensions are taken as the low-frequency band features of the logarithmic power spectrum, and this feature is used as the input of the deep neural network classifier. Then, a certain number of training rounds are set for the deep neural network classifier for training. Finally, the best model during training is selected for testing, and the obtained test results are used as the corresponding prediction results.

[0065] S4: Fuse the prediction results of the high- and low-frequency sub-band frequency ranges of the complex spectrum through a first-level fusion algorithm to obtain the first-level fusion result;

[0066] The first-level fusion formula is as follows:

[0067]

[0068] Among them, respectively represent and the results obtained through the training and evaluation of the deep neural network classifier, S complex is the first-level fusion result, and α is the weight coefficient of the first-level fusion algorithm, which is set to 0.5 in this example.

[0069] S5: Fuse the prediction results obtained from the low-frequency band features of the logarithmic power spectrum and the first-level fusion result through a second-level fusion algorithm to obtain the final result.

[0070] Since the LPS low-frequency band features are very effective for the audio forgery detection task, the formula of the second-level fusion algorithm in this example is as follows:

[0071]

[0072] Among them It is the prediction result obtained by the deep neural network classifier for the LPS low-frequency band features. β is the weight coefficient of the secondary fusion algorithm, which is set to 0.5 in this example, and S is the final result.

[0073] In the embodiments of the present invention, referring to Figure 2 , a false audio detection system for complex spectral sub-band fusion is further provided, including:

[0074] A voice feature input module, configured to extract the complex spectral features and log power spectral features of the original voice waveform as input features;

[0075] A complex spectral feature processing module, configured to divide the complex spectral features into two sub-band frequency bands, namely low-frequency and high-frequency bands, model the high- and low-frequency bands of the complex spectrum respectively, and obtain corresponding prediction results;

[0076] A log power spectrum processing module, configured to model the low-frequency band features of the log power spectrum and obtain corresponding prediction results;

[0077] A primary fusion module, configured to fuse the prediction results obtained by the complex spectral feature processing module through a primary fusion algorithm to obtain a primary fusion result;

[0078] A secondary fusion module, configured to fuse the prediction results obtained by the log power spectrum processing module and the results obtained by the primary fusion module through a secondary fusion algorithm to obtain a final result.

[0079] In the present invention, in order to quantitatively evaluate the results of different audio forgery detection systems, EER and the minimum normalized tandem detection cost function (min-tDCF) are used as evaluation metrics. EER is the error rate when the false rejection rate (FRR) and the false acceptance rate (FAR) are equal.

[0080] Table 1

[0081]

[0082] Table 1 shows the min-tDCF and EER results of different systems proposed based on the present invention. "L" and "H" respectively represent the low-frequency sub-band and the high-frequency sub-band, "Full" represents using the full-band information, and "+" represents the fusion operation. It can be seen from Table 1 that whether it is LPS or the complex spectrum, the low-frequency sub-band shows better performance than the high-frequency sub-band. At the same time, compared with the full-band system, fusing the low- and high-frequency sub-bands of the complex spectrum can improve the performance of the false audio detection system. Among them, the secondary fusion result of "Complex(L+H)+LPS(L)" shows the best performance.

[0083] These results verify that the complex spectral sub-band fusion system proposed by the present invention is very effective for the task of detecting fake audio. This is because the method proposed by the present invention is based on the complex spectrum and can make full use of speech information.

[0084] The above are only embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be included in the patent protection scope of the present invention by the same token.

Claims

1. A method for detecting fake audio based on complex spectral subband fusion, characterized in that, It includes the following steps: S1: Extract the complex spectrum features and log power spectrum features of the original speech waveform as input features; S2: Divide the complex spectrum features into two sub-band frequency bands, namely low-frequency and high-frequency bands. Model the high-frequency and low-frequency bands of the complex spectrum respectively and obtain the corresponding prediction results; S3: Model the low-frequency band features of the log power spectrum and obtain the corresponding prediction results; S4: Fuse the prediction results of the high and low sub-band frequency bands of the complex spectrum through a first-level fusion algorithm to obtain the first-level fusion result; define the low-frequency sub-band frequency band of the complex spectrum feature as the high-frequency sub-band frequency band of the complex spectrum feature as The first-level fusion formula is as follows: Among them, The distribution represents and The result obtained through the training and evaluation of a deep neural network classifier, S complex Is the first-level fusion result, and α is the weight coefficient of the first-level fusion algorithm; S5: Fuse the prediction results obtained from the low-frequency band features of the log power spectrum and the first-level fusion results through a second-level fusion algorithm to obtain the final result. The formula of the second-level fusion algorithm is as follows: Among them, is the prediction result obtained by the deep neural network classifier for the low-frequency band characteristics of the logarithmic power spectrum, β is the weight coefficient of the secondary fusion algorithm, and S complex is the primary fusion result, and S is the final result.

2. The method for detecting false audio based on complex spectral subband fusion according to claim 1, wherein In step S1, extracting the complex spectrum features of the original speech waveform includes the following steps: S101: Use the short-time Fourier transform (STFT) to convert the time-domain speech into the T-F domain; X r [t, f] + i * X i [t, f] = STFT(X[k]) (1) where \(x[k]\) represents the original speech waveform in the time domain, \(k\) is the time index of the speech signal, and are the corresponding real and imaginary parts of the STFT, \(t\) is the index of the time frame number, and \(f\) is the index of the frequency unit; S102: Concatenate the real part and the imaginary part of the STFT together to obtain the required complex spectrum features, expressed as: X complex = stack(X r , X i ) ∈ ℝ 2×F×T (2) where stack represents the concatenation operation, and F and T are the frequency and the number of time frames respectively.

3. The method for detecting false audio based on complex spectral sub-band fusion according to claim 1, characterized in that In step S1, extracting the log power spectrum features of the original speech waveform includes the following steps: S111: Convert the original speech waveform into a complex spectrum through the short-time Fourier transform (STFT): D = STFT(X[k]) (3) where x[k] represents the original speech waveform in the time domain, k is the time index of the speech signal, and D represents the converted complex spectrum; S112: Perform the absolute value operation and the logarithm operation on the complex spectrum in sequence to obtain the log power spectrum features: LPS = log(abs(D)) (4) where abs and log represent the absolute value operation and the logarithm operation respectively, and LPS is the required log power spectrum features.

4. The false audio detection method based on complex spectral sub-band fusion according to claim 1, characterized in that The frequency range of the low-frequency sub-band is 0 - 4 KHz, and the frequency range of the high-frequency sub-band is 4 - 8 KHz.

5. A false audio detection system for complex spectral sub-band fusion, characterized in that, It includes: A speech feature input module for extracting the complex spectrum features and log power spectrum features of the original speech waveform as input features; A complex spectrum feature processing module for dividing the complex spectrum features into two sub-band frequency bands, namely low-frequency and high-frequency bands, modeling the high-frequency and low-frequency bands of the complex spectrum respectively, and obtaining the corresponding prediction results; A log power spectrum processing module for modeling the low-frequency band features of the log power spectrum and obtaining the corresponding prediction results; The first-level fusion module is used to fuse the prediction results obtained by the complex spectral feature processing module through the first-level fusion algorithm to obtain the first-level fusion result; define the low-frequency sub-band frequency band of the complex spectral feature as The high-frequency sub-band frequency band of the complex spectral feature is The first-level fusion formula is as follows: Among them, The distribution represents and The result obtained through the training and evaluation of a deep neural network classifier, S complex is the first-level fusion result, and α is the weight coefficient of the first-level fusion algorithm; A second-level fusion module for fusing the prediction results obtained from the log power spectrum processing module and the results obtained from the first-level fusion module through a second-level fusion algorithm to obtain the final result. The formula of the second-level fusion algorithm is as follows: Among them, is the prediction result obtained by the deep neural network classifier for the low-frequency band features of the logarithmic power spectrum, β is the weight coefficient of the secondary fusion algorithm, and S complex is the primary fusion result, and S is the final result.

Citation Information

Patent Citations

  • Single-channel speech enhancement method based on joint dictionary learning and sparse representation

    CN111508518A

  • Single-channel speech enhancement method and system for neural network sub-band modeling, and storage medium

    CN111986660A