Voice detection method based on logarithmic graph Fourier transform feature extraction

Through the combination of logarithmic Fourier transform and graph convolution network, the high-order structural information of speech signals is extracted, which solves the problem of difficulty in capturing the high-order structure of speech signals in the prior art, and significantly improves the performance of playback speech detection.

CN119993192AActive Publication Date: 2025-05-13NANJING UNIV OF POSTS & TELECOMM
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510160462.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-05-13
Estimated Expiration
2045-02-13

AI Technical Summary

Technical Problem

The existing playback speech detection method is difficult to effectively capture the high-order structural information implicit in the voice signal, especially in complex speech data scenarios, resulting in limited detection performance.

Method used

By modeling the speech signal into a graph structure, and using the logarithmic graph Fourier transform to extract frequency domain features, and combining the graph convolutional network to capture higher-order structural information, a more robust playback speech detection system is built.

Benefits of technology

It significantly improves the discriminant and robustness of playback speech detection, can more comprehensively describe the complex dynamic structure of speech signals, and enhances the comprehensiveness and accuracy of feature representation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993192A_ABST
    Figure CN119993192A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice verification, in particular to a voice detection method based on Fourier transform feature extraction of a logarithmic graph, which comprises the following steps of: constructing a translation operator of a voice graph, describing attenuation of a dependency relationship between voice samples by using an exponential function, generating a Laplacian matrix of the graph, and extracting the Fourier transform feature of the logarithmic graph. And representing the voice signal as an undirected graph to capture an intra-frame and inter-frame structural relationship; the method comprises the following steps: mapping a sample value of a voice signal into a graph node signal, converting the voice signal from a time domain to a graph frequency domain, extracting frequency domain features, and forming enhanced feature representation by synchronously combining intra-frame and inter-frame oscillation analysis and combining time domain features; and generating a detection score to judge whether the voice signal has a playback attack or belongs to normal voice. According to the invention, by introducing logarithmic graph Fourier transform and graph signal processing methods, the limitation of the prior art in playback voice detection is effectively solved, and the comprehensiveness and discrimination capability of feature extraction and the performance of a detection system are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech verification, in particular to a speech detection method based on logarithmic graph Fourier transform feature extraction. Background Art

[0002] In speech signal processing, playback speech detection is an important task, which aims to distinguish real speech from playback speech. Playback speech is often used by attackers as a way to deceive voice authentication systems, which poses a serious threat to the security of voice authentication systems. In this field, traditional methods mostly rely on the spectral characteristics and statistical properties of speech signals to build detection models. However, these methods usually have difficulty in effectively capturing the high-order structural information implicit in speech signals, especially in scenarios involving complex speech data.

[0003] In recent years, feature extraction methods based on spectrum analysis, such as Mel-frequency cepstral coefficients and constant-Q cepstral coefficients, have been widely used in playback speech detection tasks. These methods analyze the spectral characteristics of speech signals and extract low-order features that reflect the time-frequency information of speech for training classification models. However, these low-order features are difficult to fully describe the complex dynamic structure of speech signals, thus limiting the performance of the detection system.

[0004] In order to solve this problem, researchers began to try to use graph signal processing methods to model the structural information of speech signals. Graph signal processing theory provides a new perspective. By modeling speech signals as graph structures, the dependencies of speech signals in time and space dimensions can be captured. Among them, the method based on Graph Fourier Transform (GFT) can effectively extract the implicit features of speech signals by performing frequency domain analysis on graph signals. In particular, log-Graph Fourier Transform (log-GFT) as an improved method can better capture the local and global features in speech signals and has a higher discriminant power for the detection of playback speech.

[0005] In addition, the detection performance is also closely related to the robustness of the selected features and the ability of the classification model. In recent years, deep learning models, especially those based on Convolutional Neural Network (CNN) and Graph Convolutional Network (GCN), have shown excellent capabilities in processing complex speech signals. As a deep learning model that can process graph-structured data, GCN can extract high-order structural information of speech signals layer by layer by aggregating the feature information of nodes and their neighbors. This makes it possible to build a more robust playback speech detection system.

[0006] Although these technologies have achieved certain results in playback speech detection, there are still many challenges. For example, existing methods often ignore the dependencies between local frames of speech signals, and thus fail to fully capture the spatiotemporal dynamic characteristics of the signal. At the same time, since the difference between playback speech and real speech may be relatively subtle, the performance of existing systems is still limited when faced with complex environments and diverse speech data.

[0007] In summary, how to extract implicit features of speech signals by combining logarithmic graph Fourier transform and capture high-order structural information using graph convolutional networks has become a key technical challenge to improve the performance of playback speech detection. This provides a clear direction for subsequent research and is also of great significance for applications in the field of speech security. Summary of the invention

[0008] The purpose of the present invention is to overcome the above problems or at least partially solve the above problems, and to propose a speech detection method based on logarithmic graph Fourier transform feature extraction.

[0009] To achieve the above object, the present invention provides the following technical solution: a speech detection method based on logarithmic graph Fourier transform feature extraction, comprising the following steps:

[0010] S1: Define the graph adjacency matrix based on the dependencies between speech samples, construct the translation operator of the speech graph through forward translation and backward translation operations, use the exponential function to describe the attenuation of the dependencies between speech samples, generate the Laplacian matrix of the graph, and represent the speech signal as an undirected graph to capture the structural relationships within and between frames;

[0011] S2: Map the sample values ​​of the speech signal to graph node signals, apply the logarithmic graph Fourier transform using the graph Laplacian matrix, transform the speech signal from the time domain to the graph frequency domain, extract the frequency domain features, and synchronously merge the intra-frame and inter-frame oscillation analysis to form an enhanced feature representation combined with the time domain features;

[0012] S3: Use the pre-trained classification model to process the enhanced feature vector, generate a detection score to determine whether the speech signal has a playback attack, and calculate the performance indicators t-DCF and EER to evaluate the system performance.

[0013] In a preferred embodiment, in step S1, the forward and backward shift operations of the speech signal are used, and the speech graph shift operator is expressed as

[0014] In a preferred embodiment, in step S1, the exponential function describing the decay of the dependency relationship between speech samples is expressed as:

[0015]

[0016] Where δ∈(0,1) is the experimental constant and ζ is the experimental threshold.

[0017] In a preferred embodiment, in step S1, the graph Laplacian matrix L of the speech in the frame e It can be expressed as L e =D e -A e , the Laplacian matrix L of the inter-frame speech graph f It can be expressed as L f =D f -A f , where the diagonal matrix is ​​designated as D e =∑ j A e (i,j) and D f =∑ j A f (i, j), the undirected graph representation of the speech graph signal is ζ=(v,L e ,L f ).

[0018] In a preferred embodiment, in step S2, the intra-frame graph Laplacian matrix L is used. em For the mth frame of speech Y m The Fourier projection of the graph is: where ψ e Indicates L e The left singular characteristic matrix of the intra-frame graph Laplacian matrix L is decomposed by singular value decomposition method. em Decompose it, that is:

[0019] SVD(L em )=ψ e ×∧ e ×Γ e

[0020] where ∧ e represents a diagonal matrix, Γ e represents the right singular eigenmatrix.

[0021] In a preferred embodiment, in step S2, by synchronously merging A em (m=1,2,3) and A f To analyze the oscillation between frames and intra-frames, it is expressed as:

[0022]

[0023] where ψ f Indicates L f The left singular characteristic matrix, that is, SVD(L f )=ψ f ×∧ f×Γ f , ∧ f represents a diagonal matrix, Γ f represents the right singular eigenvalue matrix, y m The inverse logarithmic joint graph Fourier transform is

[0024] In a preferred embodiment, in step S2, the dynamic features of the speech signal, including the frequency and intensity variation patterns of the speech, are further extracted by synchronously combining the intra-frame and inter-frame oscillation analysis.

[0025] Compared with the prior art, the present invention has the following beneficial effects:

[0026] 1. The present invention can effectively extract the high-order structural information hidden in the speech signal by modeling the speech signal as a graph structure and using the logarithmic graph Fourier transform to perform frequency domain analysis on the speech signal. Compared with the traditional spectrum feature extraction method, the present invention can better describe the complex dynamic structure of the speech signal, thereby improving the discriminative power of playback speech detection;

[0027] 2. The present invention can synchronously capture the dependency of speech signals within and between frames by constructing intra-frame graph Laplacian matrix and inter-frame graph Laplacian matrix, and using graph Fourier transform to project speech signals. This method overcomes the defect of ignoring local inter-frame dependency in the prior art, fully extracts the spatiotemporal dynamic characteristics of speech signals, and enhances the comprehensiveness and robustness of feature representation;

[0028] 3. The present invention decomposes the graph Laplace matrix by singular value decomposition, and uses the left singular feature matrix to perform graph Fourier projection on the speech signal, which can extract more discriminative frequency domain features. At the same time, combined with the inverse logarithmic joint graph Fourier transform, the accuracy and stability of feature representation are further enhanced, providing high-quality feature input for subsequent classification models;

[0029] In summary, the present invention effectively solves the limitations of the existing technology in playback speech detection by introducing logarithmic graph Fourier transform and graph signal processing methods, significantly improves the comprehensiveness of feature extraction, the discrimination ability and the performance of the detection system, and at the same time provides a new technical direction for the field of speech security, which has important theoretical value and practical application significance. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 Schematic diagram of feature extraction in an embodiment of the present invention. DETAILED DESCRIPTION

[0031] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0032] See also Figure 1 The present invention provides a technical solution: a speech detection method based on logarithmic graph Fourier transform feature extraction, characterized in that it includes a graph representation construction stage S1, a feature extraction stage S2 and a detection stage S3, and the graph representation construction stage S1 includes the following steps:

[0033] S1.1: The graph adjacency matrix is ​​defined using the dependencies between speech samples, and the translation operators of the speech graph are constructed through forward translation and backward translation operations, which capture the relationships between speech samples.

[0034] S1.2: Use an exponential function to describe the attenuation of the dependencies between speech samples, and combine the intra-frame and inter-frame adjacency matrices to generate the Laplacian matrix of the graph.

[0035] S1.3: Based on the generated graph Laplacian matrix, the speech signal is represented as an undirected graph, capturing the structural relationships within and between frames.

[0036] The feature extraction stage S2 includes the following steps:

[0037] S2.1: Based on the constructed graph representation, each sample value of the speech signal is mapped to a graph node signal, and the node signal is represented as a time series feature of the speech.

[0038] S2.2: Use the graph Laplacian matrix obtained in step S1.2 to transform the speech signal, apply the logarithmic graph Fourier transform, transform the speech signal from the time domain to the graph frequency domain, and extract the frequency domain features of the speech sample.

[0039] S2.3: Capture high-order structural information in speech signals in the frequency domain, especially implicit features related to speaker identity, including the frequency and intensity variation patterns of speech.

[0040] S2.4: Combine the frequency domain features and time domain features of the graph and form an enhanced feature representation through feature concatenation. These features more comprehensively express the phonetic properties and contextual information of the speech signal.

[0041] S2.5: Standardize or quantify the generated features so that they can be input into the subsequent anonymization system or attacker system for speech processing and analysis.

[0042] The detection stage S3 includes the following steps:

[0043] S3.1: Using the pre-trained classification model, the enhanced feature vector is input and the hidden layer representation of the classifier is extracted to further capture the fine-grained speech signal features.

[0044] S3.2: The hidden layer output of the classifier is used to generate a detection score, which measures whether the input speech signal has a playback attack or is normal speech, and the performance indicators t-DCF (Tandem Detection Cost Function) and EER (Equal Error Rate) are calculated to evaluate the robustness and accuracy of the detection system.

[0045] In the specific implementation, in step S1.1, the forward and backward shift operations of the speech signal are used, and the speech graph shift operator is expressed as

[0046] In specific implementation, in step S1.2, the exponential function describing the attenuation of the dependency relationship between speech samples is expressed as:

[0047]

[0048] Where δ∈(0,1) is the experimental constant and ζ is the experimental threshold.

[0049] In specific implementation, in step S1.3, the graph Laplacian matrix L of the speech in the frame e It can be expressed as L e =D e -A e , the Laplacian matrix L of the inter-frame speech graph f It can be expressed as L f =D f -A f , where the diagonal matrix is ​​designated as D e =∑ j A e (i,j) and D f =∑ j A f (i,j), using the designed L e and L f , the undirected graph of the speech graph signal is represented as ζ=(v,L e ,L f ).

[0050] In specific implementation, in step S2.2, the intra-frame graph Laplacian matrix L is used. em For the mth frame of speech Y m The Fourier projection of the graph is: where ψ e Indicates Le The left singular characteristic matrix of the intra-frame graph Laplacian matrix L is decomposed by singular value decomposition method. em Decompose it, that is:

[0051] SVD(L em )=ψ e ×∧ e ×Γ e

[0052] where ∧ e represents a diagonal matrix, Γ e represents the right singular eigenmatrix.

[0053] In specific implementation, in step S2.4, by synchronously merging A em (m=1,2,3) and A f To analyze the oscillation between frames and intra-frames, it is expressed as:

[0054]

[0055] where ψ f Indicates L f The left singular characteristic matrix, that is, SVD(L f )=ψ f ×∧f×Γ f , ∧ f represents a diagonal matrix, Γ f represents the right singular eigenvalue matrix, y m The inverse logarithmic joint graph Fourier transform is

[0056] The present invention provides a speech detection method based on logarithmic graph Fourier transform feature extraction, which realizes efficient analysis and detection of speech signals by constructing a speech graph representation, applying logarithmic graph Fourier transform to extract frequency domain features, and combining with a pre-trained classification model for detection. The specific implementation method describes in detail the steps of the graph representation construction phase, feature extraction phase and detection phase, including the definition of the graph adjacency matrix, the generation of the graph Laplacian matrix, the application of the logarithmic graph Fourier transform and the detection process of the classification model. The present invention can capture the high-order structural information of the speech signal, especially the implicit features related to the speaker identity, and form an enhanced feature representation by synchronously merging the intra-frame and inter-frame oscillation analysis, which significantly improves the accuracy and robustness of speech detection.

[0057] The technical solution of the present invention has a wide range of application scenarios, such as voice anti-counterfeiting detection, playback attack detection and speaker recognition. Compared with traditional methods, the present invention realizes comprehensive analysis and efficient detection of voice signals by combining graph signal processing technology and machine learning models.

[0058] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A speech detection method based on logarithmic graph Fourier transform feature extraction, characterized in that: The following steps are involved: S1: Define the graph adjacency matrix based on the dependencies between speech samples, construct the translation operator of the speech graph, use the exponential function to describe the attenuation of the dependencies between speech samples, generate the Laplacian matrix of the graph, and represent the speech signal as an undirected graph to capture the structural relationships within and between frames; S2: Map the sample values ​​of the speech signal to graph node signals, apply the logarithmic graph Fourier transform using the graph Laplacian matrix, transform the speech signal from the time domain to the graph frequency domain, extract the frequency domain features, and synchronously merge the intra-frame and inter-frame oscillation analysis to form an enhanced feature representation combined with the time domain features; S3: Use the pre-trained classification model to process the enhanced feature vector, generate a detection score to determine whether the speech signal has a playback attack, and calculate the performance indicators t-DCF and EER to evaluate the system performance.

2. The method for speech detection based on logarithmic graph Fourier transform feature extraction according to claim 1, characterized in that: In step S1, the translation operator of the speech graph is constructed by forward and backward shift operations of the speech signal, and the speech graph shift operator is expressed as:

3. The method for speech detection based on logarithmic graph Fourier transform feature extraction according to claim 2, characterized in that: In step S1, the exponential function describing the decay of the dependency relationship between speech samples is expressed as: Where δ∈(0,1) is the experimental constant and ζ is the experimental threshold.

4. The method for speech detection based on logarithmic graph Fourier transform feature extraction according to claim 3, characterized in that: In step S1, the graph Laplacian matrix L of the speech in the frame e It can be expressed as L e =D ej -A e , the Laplacian matrix L of the inter-frame speech graph f It can be expressed as L f =D f -A f , where the diagonal matrix is ​​designated as D e =∑ j A e (i,j) and D f =∑ j A f (i, j), the undirected graph representation of the speech graph signal is ζ=(v,L e ,L f ).

5. The method for speech detection based on logarithmic graph Fourier transform feature extraction according to claim 4, characterized in that: In step S2, the intra-frame graph Laplacian matrix L is used em For the mth frame of speech Y m The Fourier projection of the graph is: where ψ e Indicates L e The left singular characteristic matrix of the intra-frame graph Laplacian matrix L is decomposed by singular value decomposition method. em Decompose it, that is: SVD(L em )=ψ e ×∧ e ×Γ e where ∧ e represents a diagonal matrix, Γ e represents the right singular eigenmatrix.

6. The method for speech detection based on logarithmic graph Fourier transform feature extraction according to claim 5, characterized in that: In step S2, by merging A synchronously em (m=1,2,3) and A f To analyze the oscillation between frames and intra-frames, it is expressed as: where ψ f Indicates L f The left singular characteristic matrix, that is, SVD(L f )=ψ f ×∧ f ×Γ f , ∧ f represents a diagonal matrix, Γ f represents the right singular eigenvalue matrix, y m The inverse logarithmic joint graph Fourier transform is 7. A method for speech detection based on logarithmic graph Fourier transform feature extraction according to any one of claims 1 to 6, characterized in that: In step S2, the dynamic features of the speech signal, including the frequency and intensity variation patterns of the speech, are further extracted by synchronously combining the intra-frame and inter-frame oscillation analysis.

Citation Information

Patent Citations

  • Method and device for encoding / decoding video signal by using optimized conversion based on multiple graph-based model

    CN108353193A

  • Playback attack detection method and training method of corresponding detection model

    CN110718229A

  • Coding method, device, decoding method, device, equipment and readable storage medium

    CN113766229A

  • Voice enhancement method and device based on neural network, and electronic equipment

    CN113808607A

  • Speech emotion recognition method and system based on adaptive adjacency matrix

    CN115631770A