Anonymous speaker attack method based on graph convolutional network

By extracting the high-order structural information of F0 features in the speech signal in the graph convolution network and splicing it with the original features, combining the features extracted by the acoustic model, anonymous speech is generated and the attacker system is trained, which solves the problem of difficulty in capturing time dynamic features in the existing technology, and significantly improves the performance of the attacker system.

CN120048241AActive Publication Date: 2025-05-27NANJING UNIV OF POSTS & TELECOMM +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510192231.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-05-27
Estimated Expiration
2045-02-21

AI Technical Summary

Technical Problem

Existing voice anonymization technology is difficult to effectively capture the temporal dynamic characteristics in the speaker's identity information, resulting in poor results in the attack model restoring the speaker's identity.

Method used

Using a graph convolutional network-based method, anonymous voice is generated by extracting high-order structural information of F0 features and splicing it with the original F0 features, combining the ASR-BN features extracted by the acoustic model, and the attacker system is trained using the generated anonymous data.

Benefits of technology

By considering the temporal correlation between different frames of the F0 feature, the attacker system's attack ability against anonymous voice is significantly improved, and the speaker's identity information can be restored more effectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048241A_ABST
    Figure CN120048241A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of language transformation, in particular to an anonymous speaker attack method based on a graph convolutional network. Comprising the following steps: splicing and fusing an F0 feature and an original F0 feature as a new F0 feature; extracting features of the audio and performing vector quantization; splicing the processed F0 feature and the feature to generate an anonymized voice; calculating speaker embedding from the test utterance and the registered utterance; outputting similarity scores of anonymization test speech embedding and anonymization registration speech embedding, and judging whether the anonymization test speech embedding and the anonymization registration speech embedding belong to the same speaker or not according to the scores; evaluating the attack capability of an attacker system on an anonymization system by taking error rates such as a plurality of tests and registration word pairs and calculation as performance indexes; the time correlation between different frames of the F0 feature is considered, and the graph convolutional network and the F0 feature are utilized to cooperate with anonymous speaker identity information, so that the performance of an attacker system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of language conversion, and particularly to an anonymous speaker attack method based on a graph convolutional network. Background Art

[0002] Voice anonymization technology is a technology that processes voice signals to make the identity information of the speaker unidentifiable. In voice anonymization technology, the content of the audio is usually modified or perturbed, while trying to retain the semantic content of the voice as much as possible to ensure that the voice can still be understood or used for specific tasks, but no longer directly associated with the original speaker. However, attackers will use various techniques to try to crack these anonymization measures, restore the speaker's identity or extract valuable personal information from the anonymized voice. The goal of an anonymized speaker attack system is to analyze the anonymized voice to restore or identify the identity information of the original speaker. Even if the voice has been anonymized, the attack system will try to crack the effect of the anonymization algorithm, restore the personalized features of the speaker, such as voiceprint information, the gender and age of the speaker, etc., or associate the anonymized voice with a specific speaker. These attacks are usually achieved through various technical means (such as deep learning, feature extraction, machine learning, etc.) and aim to bypass the anonymization measures and restore the original voice features being protected.

[0003] Although these technologies have achieved certain results in experiments, there are still some challenges. In previous attack methods, the correlation of speaker features in different time frames was rarely considered, and the attack models often ignored the complexity of the temporal dynamics in the voice signal. In fact, the voice signal is highly time-sequential, and the features in different time frames are not only related to the current voice content but may also carry the identity features of the speaker. For example, the pitch, speaking speed, intonation, etc. in the voice have strong correlations in time series. Ignoring the temporality of these features may cause the attack model to be unable to effectively capture the identity information of the speaker.

[0004] In the voice signal, F0 (fundamental frequency) is an important feature describing the periodicity of the sound. F0 reflects the pitch change in the voice signal and is of great significance for tasks such as speaker recognition, emotion recognition, and speech synthesis. The extraction of F0 usually relies on signal periodic analysis methods, such as the autocorrelation method, time-frequency analysis methods, etc.

[0005] In practical applications, the extraction of F0 features often relies on some classic audio analysis tools, such as Yaapt (Yet Another Algorithm for Pitch Tracking), which is a pitch tracking method based on autocorrelation analysis and can effectively extract accurate F0 feature trajectories from speech signals. F0 not only reflects the pitch characteristics of the speaker but is also closely related to information such as the speaker's emotion and intonation. Therefore, in tasks such as sentiment analysis and speech synthesis, F0 features are often used as an important feature for modeling.

[0006] Graph Convolutional Network (GCN) is a deep learning method specifically designed for processing graph-structured data. Different from traditional Convolutional Neural Networks (CNNs) that process regular grid data (such as images), GCN can effectively process data with irregular structures, such as social networks, molecular structures, knowledge graphs, etc. The core idea of GCN is to transmit information in the adjacency relationship of the graph through convolutional operations, enabling each node to combine information from adjacent nodes during multiple layers of convolution and learn the deep features of the nodes. The basic principle of GCN comes from graph signal processing theory. It defines the relationship between nodes in the graph through an adjacency matrix and obtains information from adjacent nodes through convolutional operations. By stacking multiple layers of GCN layers, GCN can aggregate the information of adjacent nodes layer by layer, thereby obtaining the high-order structural information of the data points.

[0007] Therefore, it is very necessary to propose an anonymous speaker attack method based on graph convolutional network that considers the temporal correlation between different frames of F0 features, uses graph convolutional network and F0 features to jointly anonymize speaker identity information, and improves the performance of the attacker's system. Summary of the Invention

[0008] The purpose of the present invention is to provide an anonymous speaker attack method based on graph convolutional network, aiming to consider the temporal correlation between different frames of F0 features, use graph convolutional network and F0 features to jointly anonymize speaker identity information, and improve the performance of the attacker's system.

[0009] To achieve the above object, an anonymous speaker attack method based on graph convolutional network adopted by the present invention includes the following steps:

[0010] Use graph convolutional network to extract the high-order structural information of the F0 features obtained by the feature extractor, and splice and fuse it with the original F0 features as the new F0 features;

[0011] Use an acoustic model to extract the ASR-BN features of the audio and perform vector quantization on them;

[0012] Concatenate the processed F0 features and ASR-BN features, and generate anonymized speech through a synthesizer;

[0013] Use the generated anonymized data to train the attacker's system;

[0014] For a given trial utterance and enrollment utterance, calculate the speaker embeddings from the trial utterance and the enrollment utterance, and compare the two;

[0015] Output the similarity score of the anonymized trial utterance embedding and the anonymized enrollment utterance embedding, and determine whether they belong to the same speaker based on the score;

[0016] Calculate the equal error rate as a performance metric through multiple pairs of trial and enrollment utterances to evaluate the attack ability of the attacker's system against the anonymization system.

[0017] Among them, in the step of using a graph convolutional network to extract the high-order structural information of the F0 features obtained by the feature extractor, and concatenating and fusing it with the original F0 features as the new F0 features:

[0018] The feature extractor extracts features from the input speech waveform through a pitch tracking algorithm, and normalizes it after extraction; the frame length for extraction is 35.0 ms, the frame shift is 20 ms, and the threshold is 0.25.

[0019] Among them, in the step of using a graph convolutional network to extract the high-order structural information of the F0 features obtained by the feature extractor, and concatenating and fusing it with the original F0 features as the new F0 features:

[0020] The embedding features output by the extractor are input into the graph convolutional network, with dimensions [B, T, D];

[0021] Where B is the batch size, T is the number of time frames, and D is the total feature dimension.

[0022] Among them, in the step of using an acoustic model to extract the ASR-BN features of the audio and perform vector quantization on it:

[0023] The acoustic model consists of a pre-trained model and three TDNN-F layers; the output dimension of the pre-trained model is 1024, and the dimension of the bottleneck layer is 256.

[0024] Among them, in the step of concatenating the processed F0 features and ASR-BN features, and generating anonymized speech through a synthesizer:

[0025] The new F0 feature containing high-order structural information after the F0 feature is processed by the graph convolutional network is concatenated with the original F0 feature, and its dimension is adjusted to be the same as the original output dimension through a linear layer, and then it is concatenated with the ASR-BN feature.

[0026] Among them, in the step of using the graph convolutional network to extract the high-order structural information of the F0 feature obtained by the feature extractor and concatenating and fusing it with the original F0 feature as the new F0 feature, the processing process of the graph convolutional network for the F0 feature is as follows:

[0027] Obtain the original F0 feature; where the dimension of the original F0 feature is (B, C, T), B is the batch size, C is the feature dimension, and T is the number of time steps;

[0028] In order to align with the number of time steps of the ASR-BN feature, interpolation processing is performed on the F0 feature; let the number of time steps of BN be T bn , the interpolated F0 feature The number of time steps will be adjusted to T bn :

[0029]

[0030] The dimension of the interpolated feature is (B, C, T bn );

[0031] Input the F0 feature into the graph convolutional network to obtain the graph feature F', that is:

[0032]

[0033] Among them, W (0) and W (1) are the weight matrices in the graph convolutional network, and A k is the adjacency matrix:

[0034]

[0035] The output feature F' is concatenated with the original feature to obtain the new F0 feature, that is:

[0036]

[0037] An anonymous speaker attack method based on graph convolutional network of the present invention adopts a process of speech anonymization, attacker system training and performance evaluation; wherein the speech anonymization includes: using a graph convolutional network to extract high-order structural information of the F0 feature obtained by a feature extractor, and splicing and fusing it with the original F0 feature as a new F0 feature; using an acoustic model to extract the ASR-BN feature of the audio and performing vector quantization on it; splicing the processed F0 feature and ASR-BN feature, and generating anonymized speech through a synthesizer. The attacker system training and performance evaluation includes: using the generated anonymized data to train the attacker system; for a given test utterance and enrollment utterance, calculating the speaker embeddings from the test utterance and the enrollment utterance and comparing the two; outputting the similarity score of the anonymized test utterance embedding and the anonymized enrollment utterance embedding, and judging whether they belong to the same speaker according to the score; calculating the equal error rate as a performance metric through multiple pairs of test and enrollment utterances to evaluate the attack ability of the attacker system on the anonymized system. By considering the temporal correlation between different frames of the F0 feature, the graph convolutional network and the F0 feature are used to jointly anonymize the speaker identity information to improve the performance of the attacker system. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.

[0039] Figure 1 is a flowchart of the steps of the anonymous speaker attack method based on graph convolutional network of the present invention.

[0040] Figure 2 is a schematic flowchart of the speech anonymization of the present invention.

[0041] Figure 3 is a schematic flowchart of the attacker system training and performance evaluation of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application.

[0043] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms "a", "the", and "said" used in this application and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0044] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0045] Please refer to Figures 1 to 3 , the present invention provides an anonymous speaker attack method based on a graph convolutional network, including the following steps:

[0046] S100: Use a graph convolutional network to extract the high-order structural information of the F0 feature obtained by a feature extractor, and splice and fuse it with the original F0 feature as a new F0 feature;

[0047] S200: Use an acoustic model to extract the ASR-BN feature of the audio and perform vector quantization on it;

[0048] S300: Splice the processed F0 feature and the ASR-BN feature, and generate anonymized speech through a synthesizer;

[0049] S400: Use the generated anonymized data to train the attacker system;

[0050] S500: For a given trial utterance and enrollment utterance, calculate the speaker embeddings from the trial utterance and the enrollment utterance and compare the two;

[0051] S600: Output the similarity score of the anonymized trial utterance embedding and the anonymized enrollment utterance embedding, and determine whether they belong to the same speaker according to the score;

[0052] S700: Calculate the equal error rate as a performance metric through multiple trial and enrollment utterance pairs to evaluate the attack ability of the attacker system against the anonymization system.

[0053] In this embodiment, a process of voice anonymization, attacker system training, and performance evaluation is adopted; wherein voice anonymization includes: using a graph convolutional network to extract the high-order structural information of the F0 feature obtained by a feature extractor, and splicing and fusing it with the original F0 feature as a new F0 feature; using an acoustic model to extract the ASR-BN feature of the audio and performing vector quantization on it; splicing the processed F0 feature and ASR-BN feature, and generating anonymized speech through a synthesizer. Among them, attacker system training and performance evaluation include: using the generated anonymized data to train the attacker system; for a given test utterance and enrollment utterance, calculating the speaker embeddings from the test utterance and enrollment utterance and comparing the two; outputting the similarity score between the anonymized test utterance embedding and the anonymized enrollment utterance embedding, and judging whether they belong to the same speaker according to the score; calculating the equal error rate as a performance metric through multiple pairs of test and enrollment utterances to evaluate the attack ability of the attacker system on the anonymization system. By considering the temporal correlation between different frames of the F0 feature, using a graph convolutional network and the F0 feature to collaboratively anonymize the speaker identity information to improve the performance of the attacker system.

[0054] In voice anonymization:

[0055] The feature extractor extracts features from the input speech waveform through a pitch tracking algorithm and normalizes it after extraction; wherein the frame length during extraction is 35.0 ms, the frame shift is 20 ms, and the threshold is 0.25.

[0056] The embedding feature output by the extractor is input into a graph convolutional network, and the dimension is [B, T, D];

[0057] where B is the batch size, T is the number of time frames, and D is the total feature dimension.

[0058] The acoustic model consists of a pre-trained model and three TDNN-F layers; wherein the output dimension of the pre-trained model is 1024, and the dimension of the bottleneck layer is 256.

[0059] The new F0 feature containing high-order structural information after the F0 feature is processed by the graph convolutional network is spliced and combined with the original F0 feature, and its dimension is adjusted to be the same as the original output dimension through a linear layer, and then it is spliced with the ASR-BN feature.

[0060] The processing process of the graph convolutional network for the F0 feature is as follows:

[0061] Obtain the original F0 feature; wherein the dimension of the original F0 feature is (B, C, T), B is the batch size, C is the feature dimension, and T is the number of time steps;

[0062] To align with the time steps of the ASR-BN features, interpolation processing is performed on the F0 features; let the time steps of BN be T bn , the interpolated F0 features will have their time steps adjusted to T bn :

[0063]

[0064] The dimension of the interpolated features is (B, C, T bn );

[0065] Input the F0 features into the graph convolutional network to obtain the graph features F', that is:

[0066]

[0067] where, W (0) and W (1) are the weight matrices in the graph convolutional network, and A k is the adjacency matrix:

[0068]

[0069] The output features F' are concatenated with the original features to obtain the new F0 features, that is:

[0070]

[0071] In this embodiment, the voice to be anonymized is prepared, and the training set includes the train-clean-360 subset of the LibriSpeech corpus. In addition to the provided anonymized training data, the original train-clean-360 data is also used. LibriSpeech is an English read speech corpus sourced from audiobooks and is designed specifically for ASR research. LibriSpeech contains 960 hours of speech sampled at 16 kHz, and this data will be used for ASV and ASR evaluations. IEMOCAP is an affective audio-visual dataset and will be used for SER evaluations. IEMOCAP contains 12 hours of speech sampled at 16 kHz, corresponding to improvised and scripted two-person conversations between 5 female and 5 male British actors.

[0072] Use the graph convolutional network to extract the high-order structural information of the F0 features obtained by the feature extractor, and concatenate and fuse it with the original F0 features as the new F0 features;

[0073] The processing of the F0 features by the graph convolutional network is the key in the whole process, and its construction process includes the following steps:

[0074] The dimension of the original F0 feature is (B, C, T), where: B is the batch size, C is the feature dimension, and T is the number of time steps.

[0075] To align with the number of time steps of the ASR-BN feature, the F0 feature needs to be interpolated. Let the number of time steps of BN be T bn , the interpolated F0 feature will have its number of time steps adjusted to T bn :

[0076]

[0077] The interpolated feature dimension becomes (B, C, T bn ).

[0078] The F0 feature is input into the graph convolutional network, i.e.:

[0079]

[0080] where, W (0) and W (1) are the weight matrices in the graph convolutional network, A k is the adjacency matrix. According to the voice graph shift operator, here it is considered that each frame has a potential relationship with the surrounding k frames, i.e.:

[0081] Thus, the graph feature F' is obtained.

[0082] The feature F' output by the graph convolutional network is then concatenated with the original feature to obtain a new F0 feature, i.e.:

[0083]

[0084] Use the wav2vec2.0_tdnnf acoustic model to extract the ASR-BN feature of the audio and perform vector quantization on it;

[0085] Concatenate the processed F0 feature and the ASR-BN feature and generate anonymized speech through the Hifi-GAN synthesizer.

[0086] After considering the specification and practicing the content disclosed herein, those skilled in the art will readily think of other implementation schemes of this application. This application aims to cover any variations, uses, or adaptive changes of this application, and these variations, uses, or adaptive changes follow the general principles of this application and include the common general knowledge or conventional technical means in the technical field not disclosed in this application.

[0087] It should be understood that this application is not limited to the exact structure already described and shown in the drawings, and various modifications and changes can be made without departing from its scope.

Claims

1. An anonymous speaker attack method based on graph convolutional network, characterized in that: The steps include: The graph convolutional network is used to extract the high-order structural information of the F0 feature obtained by the feature extractor, and it is concatenated and fused with the original F0 feature as a new F0 feature; Use the acoustic model to extract the ASR-BN features of the audio and perform vector quantization on them; The processed F0 features and ASR-BN features are concatenated and anonymized speech is generated through a synthesizer; Use the generated anonymized data to train the attacker’s system; For a given trial utterance and enrollment utterance, compute speaker embeddings from the trial utterance and enrollment utterance, and compare the two; Output the similarity scores of the anonymized test utterance embedding and the anonymized registered utterance embedding, and determine whether they belong to the same speaker based on the scores; Through multiple trials and registered utterance pairs, the error rate is calculated as a performance indicator to evaluate the attack capability of the attacker's system on the anonymization system.

2. The anonymous speaker attack method based on graph convolutional network as claimed in claim 1, characterized in that: In the step of using the graph convolutional network to extract the high-order structural information of the F0 feature obtained by the feature extractor, and concatenating it with the original F0 feature as a new F0 feature: The feature extractor extracts features from the input speech waveform through a pitch tracking algorithm, and normalizes the features after extraction; wherein the frame length during extraction is 35.0 ms, the frame shift is 20 ms, and the threshold is 0.

25.

3. The anonymous speaker attack method based on graph convolutional network as claimed in claim 1, characterized in that: In the step of using the graph convolutional network to extract the high-order structural information of the F0 feature obtained by the feature extractor, and concatenating it with the original F0 feature as a new F0 feature: The embedded features output by the extractor are input into the graph convolutional network with the dimension [B, T, D]; Where B is the batch size, T is the number of time frames, and D is the sum of feature dimensions.

4. The anonymous speaker attack method based on graph convolutional network as claimed in claim 1, characterized in that: In the step of extracting the ASR-BN features of the audio using the acoustic model and vector quantizing it: The acoustic model consists of a pre-trained model and three TDNN-F layers; the output dimension of the pre-trained model is 1024, and the dimension of the bottleneck layer is 256.

5. The anonymous speaker attack method based on graph convolutional network as claimed in claim 1, characterized in that: In the step of concatenating the processed F0 features and ASR-BN features and generating anonymized speech through a synthesizer: The new F0 feature containing high-order structural information after the F0 feature is processed by the graph convolutional network is concatenated with the original F0 feature, and its dimension is adjusted to be consistent with the original output dimension through a linear layer, and then concatenated with the ASR-BN feature.

6. The anonymous speaker attack method based on graph convolutional network as claimed in claim 1, characterized in that: In the step of using the graph convolutional network to extract the high-order structural information of the F0 feature obtained by the feature extractor and concatenating it with the original F0 feature as the new F0 feature, the graph convolutional network processes the F0 feature as follows: Get the original F0 feature; the dimension of the original F0 feature is (B, C, T), B is the batch size, C is the feature dimension, and T is the number of time steps; In order to align with the time step of ASR-BN feature, the F0 feature is interpolated; let the time step of BN be T bn , the F0 feature after interpolation The number of time steps will be adjusted to T bn : The feature dimension after interpolation is (B, C, T bn ); Input the F0 feature into the graph convolutional network to obtain the graph feature F', that is: Among them, W (0) and W (1) is the weight matrix in the graph convolutional network, A k is the adjacency matrix: The output feature F' is different from the original feature The new F0 feature is obtained by splicing, namely:

Citation Information

Patent Citations

  • Multi-modal emotion recognition method based on graph convolutional network

    CN116229225A

  • End-to-end speech synthesis method, device, equipment and medium

    CN116469375A

  • Intelligent access control management method and system based on multi-mode identification and Internet of Things technology

    CN118968665A

  • Voice changer

    US20220130372A1

  • Generation of optimized spoken language understanding model through joint training with integrated acoustic knowledge-speech module

    US20220230629A1