Physical domain identity camouflage system and method based on voiceprint recognition adversarial samples
By constructing a subphoneme-level perturbation dictionary and a real-time phoneme processor, generating and injecting adversarial samples with channel robustness and model migration capabilities, the problem that the prior art cannot effectively disguise in the real physical domain is solved, and efficient identity disguise for voiceprint recognition systems is achieved.
Patent Information
- Application Number
- CN202210423843.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-21
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2042-04-21
AI Technical Summary
Existing adversarial sample-based identity masquerading techniques cannot be effectively implemented in real physical domain scenarios, especially in the face of black box target systems and complex channel environments, lacking the transferability of streaming adversarial perturbation generation, channel interference resistance, and attack unknown models.
By constructing subphoneme-level perturbation dictionary, phoneme recognizer, adversarial sample generator, cross-channel enhancer, and integrated classifier in the offline training section, adversarial samples with channel robustness and model migration capabilities are generated, and real-time synchronous injection is leveraged in the online camouflage section using real-time phoneme aligners and real-time phoneme predictors.
It realizes effective identity camouflage of voiceprint recognition system in the real physical domain, improves channel robustness and model migration capabilities of adversarial perturbations, and ensures efficient camouflage performance in different devices and environments.
Smart Images

Figure CN114783447B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of voiceprint recognition and adversarial samples, and in particular to a physical domain identity camouflage system and method based on voiceprint recognition adversarial samples. Technical Background
[0002] In recent years, voice has become one of the most important human-computer interaction interfaces with its natural, human-centered experience. The rapid development and deployment in personal and production fields such as smart assistants, smart homes, and automatic navigation have further stimulated voiceprint recognition to become an emerging biometric technology. With the development of deep learning technology, voiceprint-based identity authentication technology has achieved significant performance improvements, which has led to the widespread application of voiceprint recognition technology. An industry research report shows that the global voiceprint recognition market size will reach US$10.7 billion in 2020 and is expected to exceed US$27.1 billion in 2026, fully demonstrating the broad development prospects of voiceprint recognition technology. However, these deep learning-based solutions are vulnerable to attacks based on adversarial samples due to their inherent neural network structural characteristics. Adversarial sample attacks are a type of usability attack that exploits the structural characteristics of the target neural network model and adds directional, human-imperceptible tiny perturbations to the input samples, causing the model numerical deviation to continue to amplify during the transmission process, ultimately leading to the target model making wrong decisions or even directional decisions. This phenomenon has attracted widespread public attention and a lot of research in academia. The present invention utilizes subphoneme-level adversarial perturbations to construct streaming adversarial perturbations that are synchronized with human voice in real time, and combines them with channel enhancement technology and mobility enhancement technology to enable the perturbations to mislead the black-box target system in the physical space, and ultimately achieve directional camouflage of the original human voice in the physical domain to deceive the voiceprint recognition system.
[0003] The existing identity disguise methods based on adversarial samples for voiceprint recognition systems (hereinafter referred to as "voiceprint systems") have been able to achieve speaker identity disguise attacks on voiceprint recognition systems based on mainstream adversarial sample generation methods. However, these technologies mainly rely on pure digital systems and white-box models with transparent structures. They do not fully consider the new challenges faced in real physical scenarios, and their implementation forms are not compatible with real application scenarios. The interference to physical systems is limited, which leads to the inability to be implemented in real physical domain scenarios. Specifically, the existing technologies lack the following three capabilities. (1) Streaming adversarial perturbation generation and injection: In order to avoid attracting the attention of people around, the disguiser cannot directly play adversarial samples, but should generate perturbations in real time based on human voices and physically inject them; (2) Channel interference resistance: The physical propagation process introduces complex channel interference into adversarial samples, so the attack needs to be cross-channel, that is, it can resist device and environmental interference; (3) Transferability of attack unknown models: In the real world, the implementation details of the target voiceprint system are often unknown to the disguiser, which indicates the existence of a strict black box setting. In summary, the existing identity disguise technology based on adversarial samples is too idealistic for the scenarios and cannot adapt to the needs of real physical domain scenarios. Summary of the invention
[0004] The present invention improves the technical solutions of the prior art and provides a physical domain identity disguise system and method based on voiceprint recognition adversarial samples. The present invention is implemented through the following technical solutions:
[0005] The present invention discloses a physical domain identity camouflage system based on voiceprint recognition adversarial samples, the system includes an offline training part and an online camouflage part:
[0006] The offline training part includes a subphone level perturbation dictionary, a phoneme recognizer, an adversarial sample generator, a voiceprint classifier, a system optimizer, and a training corpus. After the signal is input from the training corpus into the phoneme recognizer, an aligned speech with phoneme alignment information is output. The adversarial sample generator superimposes the subphone level perturbation in the subphone level perturbation dictionary onto the aligned speech input into the adversarial sample generator according to the phoneme information to generate an adversarial sample. The adversarial sample is forward propagated in the voiceprint classifier, and the output is back-propagated through the system optimizer to optimize the subphone level and perturbation dictionary.
[0007] The online camouflage part relies on a portable real-time camouflage device. The software composition of the camouflage device includes a phoneme recognizer, a real-time phoneme aligner and a real-time phoneme predictor. The real-time voice generated by the disguiser is input into the phoneme recognizer to generate a real-time phoneme sequence. The real-time phoneme aligner locates the phoneme at the current moment according to the real-time phoneme sequence, derives the phoneme sequence to be played, and inputs the real-time phoneme predictor. The real-time phoneme predictor determines the specific duration of each phoneme in the sequence based on the real-time phoneme sequence and the phoneme sequence to be played, and plays sub-phoneme-level perturbations synchronously with the human voice based on the predicted results, thereby synthesizing adversarial samples with channel robustness and model robustness in the physical domain, and finally realizing identity camouflage for the voiceprint recognition system.
[0008] As a further improvement, the offline training part described in the present invention also includes a cross-channel enhancer located between the sub-phoneme level perturbation dictionary and the adversarial sample generator, which is used to enhance the channel robustness of the sub-phoneme level perturbation. The cross-channel enhancer utilizes the Maximum Length Sequence signal to collect channel impulse response collection, and the collection process also includes different environments, different devices and different distance conditions.
[0009] As a further improvement, the voiceprint classifier described in the present invention is an integrated classifier, which is used to enhance the cross-model migration capability of subphone-level perturbations, and simultaneously input adversarial samples into multiple pre-trained voiceprint models with different model architectures and model training sets for forward propagation, and sum the output scores through a weighted average operation; wherein the model architecture is one or more of d-vector, x-vector and DeepSpeaker, the model training set is a subset of VoxCeleb1 / VoxCeleb2, and the weighted average operation uses an attention mechanism to dynamically adjust the weights of each model output during the iteration process; the camouflage device is an embedded device equipped with a microphone, a speaker, and a processing chip hardware device.
[0010] The present invention also discloses a physical domain identity disguise method based on voiceprint recognition adversarial samples, which specifically includes the following steps:
[0011] Offline training part:
[0012] 1) The subphone level perturbation dictionary provides a matching subphone level perturbation of 10-20ms length for each phoneme used, which is initialized to a random perturbation that conforms to the normal distribution;
[0013] 2) To enhance the perturbation’s ability to resist channel interference, before being superimposed on the speech, the subphone-level perturbation is augmented by a cross-channel enhancer that simulates the channel states of a variety of different devices and room environments based on previously collected channel impulse responses;
[0014] 3) In order to correctly superimpose subphonemic perturbations on speech, a phoneme recognizer is used to extract phoneme information from each item in the training set corpus. The phoneme information includes the phoneme type and start and end timestamps, and is aligned with the original speech.
[0015] 4) The adversarial sample generator superimposes the data-augmented subphoneme-level perturbation and the aligned speech by phoneme, and the superposition method is to repeatedly fill the subphoneme-level perturbation until the entire phoneme is filled, and outputs the adversarial sample;
[0016] 5) In order to improve the cross-model migration capability, adversarial samples are input into multiple voiceprint recognition models through an integrated classifier, and the outputs of multiple models are integrated into one in a weighted form;
[0017] 6) According to the recognition results of the integrated classifier, the system optimizer solves the system optimization problem and iteratively updates the sub-phoneme level perturbation dictionary, and finally obtains a trained sub-phoneme level perturbation dictionary.
[0018] Online camouflage section:
[0019] 7) The impersonator records the same speech as the text in advance and uses a phoneme recognizer to extract the phoneme sequence, including all the phonemes and duration in the speech, as a standard phoneme sequence for reference in the real-time impersonation process;
[0020] 8) The disguise process begins. The disguiser holds the disguise device and speaks the preset password. The disguise device receives the voice in real time through the microphone.
[0021] 9) The speech signal is recognized as a real-time phoneme sequence by a phoneme recognizer;
[0022] 10) The real-time phoneme sequence is aligned with a pre-given standard phoneme sequence by a real-time phoneme aligner, thereby obtaining the phoneme sequence to be played that the speaker will say next;
[0023] 11) The real-time phoneme predictor estimates the duration of the phonemes in the to-be-played phoneme sequence based on the real-time phoneme sequence, the standard phoneme sequence and the to-be-played phoneme sequence whose duration is to be estimated;
[0024] 12) According to the phoneme sequence and the duration of each phoneme, the camouflage device accurately plays the corresponding sub-phoneme-level adversarial perturbation through the speaker, and finally realizes the online synchronization process with the real-time voice, achieving the purpose of physical domain streaming camouflage attack.
[0025] As a further improvement, the system optimization problem used in step 6) of the present invention is as follows:
[0026]
[0027]
[0028] 13) Where x is the input sample, y t is the target label, P is the subsonic and perturbation dictionary, G(·,·) is the adversarial sample generator, the classifier outputs the confidence that the input sample is speaker i, θ is the upper limit of the threshold for the system to judge that the input sample is a disguised sample, and L s,θ (x,y t ) is the sample x judged as y under the system S with a threshold of θ t The loss function is To obtain the expected operation, μ, χ, and σ are the distribution of training samples, the distribution of channel impulse responses, and the model distribution used for the ensemble model, respectively, and α is the weight factor of each classifier in the ensemble classifier.
[0029] As a further improvement, the phoneme recognizer in step 3) and step 9) of the present invention is used to extract the phoneme sequence and its time information in the speech, based on the bidirectional RNN neural network model architecture, and its workflow is: the input audio is framed, and each frame is subjected to 26-dimensional MFCC extraction and then input into the bidirectional RNN network, the input layer is of size 26×1, the hidden layer uses a GRU unit of size 256, the output layer size is 40 (corresponding to the phonemes used), and the output result is processed by merging repeated phonemes and correcting erroneous phonemes to output a phoneme sequence.
[0030] As a further improvement, the implementation process of the real-time phoneme aligner described in the present invention is: first, a reference phoneme sequence is constructed from the phoneme information extracted from the pre-recorded text speech that is the same as the real-time speech, a phoneme recognizer is used to continuously perform phoneme recognition on the recorded real-time speech to obtain a real-time phoneme sequence, and a sliding window mechanism is used to align the recognized phoneme sequence with the reference phoneme sequence to confirm the phoneme sequence in the next speaker's speech.
[0031] As a further improvement, the sliding window mechanism described in the present invention utilizes long-term and short-term sliding windows, wherein the long-term sliding window is used to determine the search range of the short-term sliding window, and the short-term sliding windows exist in pairs in the reference phoneme sequence and the real-time phoneme sequence, and are used to compare the matching degree of the window contents.
[0032] As a further improvement, the implementation process of the real-time phoneme predictor described in the present invention is: based on the relative relationship between the real-time phoneme sequence and the benchmark phoneme sequence, the exponentially weighted moving average algorithm is used to estimate the current speaking rate, and then the phoneme duration in the benchmark phoneme sequence is inversely calculated with the current speaking rate to obtain the duration of the phonemes in the real-time speech.
[0033] The beneficial effects of the present invention are as follows:
[0034] The present invention proposes a physical domain identity disguise system and method based on voiceprint recognition adversarial samples. The existing technical method constitutes a one-time adversarial perturbation for a single voice, and needs to obtain a complete voice before generating the perturbation. It has the problems of poor real-time performance, poor universality, low generation efficiency, etc., and is not suitable for the real streaming physical domain attack scenario. The present invention innovatively proposes a real-time streaming disguise attack method that separates the perturbation and generation process from the application process, uses a real-time phoneme aligner and a real-time phoneme predictor to predict and locate the phonemes in the real-time voice, and generates fine-grained universal sub-phoneme-level adversarial perturbations at the phoneme level, so that the sub-phoneme-level adversarial perturbations generated once can be applied to the streaming voice in real time, and finally realize the disguise attack form adapted to the real physical domain scenario. In the real-time evaluation, the average time overhead of each real-time synchronization of the present invention is 0.11s, which shows that the synchronization mechanism of the present invention can achieve good real-time performance at a synchronization interval of 0.5s; the median of the phoneme delay is 50ms, and more than 75% of the phoneme delays are less than 100ms, which has good synchronization performance. The average phoneme hit rate of the real-time phoneme prediction mechanism of the present invention reaches 61.3%, and the phonemes and their occurrence times are effectively predicted.
[0035] In order to meet the need to reduce the perturbation generation process's need for known target model details, the existing technical method uses query-based gradient estimation technology to optimize adversarial perturbations. This method requires tens of thousands of query requests to the target system to train a single adversarial perturbation, which has serious inefficiency problems and is easily perceived by the target system, and is not practically operable. The present invention proposes an integrated classifier-based method to enhance the cross-model migration capability of adversarial perturbations. The perturbation generation process does not need to know the details of the target model, nor does it need to query the target system. By introducing voiceprint recognition models of various different architectures and data sets into the perturbation training process, the perturbation can effectively learn and discover security vulnerabilities that are susceptible to perturbations in the speech itself and in the voiceprint model, thereby having a generalized attack capability on the voiceprint recognition system and realizing a black box attack that is more in line with real attack scenarios. In the transferability evaluation, the present invention achieved 94.8%, 89.5% and 85.5% ASR under three different attack types: cross-training data, cross-model architecture, and cross-data and architecture, respectively, which are 27.9%, 29.4% and 37.2% higher than the single attack model, respectively, indicating that compared with a single model, the integrated model helps to improve the robustness of attacking different black-box models.
[0036] Existing technical methods generate adversarial perturbations for purely digital systems, which cannot resist channel interference from devices, environments, etc. during physical playback, thus causing serious performance degradation in physical domain scenarios. The present invention fully considers channel interference phonemes, introduces a cross-channel enhancer, and performs data augmentation on subphoneme-level adversarial perturbations from two dimensions: environment and device, effectively improving the channel robustness of the perturbation, and ultimately ensuring that it maintains good performance during physical playback. In the channel robustness evaluation, the present invention achieved 85.5%, 91.0%, 86.9% and 97.3% ASR on four device models, respectively, and the average ASR was 44.4% higher than the control group without enhancement; at the same time, the present invention achieved 90.5%, 91.3% and 88.5% ASR in three different room environments, respectively, which were 56.1%, 53.2% and 63.7% higher than the control group, which fully proves that the channel enhancement technology proposed in the present invention does improve the robustness of adversarial interference in physical camouflage.
[0037] In the comprehensive performance evaluation, the present invention achieves a lower MCD than the existing FakeBob, while the ASR is improved by 15.5%. It reaches 80.5%, 85.5% and 90.5% ASR on the three mainstream voiceprint recognition models of d-vector, x-vector and deepSpeaker respectively, which fully proves the effectiveness of the present invention on different systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is a system framework diagram of the present invention;
[0039] Figure 2 Generate example graphs for adversarial examples based on subphone-level perturbations;
[0040] Figure 3 Schematic diagram of integrated classifier;
[0041] Figure 4 is a schematic diagram of a real-time phoneme aligner;
[0042] Figure 5 The streaming synchronization performance diagram of the present invention;
[0043] Figure 6 A performance graph of streaming real-time camouflage of the present invention;
[0044] Figure 7 is a performance diagram of the present invention under different channel environments;
[0045] Figure 8 is a performance diagram of the present invention under different models;
[0046] Fig. 9 It is a noise and hearing test result diagram of the present invention. DETAILED DESCRIPTION
[0047] The technical solution of the present invention is further described below through specific implementation cases:
[0048] The present invention discloses a physical domain identity disguise system and method based on voiceprint recognition adversarial samples, which is carried out in the following attack scenario: the disguiser intends to impersonate a legitimate user to access a target voiceprint system device through a physical space in order to retrieve the user's voice information or activate sensitive voice commands to achieve the purpose of identity disguise. Assume that the disguiser is not registered in the target system, so under normal circumstances, it should be regarded as an illegal user and denied login. Limited by portability, the disguiser can only carry a small disguise device equipped with a speaker and a microphone for playing adversarial interference. In addition, the disguiser has no prior knowledge of the target voiceprint system, including voiceprint systems, signal processing technology, etc. This indicates that the disguiser should launch a black box attack. On the other hand, it is assumed that the disguiser can collect some voice samples of the target user from public conversations or social media for training. However, it should be clear that the collected samples do not need to cover the text used in further disguise. During the disguise process, the disguiser is not restricted by space, that is, there may be other people around the disguiser and the target voiceprint system. In order to avoid attracting the attention of others, the disguiser cannot play the voice of the legitimate user without live speech, which is unnatural and easily attracts the attention of people around.
[0049] The present invention discloses a physical domain identity camouflage system based on voiceprint recognition adversarial samples, the system comprising an offline training part and an online camouflage part. Figure 1The system framework diagram of the present invention. The offline training part includes a subphone level perturbation dictionary, a phoneme recognizer, an adversarial sample generator, a cross-channel enhancer, an integrated classifier, a system optimizer and a training corpus. The speech signal is input into the phoneme recognizer from the training corpus recorded in advance from the pretender, and then an aligned speech with phoneme alignment information is output; the adversarial sample generator augments the subphone perturbation in the subphone level perturbation dictionary with a cross-channel enhancer and then superimposes it into the aligned speech according to the phoneme information to generate an adversarial sample; the adversarial sample is forward propagated in the integrated classifier, and the output is back-propagated through the system optimizer to optimize the subphone level and perturbation dictionary. The online camouflage part relies on a portable real-time camouflage device, which includes a phoneme recognizer (with the same structure as the phoneme recognizer described in the offline training part), a real-time phoneme aligner and a real-time phoneme predictor. The real-time speech generated by the imposter is input into the phoneme recognizer to generate a real-time phoneme sequence. The real-time phoneme aligner locates the phoneme at the current moment according to the real-time phoneme sequence, derives the phoneme sequence to be played, and inputs the real-time phoneme predictor. The real-time phoneme predictor determines the specific duration of each phoneme in the sequence based on the real-time phoneme sequence and the phoneme sequence to be played, and plays sub-phoneme-level perturbations synchronously with the human voice based on the predicted results, thereby synthesizing adversarial samples with channel robustness and model robustness in the physical domain, and finally realizing identity disguise for the voiceprint recognition system.
[0050] The present invention discloses a physical domain identity camouflage method based on voiceprint recognition adversarial samples. The method is mainly divided into an offline training stage and an online camouflage stage, and the steps are as follows:
[0051] Offline training phase
[0052] 14) The subphone level perturbation dictionary provides a matching subphone level perturbation of 10-20ms length for each phoneme used, which is initialized to a random perturbation conforming to a normal distribution;
[0053] 15) To enhance the perturbation’s ability to resist channel interference, before being superimposed on the speech, the subphone level perturbation is augmented by a cross-channel enhancer that simulates the channel states of various devices and room environments based on previously collected channel impulse responses;
[0054] 16) In order to correctly superimpose the subphonemic level disturbance on the speech, a phoneme recognizer is used to extract phoneme information from each piece of the training set corpus, wherein the phoneme information includes the phoneme type and the start and end timestamps, and is aligned with the original speech to form speech;
[0055] 17) The adversarial sample generator superimposes the data-augmented subphoneme-level perturbation and the aligned speech by phoneme, and the superposition method is to repeatedly fill the subphoneme-level perturbation until the entire phoneme is filled, and outputs the adversarial sample;
[0056] 18) In order to improve the cross-model migration capability, adversarial samples are input into multiple voiceprint recognition models through an integrated classifier, and the outputs of multiple models are integrated into one in a weighted form;
[0057] 19) Based on the recognition results of the integrated classifier, the system optimizer calculates the target loss and iteratively updates the sub-phoneme level perturbation dictionary by back-propagating the gradient, and finally obtains a trained sub-phoneme level perturbation dictionary.
[0058] Online camouflage stage
[0059] 20) The impersonator records the speech of the same text in advance and uses a phoneme recognizer to extract the phoneme sequence, including all the phonemes and duration in the speech, as a standard phoneme sequence for reference in the real-time impersonation process.
[0060] 21) The disguise process begins. The disguiser holds the disguise device and speaks the preset password. The disguise device receives the voice in real time through the microphone.
[0061] 22) The speech signal is recognized as a real-time phoneme sequence by the phoneme recognizer
[0062] 23) The real-time phoneme sequence is aligned with the pre-given standard phoneme sequence by the real-time phoneme aligner, thereby obtaining the phoneme sequence to be played that the speaker will say next.
[0063] 24) The real-time phoneme predictor estimates the duration of the phonemes in the to-be-played phoneme sequence based on the real-time phoneme sequence, the standard phoneme sequence and the to-be-played phoneme sequence whose duration is to be estimated.
[0064] 25) According to the phoneme sequence and the duration of each phoneme, the camouflage device can accurately play the corresponding sub-phoneme-level adversarial perturbation, and finally realize the online synchronization process with the real-time voice, achieving the purpose of physical domain streaming camouflage attack.
[0065] The following is a detailed description of the subphone-level adversarial perturbation dictionary and its generation process, which does not yet include integrated classifiers and cross-channel enhancers. Subphone-level perturbations are trained, fixed-length, limited-amplitude audio noises that can be quickly generated as adversarial samples by superimposing them on corresponding phonemes in speech. The subphone-level perturbation dictionary is a dictionary container containing 40 subphone-level perturbations, where each perturbation corresponds one-to-one to the 39 phonemes shown in Table 1, and an additional perturbation corresponds to inter-speech pauses. The length of each perturbation is 10-20ms. Define P as a dictionary of subphone-level perturbations, with the key values being 40 phoneme symbols and the real values being the corresponding fixed-length audio perturbations. In order to interfere with the target voiceprint system S, first generate a speech-level adversarial sample x′ for the input audio x based on P. Specifically, assume that there are m different phonemes in x, corresponding to p0,…,p in the dictionary P.m-1 For each phoneme i, by repeating the corresponding subphone level perturbation p i And inject it into the corresponding position in x to generate adversarial sample x′, Figure 2 Generate example graphs for adversarial examples based on subphone-level perturbations; in order to launch a successful disguise, the generated adversarial example x′ should make it possible to target user y t The output score It should be the largest among all registered users and greater than the preset threshold θ. Therefore, the loss function is designed as follows:
[0066]
[0067] In order to generate P for phonemes from different voices so that it has generalized interference capabilities, the training corpus μ is further used for perturbation optimization instead of a single sample. The system optimization problem to be solved by the system optimizer is as follows:
[0068]
[0069] in, is the expectation of the loss function, G(x,P) represents the generating function of adversarial sample x′ generated from original audio x based on perturbation dictionary P, that is, x′=G(x,P), ∈ is the amplitude threshold. By solving the above optimization problem, a sub-phoneme-level adversarial perturbation dictionary can be obtained. Based on this dictionary, adversarial interference of speech with different text content can be achieved by injecting perturbations into each corresponding phoneme, thereby achieving real-time camouflage.
[0070] Table 1 Commonly used phonemes in speech recognition
[0071]
[0072] The following is a specific description of the cross-channel enhancer. In order to enhance the physical channel robustness of the above-mentioned adversarial perturbations, a cross-channel enhancer is added to the above-mentioned generation process to enhance the channel of the sub-phoneme-level perturbations. The basic idea of this method is to model the channel interference into a channel impulse response c, and then inversely obtain the signal after the channel interference through a convolution operation. In order to obtain the channel impulse response, the present invention adopts one of the basic acoustic measurement methods in the ISO standard to measure the impulse response based on the MLS signal. Specifically, an MLS signal (i.e., a specific pseudo-random binary signal) is played by a transmitter, then propagated in a specific environment, and finally received by a receiver. By performing a convolution operation on the received signal and the transmitted MLS signal, the CIR of the physical propagation system (including the transceiver device model and environment) can be derived. Using this method, the imposter only needs to collect one sample in each environment under each device model to generate different channel responses. After obtaining the unit impulse responses of different channels, before injecting the sub-phoneme-level perturbation into the speech, the adversarial perturbation at the speech level is first enhanced by a convolution operation. Therefore, the system optimization problem to be solved by the system optimizer is converted into the following form
[0073]
[0074] Y(·,c) is the speech enhanced with the unit impulse response c, and χ is a unit impulse response distribution collected from multiple real transceiver devices and multiple environments. By using the collected unit impulse responses to enhance the perturbation, it is possible to achieve successful identity masquerade across various device models and different environments, and finally initiate successful identity masquerade in the physical domain.
[0075] The following is a specific description of the integrated classifier. In addition to channel robustness, the present invention also discloses an integrated classifier technology, which enhances the portability of sub-phoneme-level adversarial perturbations to meet the needs of black-box attacks on different voiceprint systems. Specifically, this technology integrates the outputs of multiple voiceprint systems in the perturbation optimization process. When a sound sample x is input into the optimization process, the corresponding adversarial perturbation is generated based on P; then, the generated adversarial sample is input to n voiceprint models selected from a predefined voiceprint model set σ, rather than just one model, to obtain various outputs. After that, the n outputs are aggregated into a whole by weighted summation. Taking into account the different contributions of different voiceprint models in the set, the present invention further introduces an attention mechanism for adjusting the weights of each model in real time, that is, by iterating the attention coefficient Dynamically adjust each model S i The composition ratio of Figure 3 is a schematic diagram of an integrated classifier; therefore, the system optimization problem to be solved by the system optimizer is converted into the following form:
[0076]
[0077] By integrating learning and attention mechanisms, the calibrated perturbations expand their targets from a single voiceprint model to various voiceprint models, thus achieving black-box attack capabilities.
[0078] The system optimizer is described in detail below. The system optimizer used in the present invention is an Adam optimizer based on the Pytorch platform, which uses mini batch technology for training optimization.
[0079] The following is a detailed description of the phoneme recognizer. The phoneme recognizer is used to extract the phoneme sequence and the duration corresponding to each phoneme from the speech signal. The phoneme recognizer used in the offline training stage and the online camouflage stage belongs to the same neural network system, and its workflow is: based on the bidirectional RNN neural network model architecture, its workflow is: the input audio is framed, and 26-dimensional MFCC is extracted for each frame and then input into the bidirectional RNN network. The input layer is of size 26×1, and the hidden layer uses a GRU unit of size 256. The output layer size is 40 (corresponding to the phonemes used). The output result is processed by merging repeated phonemes and correcting erroneous phonemes before outputting the phoneme sequence.
[0080] The following is a specific description of the real-time phoneme aligner. After completing the generation of sub-phoneme-level, cross-channel and transferable adversarial interference, the pretender needs to use the disguise device to inject disturbances into real-time speech. The premise of real-time injection of disturbances is to accurately predict the type of subsequent phonemes and locate their appearance time in the speech. In order to locate and predict phonemes, the present invention proposes a reference phoneme sequence, which is derived in advance from the voice text set by the pretender before the attack, and is composed of the phonemes used in the input voice and their timestamp information, which is used as a reference for alignment. The disguise device aligns the latest recorded voice with the reference phoneme sequence set in advance by the pretender, and identifies the correct type of the subsequent synchronized phonemes. The reference phoneme sequence can be aligned with the last recorded phoneme in each synchronized real-time voice. Figure 4As a schematic diagram of a real-time phoneme aligner, the phoneme sequence of a complete speech can be divided into a recorded phoneme sequence and a to-be-played phoneme sequence. In order to align the reference phoneme sequence with the most recently recorded phoneme in the recorded phoneme sequence, a long-term sliding window and a short-term sliding window are introduced. The long-term window determines a large-scale search interval, which starts from the last aligned phoneme and has a preset duration of L, and the interval should contain the phoneme alignment position. Then, a short-term window is used to more accurately locate the latest recorded phoneme. In particular, a short-term window slides in the long-term window of the entire reference phoneme sequence, and the Levenshtein distance is applied to measure its similarity with the window covering the latest recorded phoneme. Only when the distances between all measured phonemes are optimal and do not exceed the preset threshold, the device will take the last phoneme in the corresponding short-term window as the alignment position in the speech. Otherwise, the short-term window in the recording is pulled back by one phoneme and the above process is repeated. Since the speech rate of the pretender is not fixed, the above alignment will be repeated in real time to avoid accumulated errors in the alignment.
[0081] The following is a specific explanation of the real-time phoneme predictor. Although the estimated phoneme sequence can be determined by referring to the benchmark phoneme sequence, the synchronization of sub-phoneme level perturbations still requires the determination of the duration of the phoneme. However, this duration is still unknown in the estimated sequence. Therefore, the real-time phoneme predictor applies a phoneme duration estimation method based on the EWMA algorithm, which dynamically adjusts the duration of the phoneme-level adversarial interference to track real-time streaming speech. The basic idea of phoneme duration estimation is to standardize the speaking speed of different speech to the same scale as the benchmark phoneme sequence. Specifically, first, the speaking speed v relative to the benchmark phoneme sequence is derived for the i-th phoneme. i ,Right now in is the duration of the ith phoneme in the reference phoneme sequence, is the duration of the corresponding phoneme in the recorded speech. The duration of is deduced to a cumulative speaking rate based on the previous k phonemes based on EWMA, that is,
[0082]
[0083] Where β is a preset weight. Considering that human speaking behavior is similar to an LTI system, the speed should remain stable over a period of time. Therefore, according to the estimated speed, the phoneme duration in the reference phoneme sequence can be readjusted to the duration of the phoneme to be played, that is,
[0084]
[0085] where M is the size of the predicted phoneme segment. Afterwards, the duration of the next phoneme before the next calibration can be estimated and the number of corresponding subphone-level perturbations can be determined. In particular, assuming that the duration of each generated subphone-level perturbation is As each subsequent phoneme is predicted as p over time According to equation (8), the masquerading device repeats the corresponding subphone-level perturbation The time is perturbed at the phoneme level and then played back through the loudspeaker.
[0086] In order to verify the technical effect of the present invention, the subphone level perturbation is generated by minimizing the objective function on an AMAX server (Intel Xeon Silver 4210R, 256GB RAM, NVIDIA RTX A6000). By default, the amplitude threshold ∈=0.02 is set, and the duration of the subphone level perturbation is 12.5ms. 24 CIRs were measured using 8 different recording device models and 3 different environments for data enhancement. In addition, 3 mainstream voiceprint model architectures (d-vector, x-vector and Deep Speaker) and 3 training data sets (taken from Voxceleb1 / 2 data sets) were used. A total of 9 different voiceprint models were trained, and during the test, one of them was taken as the target system, and 4 white box models with different structures and training data sets were used for model integration training perturbation. The present invention was deployed on Seeed ReSpeaker Core v2 as a disguised device for the pretender. In order to control the experimental variables, the speaker EDIFIER M230 was used to play the human voice instead of simply speaking by the person. The target voiceprint system is deployed on a Lenovo Xiaoxin Pro 13 with an external receiving front end (i.e., omnidirectional microphone RunPu M10W, lavalier microphone TAKSTAR TCM-340). In each experiment, the M230 voice is played as the real-time voice of the pretender. During the voice playback, the pretender device receives the signal through the lavalier microphone fixed near the M230, and then derives the corresponding perturbation based on the signal. After that, the perturbation is widely projected using two speakers (i.e., JBL CLIP 3 and HP DHS) for injection into the voice. In order to simulate physical domain attacks, the M230 used to play the pretender's voice is 50cm (40cm×30cm) away from the front end of the voiceprint system, while the other speaker (i.e., JBL CLIP 3 or HP DHS) is 5cm away. In order to eliminate the cumulative error, the human voice perturbation is calibrated every 0.5s. In addition, the delay is set to 0.2-0.4s to compensate for the time cost of data processing (about 0.17-0.23s, measured from the implementation) and signal transmission (the value follows the uniform distribution U(0.05,0.1)). The experiments were conducted in three different indoor environments: laboratory (7.4×5.6m 2 , 38.7dBA), office (18.0×6.0m 2 , 43.1dBA) and reading room (3.1×4.4m 2 , 37.2dBA).
[0087] As a physical disguise, the present invention allows the disguiser to authenticate by speaking while playing the corresponding interference in real time. Therefore, its performance depends largely on the accuracy of the synchronization between the interference and the sound. Figure 5 The streaming synchronization performance diagram of the present invention; Figure 5 (a) shows the cumulative distribution function (CDF) of the absolute phoneme delay of the present invention. It can be found that the median of the phoneme delay is 50ms, and more than 75% of the phoneme delays are less than 100ms, which indicates that the alignment delay is acceptable. Considering that synchronization is performed periodically (for example, every 0.5s in the implementation of), the delay of most phonemes is less than 50ms. Despite the presence of prediction errors, the synchronization process also introduces phoneme delays. Considering conventional synchronization, such time overhead may significantly affect the performance of phoneme alignment. Therefore, the time overhead of synchronization is also evaluated, and the results show that the maximum, minimum and average time overheads are 0.13s, 0.06s and 0.11s, respectively, which are less than the synchronization period (i.e., 0.5s). Therefore, the designed synchronization mechanism can effectively support the proposed physical attack. On the other hand, Figure 5 (b) shows the hit rate CDF of the present invention. The hit rates range from 41.0% to 80.9%, with the topic value and average value being 61.2% and 61.3% respectively, indicating that the interference generated by the present invention correctly injects more than 60% of the phonemes on average, effectively predicting the phonemes and the time of their occurrence. In addition, the performance under different synchronization mechanisms is also evaluated. In addition to the mechanism proposed in the present invention, two other mechanisms are implemented for comparison, namely: (1) True reference phoneme sequence: using the accurate phoneme sequence and duration as the reference phoneme sequence, (2) Synchronize once: synchronization is performed only once. Figure 6 is a performance diagram of the streaming real-time camouflage of the present invention, Figure 6 The hit rate and attack success rate (ASR) using the present invention and the other two mechanisms are shown. Compared with synchronization once, the average ASR and hit rate of the present invention are 29.7% and 35.1% higher, respectively. This result shows that the synchronization mechanism of the present invention plays an important role in improving the camouflage performance of the physical domain. On the other hand, it can be found that the hit rate of the present invention is 21.8% lower than the true benchmark phoneme sequence, but the ASR only decreases by 2.4%. This is because even if a phoneme is not accurately estimated, the corresponding disturbance may still be injected near the phoneme due to the regular arrangement. This result further proves that the synchronization performance of the present invention is good.
[0088] The robustness of physical camouflage to various channel interferences is another important feature. The performance of the present invention under different device models and environmental channels is evaluated. In the experiment, two other data augmentation mechanisms are implemented for comparison, namely (1) targeted enhancement: only one device model pair (i.e., CLIP3-M10W) and one environment (i.e., laboratory) are enhanced. (2) No enhancement: no data augmentation technology is used to generate disturbances. Figure 7 is a performance diagram of the present invention under different channel environments, Figure 7(a) shows the ASR under four different device model pairs (i.e., CLIP3 and DHS as speakers, M10W and TCM340 as receivers). It can be seen that the present invention achieves ASR of 85.5%, 91.0%, 86.9% and 97.3% on the four device models, respectively. The ASR of M10W as a receiver is slightly lower because M10W is a conference microphone with higher sensitivity and larger sensing range, and therefore introduces more ambient noise. In addition, the average ASR of the present invention is 44.4% higher than that without enhancement. This result shows that channel enhancement does improve the robustness of adversarial interference in physical camouflage. For targeted enhancement, its ASR is 97.8% on the device model (i.e., CLIP3-M10W), but drops rapidly to 53.9%, 51.6% and 42.5% under other unknown device models, respectively. This shows that introducing enough channel responses to enhance subphone-level perturbations can significantly improve performance. Figure 7 (b) shows the ASR in three different environments (i.e., laboratory, office, and research). It can be observed that the ASR is 90.5%, 91.3%, and 88.5%, respectively, which is higher than the normal values of 56.1%, 53.2%, and 63.7%, respectively. This shows that the present invention is also robust to different environments. For targeted enhancement, the ASR in the three environments is 92.3%, 85.4%, and 81.5%, respectively, with smaller differences. Please note that this may be caused by the relatively stable environment selected in the experiment. More complex environments may introduce more noise. However, as long as the channel enhancement technology is used, the present invention can maintain robustness to different channels.
[0089] The transferability of the present invention in different models was further evaluated. Figure 8 is a performance diagram of the present invention under different models; Figure 8 (a) shows the ASR of sub-phoneme-level adversarial perturbations trained using an ensemble model and a single model under different transfer attack types (i.e., white box, cross-training data, cross-model architecture, and both data and architecture). It can be seen that under the three different transfer attack types, the ASR using the ensemble model is 94.8%, 89.5%, and 85.5%, respectively, which are 27.9%, 29.4%, and 37.2% higher than the ASR of the single attack model, respectively. This shows that compared with a single model, the ensemble model helps to improve the robustness of attacking different black-box models. In addition, it can be observed that cross-structure and dataset attacks perform the worst compared to other transfer attack types. However, even in this case that is closest to the real physical attack, the present invention can achieve more than 80% ASR, thereby verifying its effectiveness. Figure 8(b) shows the ASR under the integration of different numbers of models. It can be seen that under 2, 3 and 4 models, the average ASR is 69.7%, 82.2% and 85.5% respectively, and shows an increasing trend with the increase in the number of models. But the results also show that the increase in ASR gradually decreases with the increase in the number of models. This is because the integrated model can only construct a more generalized solution space, rather than the exact solution space of the target model. But even in such an incomplete space, the present invention achieves an average ASR of 85.5% under 4 ensemble models, proving its robustness under the black box model.
[0090] Finally, the noise level generated by the present invention during operation and the degree of human ear perception caused by it were evaluated through objective experiments. First, the SPL in the surrounding environment was measured when the disturbance was played, where the SPL was measured by a decibel meter SMARTSENSOR AR844 (30-80dB, A weighting). Fig. 9 This is a noise and hearing test result diagram of the present invention, Fig. 9 (a) shows the SPL distribution around the camouflage device at different distances and angles. It can be seen that the SPL in front of the camouflage device (i.e., approximately 0° to 30°) is higher than that at other angles. When the angle is greater than 30°, the sound pressure level at more than 1 meter is lower than 39.2dB, which is only 0.5dB higher than the surrounding environment. This result shows that only when the surrounding person appears at a specific angle to the camouflage device can he / she perceive the interference. In addition, from the perspective of distance, the maximum SPL 45.1dB occurs at a distance of about 0.5m. But when the distance increases to 2 meters, the SPL decays rapidly to 38.9dB. Considering the common social distance of the WHO (i.e., 1 meter), such a small SPL is difficult to be perceived by the surrounding people. In addition, the disguiser can consciously control the direction of his / her device to avoid the attention of the surrounding people. In addition to using SPL, the audibility of adversarial samples in the physical domain is further evaluated using Mel Cepstral Distortion (MCD). Another typical mainstream work FakeBob is introduced here as a benchmark. Then, the MCDs of the present invention and the other two baselines are derived with the same original speech as reference. Fig. 9 (b) shows the present invention and MCDs respectively. It can be seen that the average MCDs of the three methods are 2.45dB, 2.24dB and 4.15dB respectively. Compared with the original noise, the MCD of the present invention increases by 0.21dB, indicating that the distortion caused by the present invention is similar to the distortion of the ambient noise. On the other hand, the average MCD of the present invention is 1.7dB lower than that of FakeBob. These results further prove that the present invention will not be noticed by the surrounding people at a normal distance.
[0091] Finally, it should be noted that the above examples are only specific embodiments of the present invention. Obviously, the present invention is not limited to the above examples, and there are many variations. All variations that can be directly derived or associated with the content disclosed by ordinary technicians in this field should be considered as the protection scope of the present invention.
Claims
1. A physical domain identity disguise system based on voiceprint recognition adversarial samples, characterized in that: The system includes an offline training part and an online camouflage part: The offline training part includes a subphone level perturbation dictionary, a phoneme recognizer, an adversarial sample generator, a voiceprint classifier, a system optimizer and a training corpus; after the signal is input from the training corpus into the phoneme recognizer, an aligned speech carrying phoneme alignment information is output; the adversarial sample generator superimposes the subphone level perturbation in the subphone level perturbation dictionary onto the aligned speech input into the adversarial sample generator according to the phoneme information to generate an adversarial sample; the adversarial sample is forward propagated in the voiceprint classifier, and the output is back-propagated through the system optimizer to optimize the subphone level and perturbation dictionary; The online disguise part relies on a portable real-time disguise device, and the software composition of the disguise device includes a phoneme recognizer, a real-time phoneme aligner and a real-time phoneme predictor; the real-time voice generated by the disguiser is input into the phoneme recognizer to generate a real-time phoneme sequence, and the real-time phoneme aligner locates the phoneme at the current moment according to the real-time phoneme sequence, derives the phoneme sequence to be played, and inputs the real-time phoneme predictor. The real-time phoneme predictor determines the specific duration of each phoneme in the sequence based on the real-time phoneme sequence and the phoneme sequence to be played, and plays sub-phoneme-level perturbations synchronously with the human voice according to the prediction results, thereby synthesizing adversarial samples with channel robustness and model robustness in the physical domain, and finally realizing identity disguise for the voiceprint recognition system; The voiceprint classifier is an integrated classifier, which is used to enhance the cross-model migration capability of subphone-level perturbations. Adversarial samples are simultaneously input into multiple pre-trained voiceprint models with different model architectures and model training sets for forward propagation, and the output scores are summed through a weighted average operation; wherein the model architecture is one or more of d-vector, x-vector and DeepSpeaker, the model training set is a subset of VoxCeleb1 / VoxCeleb2, and the weighted average operation uses an attention mechanism to dynamically adjust the weights of each model output during the iteration process; the camouflage device is an embedded device equipped with a microphone, a speaker and a processing chip hardware device.
2. The physical domain identity disguise system based on voiceprint recognition adversarial samples according to claim 1 is characterized in that: The offline training part also includes a cross-channel enhancer located between the sub-phoneme level perturbation dictionary and the adversarial sample generator, which is used to enhance the channel robustness of the sub-phoneme level perturbation. The cross-channel enhancer uses the Maximum Length Sequence signal to collect channel impulse response, and the collection process also includes different environments, different devices and different distance conditions.
3. A physical domain identity disguise method based on voiceprint recognition adversarial samples, characterized in that: The specific steps include: Offline training part: 1) The subphone level perturbation dictionary provides a matching subphone level perturbation of 10-20ms length for each phoneme used, which is initialized to a random perturbation that conforms to the normal distribution; 2) To enhance the perturbation’s ability to resist channel interference, before being superimposed on the speech, the subphone-level perturbation is augmented by a cross-channel enhancer that simulates the channel states of a variety of different devices and room environments based on previously collected channel impulse responses; 3) In order to correctly superimpose the subphonemic level disturbance on the speech, a phoneme recognizer is used to extract phoneme information from each item in the training set corpus. The phoneme information includes the phoneme type and the start and end timestamps, and is aligned with the original speech. 4) The adversarial sample generator superimposes the data-augmented subphoneme-level perturbation and the aligned speech by phoneme, and the superposition method is to repeatedly fill the subphoneme-level perturbation until the entire phoneme is filled, and outputs the adversarial sample; 5) In order to improve the cross-model migration capability, adversarial samples are input into multiple voiceprint recognition models through an integrated classifier, and the outputs of multiple models are integrated into one in a weighted form; 6) According to the recognition results of the integrated classifier, the system optimizer solves the system optimization problem and iteratively updates the subphone level perturbation dictionary, and finally obtains a trained subphone level perturbation dictionary; Online camouflage section: 7) The imposter records the same speech as the text in advance and uses the phoneme recognizer to extract the phoneme sequence. Includes all the phonemes and durations in speech as standard phoneme sequences for reference during real-time camouflage; 8) The disguise process begins. The disguiser holds the disguise device and speaks the preset password. The disguise device receives the voice in real time through the microphone. 9) The speech signal is recognized as a real-time phoneme sequence by a phoneme recognizer; 10) The real-time phoneme sequence is aligned with a pre-given standard phoneme sequence by a real-time phoneme aligner, thereby obtaining the phoneme sequence to be played that the speaker will say next; 11) The real-time phoneme predictor estimates the duration of the phonemes in the to-be-played phoneme sequence based on the real-time phoneme sequence, the standard phoneme sequence and the to-be-played phoneme sequence whose duration is to be estimated; 12) According to the phoneme sequence and the duration of each phoneme, the camouflage device accurately plays the corresponding sub-phoneme-level adversarial perturbation through the speaker, and finally realizes the online synchronization process with the real-time voice, achieving the purpose of physical domain streaming camouflage attack.
4. The physical domain identity disguise method based on voiceprint recognition adversarial samples according to claim 3 is characterized in that: The phoneme recognizer in step 3) and step 9) is used to extract the phoneme sequence and its time information in the speech. Based on the bidirectional RNN neural network model architecture, its workflow is as follows: the input audio is framed, and each frame is subjected to 26-dimensional MFCC extraction and then input into the bidirectional RNN network. The input layer has a size of 26×1, and the hidden layer uses a GRU unit with a size of 256. The output layer has a size of 40, corresponding to the phonemes used. The output result is processed by merging repeated phonemes and correcting erroneous phonemes to output a phoneme sequence.
5. The physical domain identity disguise method based on voiceprint recognition adversarial samples according to claim 3 or 4 is characterized in that: The implementation process of the real-time phoneme aligner is as follows: first, a reference phoneme sequence is constructed from the phoneme information extracted from the pre-recorded text speech that is the same as the real-time speech, a phoneme recognizer is used to continuously perform phoneme recognition on the recorded real-time speech to obtain a real-time phoneme sequence, and a sliding window mechanism is used to align the recognized phoneme sequence with the reference phoneme sequence to confirm the phoneme sequence in the next speaker's speech.
6. The physical domain identity disguise method based on voiceprint recognition adversarial samples according to claim 5 is characterized in that: The sliding window mechanism utilizes long-term and short-term sliding windows, wherein the long-term sliding window is used to determine the search range of the short-term sliding window, and the short-term sliding windows exist in pairs in the reference phoneme sequence and the real-time phoneme sequence, and are used to compare the matching degree of the window contents.
7. The physical domain identity disguise method based on voiceprint recognition adversarial samples according to claim 3 or 6 is characterized in that: The implementation process of the real-time phoneme predictor is as follows: based on the relative relationship between the real-time phoneme sequence and the reference phoneme sequence, the exponentially weighted moving average algorithm is used to estimate the current speaking rate, and then the phoneme duration in the reference phoneme sequence is inversely calculated with the current speaking rate to obtain the duration of the phonemes in the real-time speech.
Citation Information
Patent Citations
Phoneme-level voiceprint recognition confrontation sample construction system and method based on neural network generative model
CN114093371A