A voice backdoor verification method and device based on room impulse response

By constructing a dynamic trigger based on room impulse response and using a conditional variational autoencoder generative adversarial network generator to synthesize the dynamic trigger for poisoning training of the speech model, the problem of insufficient concealment and robustness of speech backdoor testing in the digital space in existing technologies is solved, and the verification of effective backdoor activation in real environment is realized.

CN116597811BActive Publication Date: 2025-11-28ZHEJIANG UNIV +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202310533603.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-12
Publication Date
2025-11-28
Estimated Expiration
2043-05-12

AI Technical Summary

Technical Problem

Existing voice backdoor attack tests are conducted in digital space, which fails to effectively simulate the real physical environment, resulting in insufficient concealment and robustness, and making it impossible to effectively assess the security of voice systems.

Method used

By acquiring the physical spatial attribute information of the target speech model, a conditional vector of the room impulse response is constructed. A dynamic trigger is generated using a conditional variational autoencoder generative adversarial network. The room impulse response signal is synthesized as a dynamic trigger. Clean speech samples are poisoned for training. The poisoned speech samples are used to enhance the target speech model. The backdoor is activated to verify its vulnerability.

Benefits of technology

It improves the concealment and robustness of voice backdoors in physical space, enables effective activation of backdoors in real-world environments, reduces testing costs, and enhances the reliability of backdoor attack testing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116597811B_ABST
    Figure CN116597811B_ABST
Patent Text Reader

Abstract

The application discloses a speech backdoor verification method based on a room impulse response, comprising: obtaining clean speech samples of a target speech model and attribute information of a physical space where the target speech model is located; setting acoustic parameters according to the attribute information, and constructing a conditional vector of the room impulse response according to the acoustic parameters; inputting the conditional vector and a randomly sampled hidden vector after splicing into a room impulse response generator to synthesize a room impulse response signal as a dynamic trigger; using the dynamic trigger to poison the clean speech samples as poisoned speech samples, and using the poisoned speech samples and the clean speech samples to train the target speech model, so that the target speech model is infected and injected with a backdoor; after deploying the infected target speech model, normally speaking to emit speech to trigger the backdoor, thereby verifying the backdoor vulnerability of the target speech model, the method effectively improves the concealment and robustness of the speech backdoor, and thus provides a real and reliable backdoor attack test.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of computer artificial intelligence security, and particularly relates to a voice backdoor verification method and device based on room impulse response. BACKGROUND

[0002] In recent years, deep learning technology has endowed modern voice systems (such as voice content recognition and identity recognition) with strong cognitive ability. However, recent research has revealed that the high dependence of modern voice systems on training data can lead to serious backdoor security threats. Such backdoor attacks rely on carefully designed triggers and can establish a strong false association with the target output by polluting a small amount of training data, thereby misleading the system to output incorrect results. This attack does not affect the normal use of the system, and therefore has strong concealment, seriously threatening the security of artificial intelligence systems.

[0003] In order to evaluate and improve the security and robustness of voice systems, pre-emptive backdoor vulnerability testing is needed. However, existing backdoor attack testing research has only verified the feasibility of voice backdoor learning in the digital space, that is, by designing different triggers (such as single-frequency sound, ultrasonic wave, background sound, and adversarial samples) in the time domain and frequency domain to implant backdoors into voice command recognition or voice identity recognition systems. These studies assume that the trigger is directly injected into the target voice system after network transmission without any distortion, which greatly deviates from the real attack testing environment. In addition, the existing digital domain triggers have poor concealment and cannot resist channel interference, resulting in insufficient backdoor attack testing capability.

[0004] Therefore, how to improve the concealment and robustness of voice backdoors in the physical space to provide a real and reliable attack testing scheme is a problem to be solved at present.

[0005] Patent application No. CN114462031A discloses a backdoor attack method, which includes obtaining original sample data of a target model; selecting target sample data with a target class label as the target class from the original sample data as poisoning data, and setting a trigger in the original text in the target sample data to obtain the poisoning data; training the target model using the original sample data and the poisoning data as training data to obtain a test model; wherein the test model is used for attack testing on test data with a set trigger.

[0006] The patent document with the publication number CN113946687A discloses a label consistent text backdoor attack method, which comprises: generating trigger words through a semantic primitive library for a target data set; using an adversarial perturbation method and a hidden keyword method based on a black box condition to perturb original input samples; generating poisoned samples by adding the generated trigger words to the perturbed sentences through a trigger word replacement method based on semantic primitives, and training a target model with the poisoned data set; in the reasoning stage, the trigger words are added to the test sentences by the semantic replacement method, so as to induce the target model to predict the target category. The present application sets

[0007] When the technical solutions disclosed in the above two patent documents are applied to a speech recognition system, since the trigger has no distortion, it greatly deviates from the real attack test environment, resulting in insufficient backdoor attack test capability. SUMMARY

[0008] In view of the above, the purpose of the present application is to provide a speech backdoor verification method and device based on room impulse response, which realizes the implantation and activation of speech backdoors in physical space, effectively improves the concealment and robustness of speech backdoors, and thus provides real and reliable backdoor attack tests.

[0009] To achieve the above-mentioned purpose of the application, the speech backdoor verification method based on room impulse response provided by the embodiment comprises the following steps:

[0010] Obtain clean speech samples of a target speech model, and obtain attribute information of a physical space where the target speech model is located;

[0011] Set acoustic parameters according to the attribute information;

[0012] Construct a conditional vector of room impulse response according to the acoustic parameters;

[0013] Concatenate the conditional vector with a randomly sampled hidden vector and input it into a pre-trained room impulse response generator to synthesize a batch of room impulse response signals as dynamic triggers;

[0014] Use the dynamic triggers to poison the clean speech samples as poisoned speech samples, and use the poisoned speech samples and the clean speech samples to enhance the training of the target speech model, so that the target speech model is infected and injected with a backdoor;

[0015] After deploying the infected target speech model, speak normally in the physical space where the target speech model is located to trigger the backdoor, thereby verifying the backdoor vulnerability of the target speech model.

[0016] Preferably, the attribute information of the physical space includes room size and wall material.

[0017] Preferably, the setting the acoustic parameters according to the attribute information comprises:

[0018] The acoustic parameters include a reverberation time, i.e. a time required for a sound field in the room to decay to a certain degree from an initial value, when a 60dB sound pressure level is selected, the reverberation time RT 60 The specific calculation is:

[0019]

[0020] Wherein, V=L×W×H is the room volume, S=2(L×W+L×H+W×H) is the room surface area, c is the sound speed, and a represents the average sound absorption coefficient of the room surface, which is calculated as:

[0021]

[0022] Wherein, S i is the area of the i-th face of the room, I is the total number of faces of the room, and a i is the sound absorption coefficient corresponding to the i-th face, which is determined by the wall material.

[0023] Preferably, the constructing the condition vector of the room impulse response according to the acoustic parameters comprises:

[0024] The condition vector c is constructed according to the room size (L, W, H), the target voice model position (x t , y t , z t ), the speaker position (x a , y a , z a ), and the room reverberation time RT 60 as the acoustic parameter, and is expressed as:

[0025] c=[L,W,H,x a ,y a ,z a ,x t ,y t ,z t ,RT 60 ].

[0026] Preferably, the room impulse response generator is trained based on a conditional variational autoencoder-generative adversarial network architecture, and specifically comprises:

[0027] The conditional variational autoencoder - generative adversarial network architecture comprises an encoder, a generator and a discriminator, wherein the encoder maps a real room impulse response signal to a latent vector by learning a conditional probability distribution, the latent vector is spliced with a condition vector and input to the generator, the generator reconstructs a synthesized room impulse response signal by learning another conditional probability distribution, and the discriminator discriminates between the input real room impulse response signal and the synthesized room impulse response signal.

[0028] The conditional variational autoencoder - generative adversarial network architecture is adversarially trained, and the loss function used in adversarial training comprises a reconstruction loss, a KLD loss, an adversarial loss and a gradient penalty loss, and after training, the generator with optimized parameters is used as a pre-trained room impulse response generator.

[0029] Preferably, the loss function is represented as:

[0030]

[0031]

[0032] The reconstruction loss is represented as:

[0033] The KLD loss is represented as:

[0034] The part of the adversarial loss is represented as:

[0035] The other part of the adversarial loss is represented as:

[0036] The gradient penalty loss is represented as:

[0037] wherein h represents a real room impulse response signal, h' represents a reconstructed synthesized room impulse response signal, represents the square of the L2 norm, z represents a latent vector, p(z|h) represents a conditional probability distribution learned by the encoder, and N(z|0,I) represents a standard Gaussian distribution with mean 0 and variance I, represents the KL divergence, represents the expectation, and D(·) represents the discrimination result output by the discriminator, represents the derivative of . ​

[0038] Preferably, the use of dynamic triggers to poison clean speech samples as poisoned speech samples comprises:

[0039] The dynamic trigger is normalized, the normalized dynamic trigger is flipped along the time axis, and the normalized dynamic trigger is time-domain weighted summed with the clean speech to obtain a poisoned speech sample, and the category label of the poisoned speech sample is modified to the target result label.

[0040] Preferably, the speaking normally in the physical space where the target speech model is located to trigger the backdoor to verify the backdoor vulnerability of the target speech model comprises:

[0041] The speech emitted when speaking normally in the physical space where the target speech model is located is automatically brought into the dynamic trigger to form a test sample during the propagation process of the speech, and the test sample activates the hidden backdoor when received by the target speech model, and finally verifies whether there is a backdoor vulnerability according to the output result of the target speech model.

[0042] To achieve the above-mentioned purposes, the embodiment further provides a speech backdoor verification device based on room impulse response, comprising an acquisition module, an acoustic parameter setting module, a condition vector generation module, a dynamic trigger generation module, a data enhancement type poisoning module, and an injection-free backdoor activation module.

[0043] The acquisition module is used to acquire clean speech samples of a target speech model, and obtain attribute information of a physical space where the target speech model is located;

[0044] The parameter setting module is used to set acoustic parameters according to the attribute information;

[0045] The condition vector generation module is used to construct a condition vector of the room impulse response according to the acoustic parameters;

[0046] The dynamic trigger generation module is used to splice the condition vector and a randomly sampled hidden vector, and input the spliced vector into a pre-trained room impulse response generator to synthesize a batch of room impulse response signals as dynamic triggers;

[0047] The data enhancement type poisoning module is used to use the dynamic triggers to poison the clean speech samples as poisoned speech samples, and use the poisoned speech samples and the clean speech samples to train the target speech model, so that the target speech model is infected and injected with a backdoor;

[0048] The injection-free backdoor activation module is used to deploy the infected target speech model, speak normally in the physical space where the target speech model is located to emit speech to trigger the backdoor, and verify the backdoor vulnerability of the target speech model.

[0049] The speech backdoor verification method based on room impulse response can effectively improve the concealment and robustness of the speech backdoor in the physical space, thereby performing preliminary backdoor vulnerability evaluation and verification on the target speech model of the speech system, and has at least the beneficial effects of:

[0050] 1) Robust physical backdoor: The sound channel itself is used as an injection path, that is, the room impulse response of the speaking voice is used as a natural dynamic trigger, which reduces the influence of physical channel interference and effectively improves the robustness of the dynamic trigger in the physical space.

[0051] 2) Concealed trigger design: The poisoned speech sample and the clean speech sample are transmitted through the same digital path and physical channel and received by the target speech model of the target speech system, and the dynamic trigger based on the room impulse response behaves as natural reverberation, thus having strong concealment.

[0052] 3) Injection-free backdoor activation: The backdoor can be activated by directly issuing a voice in the physical space without any transceiving equipment or operation, and this injection-free activation method can reduce the cost of backdoor triggering and improve the efficiency of backdoor vulnerability verification testing. BRIEF DESCRIPTION OF DRAWINGS

[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0054] Figure 1 is a flowchart of the speech backdoor verification method based on room impulse response provided by the embodiment;

[0055] Figure 2 is a flowchart of the speech backdoor verification method based on room impulse response provided by the embodiment;

[0056] Figure 3 is a schematic diagram of the acoustic parameter setting provided by the embodiment;

[0057] Figure 4 is an architecture diagram of the conditional variational autoencoder-generative adversarial network provided by the embodiment;

[0058] Figure 5 is a structural schematic diagram of the speech backdoor verification device based on room impulse response provided by the embodiment. DETAILED DESCRIPTION

[0059] In order to make the objects, technical solutions and advantages of the present application clearer, further detailed description will be given to the present application in combination with the drawings and examples. It should be understood that the specific implementation described herein is only used to explain the present application, and does not limit the protection scope of the present application.

[0060] Figure 1 is a flowchart of the voice backdoor verification method based on room impulse response provided by the embodiment; Figure 2 is a flowchart of the voice backdoor verification method based on room impulse response provided by the embodiment. As shown in Figure 1 and Figure 2 , the voice backdoor verification method based on room impulse response provided by the embodiment comprises the following steps:

[0061] S110, obtaining clean speech samples of a target speech model, and obtaining attribute information of a physical space where the target speech model is located.

[0062] In the embodiment, the target speech model is deployed in a target speech system, and is mainly used for speech recognition according to input data and outputting a recognition result. The attribute information of the physical space where the target speech model is located includes room size (L, W, H) and wall material, wherein L, W and H represent the length, width and height of the room. The wall material is the wall material information of the four walls, floor and ceiling.

[0063] In the embodiment, the target speech model position (x t ,y t ,z t ), the speaker position (x a ,y a ,z a ) and the expected target speech model position are also obtained.

[0064] S120, setting acoustic parameters according to the attribute information.

[0065] In the embodiment, in order to describe the acoustic reverberation characteristics of the physical space, the room 60 dB reverberation time RT 60 is estimated according to the attribute information by Eyring formula, and is calculated as:

[0066]

[0067] Wherein, V=L×W×H is the room volume, S=2(L×W+L×H+W×H) is the room surface area, c is the sound speed, which can be specifically c=343m / s for the sound speed at 20℃, and a represents the average sound absorption coefficient of the room surface, which is calculated as:

[0068]

[0069] Among them, S i Let be the area of ​​the i-th face of the room, I be the total area of ​​the room, and α be the area of ​​the ith face. i Let be the sound absorption coefficient corresponding to the i-th surface. This sound absorption coefficient is determined by the wall material and is obtained by querying a publicly available building material acoustics database.

[0070] S130, construct the conditional vector of the room impulse response based on acoustic parameters.

[0071] In the embodiment, when constructing the conditional vector of the room impulse response, the room dimensions (L, W, H) and the target speech model location (x) are considered. t ,y t ,z t Speaker position (x) a ,y a ,z a ), and the room reverberation time RT as an acoustic parameter 60 Construct a condition vector c, such as Figure 3 As shown, it is represented as:

[0072] c = [L, W, H, x] a ,y a ,z a ,x t ,y t ,z t ,RT 60 ].

[0073] S140 concatenates the conditional vector with the randomly sampled latent vector and inputs it into a pre-trained room impulse response generator to synthesize a batch of room impulse response signals as dynamic triggers.

[0074] In this embodiment, to generate a room impulse response signal that conforms to the conditional vector description, a room impulse response generator is pre-trained to learn the overall data distribution and local acoustic structure of real room impulse response signals. Specifically, the room impulse response generator is trained based on a conditional variational autoencoder-generative adversarial network architecture.

[0075] like Figure 4 As shown, the Conditional Variational Autoencoder-Generative Adversarial Network (CVAE-GAN) architecture includes an encoder E, a generator G, and a discriminator D. The encoder E learns the latent vector z of the real room impulse response (RIR) h by learning the conditional probability distribution p(z|h). Specifically, according to the variational inference principle, the encoder E learns the mean vector μ and variance vector σ of the latent vector z through a deep neural network, and then uses the reparameter technique to construct the latent vector z: z=μ+σ⊙ε, where ε~N(0,I).

[0076] The generated latent vector z is concatenated with the condition vector c and input to the generator G, which reconstructs the synthesized room impulse response signal h' by learning another conditional probability distribution q(h'|z, c), and then inputs the real room impulse response signal h and the reconstructed synthesized room impulse response signal (RIR) h' to the discriminator D, which discriminates between the input real room impulse response signal h and the synthesized room impulse response signal h'. Through adversarial training, the generator G strives to generate realistic room impulse response signals to deceive the discriminator D, while the discriminator D strives to distinguish the synthesized room impulse response signal from the generator G from the real room impulse response signal.

[0077] In the embodiment, the encoder E and the discriminator D use 5-layer one-dimensional convolution for dimension reduction to obtain a high-dimensional representation, the encoder E uses two parallel fully connected layers to learn the mean vector and the variance vector from the high-dimensional representation, and the discriminator D applies a global pooling layer and a fully connected layer to learn the true and false probability. The generator G uses a network architecture symmetric to the encoder E, but replaces the one-dimensional convolution layer with a one-dimensional deconvolution layer.

[0078] During adversarial training, the loss function adopted includes reconstruction loss, KLD loss, adversarial loss, and gradient penalty loss, which are specifically represented as:

[0079]

[0080]

[0081] The reconstruction loss is represented as Lr, and is calculated as:

[0082] The KLD loss is represented as LKLD, and is calculated as:

[0083] The part of the adversarial loss is represented as La, and is calculated as:

[0084] The other part of the adversarial loss is represented as Lb, and is calculated as:

[0085] The gradient penalty loss is represented as Lg, and is calculated as:

[0086] where h represents the real room impulse response signal, h' represents the reconstructed synthesized room impulse response signal, represents the square of the L2 norm, z represents the latent vector, p(z|h) represents the conditional probability distribution learned by the encoder, N(z|0, I) represents a standard Gaussian distribution with mean 0 and variance I, Denotes KL divergence, This represents the expectation, and D(·) represents the discrimination result output by the discriminator. Indicates to In Find the derivative.

[0087] After pre-training using the aforementioned loss function, the latent vectors gradually approach a standard Gaussian distribution, and the generator G can generate realistic room impulse response signals that conform to the description of the given conditional vectors. After training, the parameter-optimized generator serves as the pre-trained room impulse response generator.

[0088] Based on the pre-trained room impulse response generator, given the conditional vector of the target physical space, a batch of latent vectors are randomly sampled and concatenated. Then, a batch of room impulse response signals are synthesized through the room impulse response generator G. Due to the randomness of the latent vectors, although the synthesized room impulse response signals have similar reverberation characteristics, there are still some differences. Therefore, they can be used as dynamic triggers.

[0089] S150 uses a dynamic trigger to poison clean speech samples as poisoned speech samples. It then uses the poisoned speech samples and clean speech samples to enhance the training of the target speech model, thereby infecting the target speech model and injecting a backdoor.

[0090] In this embodiment, during the data augmentation stage of constructing the target speech system, a dynamic trigger is used to poison clean speech samples as poisoned speech samples. The poisoned speech samples and clean speech samples are then used to enhance and train the target speech model, thereby infecting the target speech model and injecting a backdoor. Specifically, the poisoning process is as follows:

[0091] First, the dynamic trigger is normalized: the main impulse response of the dynamic trigger h′ is extracted, with a length of 1 second, and the signal power is normalized.

[0092] Then, speech convolution: flip the normalized dynamic trigger h′ along the time axis and perform a time-domain weighted sum with the clean speech x: T(x)=x*h′ as the poisoned speech sample;

[0093] Label modification: Modify the category label of the poisoned speech sample to the target result label.

[0094] After the poisoning is completed, the target speech model is trained on the poisoned speech sample and other clean speech samples, thereby learning the strong correlation between the dynamic trigger and the preset target result label to form a backdoor.

[0095] S160, after the infected target speech model is deployed, normal speech in the physical space where the target speech model is located is emitted to trigger the backdoor, thereby verifying the backdoor vulnerability of the target speech model.

[0096] In the embodiment, after the infected target speech model is deployed, the speech emitted when speaking normally in the physical space where the target speech model is located is automatically brought into a dynamic trigger to form a test sample by the room reverberation effect in the propagation process of the speech, the test sample activates the hidden backdoor after being received by the target speech model, and finally whether the backdoor vulnerability exists is verified according to the output result of the target speech model: if the output result is the same as the preset target result label, it is proved that the backdoor exists in the target speech system and is successfully activated.

[0097] The speech backdoor verification method based on room impulse response provided in the above embodiment takes the sound channel itself as an injection path, effectively improves the robustness of the dynamic trigger in the physical space, and the trigger based on the room impulse response is expressed as natural reverberation, which is difficult to be distinguished and detected by the human ear. In addition, the backdoor can be activated by normal speech without any transceiving equipment or operation, which greatly reduces the cost of physical backdoor attack test.

[0098] Based on the same inventive concept, the embodiment further provides a speech backdoor verification device 500 based on room impulse response, as shown in Figure 5 which comprises an acquisition module 510, an acoustic parameter setting module 520, a condition vector generation module 530, a dynamic trigger generation module 540, a data enhancement type poisoning module 550, and an injection-free backdoor activation module 560.

[0099] The acquisition module 510 is configured to acquire clean speech samples of a target speech model, and obtain attribute information of a physical space where the target speech model is located; the acoustic parameter setting module 520 sets acoustic parameters according to the attribute information; the condition vector generation module 530 constructs a condition vector of the room impulse response according to the acoustic parameters; the dynamic trigger generation module 540 inputs the condition vector and a randomly sampled hidden vector into a pre-trained room impulse response generator after splicing, synthesizes a batch of room impulse response signals as dynamic triggers; the data enhancement type poisoning module 550 poisons the clean speech samples with the dynamic triggers as poisoned speech samples, and trains the target speech model with the poisoned speech samples and the clean speech samples, so that the target speech model is infected and injected with a backdoor; and the injection-free backdoor activation module 560 emits speech in the physical space where the target speech model is located after the infected target speech model is deployed, to trigger the backdoor, thereby verifying the backdoor vulnerability of the target speech model.

[0100] It should be noted that the voice backdoor verification device based on room impulse response provided in the above embodiment should be divided into the above functional modules when performing voice backdoor verification based on room impulse response. The above functions can be completed by different functional modules according to needs. In addition, the voice backdoor verification device based on room impulse response provided in the above embodiment and the voice backdoor verification method embodiment belong to the same concept, and the specific implementation process is detailed in the voice backdoor verification method embodiment. Here, it will not be repeated.

[0101] The voice backdoor verification device based on room impulse response provided in the embodiment can successfully implant backdoors in different voice content recognition systems and identity recognition systems and activate the test in the physical space, has strong physical robustness and concealment, and can support the pre-verification of voice system backdoor vulnerability.

[0102] The specific embodiments described above have detailed the technical solutions and beneficial effects of the present application. It should be understood that the above description is only the most preferred embodiment of the present application and is not intended to limit the present application. Any modifications, supplements, and equivalent replacements made within the principle range of the present application should be included in the protection scope of the present application.

Claims

1. A voice backdoor verification method based on room impulse response, characterized in that, The method comprises the following steps: obtaining clean speech samples of a target speech model and attribute information of a physical space where the target speech model is located; setting acoustic parameters according to the attribute information; constructing a conditional vector of a room impulse response according to the acoustic parameters; concatenating the conditional vector and a randomly sampled latent vector and inputting the result into a pre-trained room impulse response generator to synthesize a batch of room impulse response signals as dynamic triggers; poisoning the clean speech samples using the dynamic triggers to obtain poisoned speech samples, including: normalizing the dynamic triggers, flipping the normalized dynamic triggers along the time axis, and performing time-domain weighted summation with the clean speech to obtain the poisoned speech samples, and modifying the class labels of the poisoned speech samples to target result labels; performing enhancement training on the target speech model using the poisoned speech samples and the clean speech samples to infect and inject backdoors into the target speech model; deploying the infected target speech model and speaking normally in the physical space where the target speech model is located to trigger the backdoors, thereby verifying the backdoor vulnerability of the target speech model.

2. The room impulse response based voice backdoor verification method of claim 1, wherein, The attribute information of the physical space includes room size and wall material.

3. The room impulse response based voice backdoor verification method of claim 2, wherein, The setting of the acoustic parameters according to the attribute information comprises: The acoustic parameter comprises the reverberation time, i.e. the time required for the sound field in a room to decay to a certain degree from an initial value, when a sound pressure level of 60 dB is chosen, the reverberation time RT 60 The specific calculation is: wherein V = L x W x H is the room volume, S = 2(L x W + L x H + W x H) is the room surface area, c is the sound speed, and a represents the average sound absorption coefficient of the room surface, which is calculated as: Wherein, S i is the area of the i-th face of the room, I is the total number of faces of the room, a i is the sound absorption coefficient corresponding to the i-th face, which is determined by the wall material.

4. The room impulse response based voice backdoor verification method of claim 1, wherein, The construction of the conditional vector of the room impulse response according to the acoustic parameters comprises: A condition vector c is constructed from the room dimensions (L, W, H), the target speech model position (x t ,y t ,z t ), the speaker position (x a ,y a ,z a ), and the room reverberation time RT 60 as an acoustic parameter, and is expressed as: c = [L, W, H, x a , y a , z a , x t , y t , z t , RT 60 ].

5. The room impulse response based voice backdoor verification method of claim 1, wherein, The room impulse response generator is trained based on a conditional variational autoencoder generative adversarial network architecture, which specifically comprises: The conditional variational autoencoder generative adversarial network architecture comprises an encoder, a generator, and a discriminator, wherein the encoder maps a real room impulse response signal to a latent vector by learning a conditional probability distribution, the latent vector is concatenated with a conditional vector and input into the generator, the generator reconstructs a synthesized room impulse response signal by learning another conditional probability distribution, and the discriminator distinguishes between the input real room impulse response signal and the synthesized room impulse response signal; The conditional variational autoencoder generative adversarial network architecture is subjected to adversarial training, and the loss function used during the adversarial training comprises a reconstruction loss, a KLD loss, an adversarial loss, and a gradient penalty loss, and after the training is completed, the generator with optimized parameters is used as the pre-trained room impulse response generator.

6. The room impulse response based voice backdoor verification method of claim 5, wherein, The loss function is represented as: denotes the reconstruction loss, computed as: represents the KLD loss, calculated as: represents a portion of the adversarial loss, computed as: represents another part of the adversarial loss, computed as: denotes the gradient penalty loss, computed as: where h denotes the real room impulse response signal, h ′ denotes the reconstructed synthetic room impulse response signal, denotes the square of the L2 norm, z denotes the latent vector, p(z|h) denotes the conditional probability distribution learned by the encoder, N(z|0, I) denotes a standard Gaussian distribution with mean 0 and variance I, denotes the KL divergence, denotes the expectation, D(·) denotes the discriminative result of the discriminator output, denotes the derivative with respect to in the equation .

7. The room impulse response based voice backdoor verification method of claim 1, wherein, The normal speaking in the physical space where the target speech model is located to trigger the backdoors, thereby verifying the backdoor vulnerability of the target speech model, comprises: During the propagation of the speech normally spoken in the physical space where the target speech model is located, the dynamic triggers are automatically added to form test samples, the test samples are received by the target speech model to activate the hidden backdoors, and finally the existence of the backdoor vulnerability is verified according to the output result of the target speech model.

8. A voice backdoor verification apparatus based on room impulse response, characterized by, The method comprises an acquisition module, an acoustic parameter setting module, a conditional vector generation module, a dynamic trigger generation module, a data enhancement poisoning module, and an injection-free backdoor activation module. The acquisition module is configured to acquire clean speech samples of a target speech model and obtain attribute information of a physical space where the target speech model is located; The parameter setting module is configured to set an acoustic parameter according to the attribute information; The condition vector generation module is configured to construct a condition vector of a room impulse response according to the acoustic parameter; The dynamic trigger generation module is configured to splice the condition vector and a randomly sampled hidden vector to input a pre-trained room impulse response generator to synthesize a batch of room impulse response signals as dynamic triggers; The data enhancement type poisoning module is configured to use the dynamic triggers to poison the clean speech samples as poisoned speech samples, including: performing normalization processing on the dynamic triggers, flipping the normalized dynamic triggers along a time axis, and performing time domain weighted summation on the normalized dynamic triggers and the clean speech to obtain the poisoned speech samples, and modifying a category label of the poisoned speech samples to a target result label; The target speech model is trained using the poisoned speech samples and the clean speech samples, so that the target speech model is infected and injected with a backdoor; The injection-free backdoor activation module is configured to deploy the infected target speech model, and then make a speech in the physical space where the target speech model is located to trigger the backdoor, so as to verify the backdoor vulnerability of the target speech model.

Citation Information

Patent Citations

  • Label-consistent text backdoor attack method

    CN113946687A

  • Backdoor attack method, related device and storage medium

    CN114462031A

  • Voice signal generation method and device

    CN108182936A

  • Convolutional confrontation sample construction method and device for voice identity anonymity

    CN115631757A