A method, device and electronic device for generating speech adversarial samples
By encoding and migrating perturbation processing of speech signals in the latent space, voice adversarial samples are generated, which solves the problem of audio quality degradation in the prior art speech adversarial sample generation, and achieves high-quality speech adversarial sample generation.
Patent Information
- Application Number
- CN202510398578.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-04-01
AI Technical Summary
The existing voice adversarial sample generation method causes audio quality to decline during the process of adding perturbation to the voice, affecting normal voice communication and communication.
The current audio signal is encoded in the latent space, select a perturbation from the constructed migratory set to add it to the latent feature encoding, and decode the perturbed encoding to generate a speech adversarial sample.
By encoding and perturbing in the latent space, the introduction of noise into the original audio data space is avoided, the problem of audio quality is solved, and the quality of voice confrontation samples is improved.
Smart Images

Figure CN119920239B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of voice security technology. Specifically, it relates to a method, device, and electronic device for generating voice adversarial samples. Background Art
[0002] With the rapid development of artificial intelligence technology, an Automatic Speech Recognition (ASR) system can efficiently convert speech into text using deep learning technology. This technology has been widely applied in industries such as financial services, healthcare, and customer service, improving work efficiency and accuracy. However, this widespread application has also brought new challenges, especially in terms of privacy protection. Therefore, how to effectively protect personal privacy while ensuring the quality of voice communication has become an urgent problem to be solved.
[0003] However, existing voice processing methods usually add general adversarial perturbations to the audio waveform to protect privacy, making the automatic speech recognition system unable to accurately recognize the speech content. However, this method often leads to a serious decline in audio quality and introduces a large amount of noise, thus affecting normal voice communication. Summary of the Invention
[0004] In view of this, the purpose of this application is to provide a method, device, and electronic device for generating voice adversarial samples to solve the problem that the existing voice adversarial sample generation method leads to a decline in audio quality.
[0005] In a first aspect, an embodiment of this application provides a method for generating voice adversarial samples, including:
[0006] Receiving an input current audio signal, encoding the current audio signal in the latent space to obtain a current latent feature encoding;
[0007] Selecting a perturbation from a constructed transferable perturbation set and adding it to the current latent feature encoding to obtain a perturbed current latent feature encoding, where the perturbations in the transferable perturbation set are trained in the latent space;
[0008] Decoding the perturbed current latent feature encoding to obtain a current voice adversarial sample.
[0009] In an optional implementation, encoding the current audio signal in the latent space to obtain a current latent feature encoding includes: inputting the current audio signal into a variational autoencoder for latent space encoding to obtain a current latent feature encoding.
[0010] In an alternative embodiment, before receiving the current audio signal as input, it further includes: for each initial perturbation, iteratively optimizing the initial perturbation with the optimization goal of minimizing the audio loss to obtain an optimized transferable perturbation, where the audio loss includes a first audio loss representing the loss between the first latent feature encoding corresponding to the target audio and the perturbation; constructing a transferable perturbation set based on multiple optimized transferable perturbations.
[0011] In an alternative embodiment, iteratively optimizing the initial perturbation with the optimization goal of minimizing the audio loss to obtain an optimized transferable perturbation includes: for the current round of optimization process, determining the first audio loss of the current round based on the first latent feature encoding and the perturbation of the previous round; determining the second audio loss of the current round based on the original audio adversarial sample of the current round, the target audio, and the perturbation of the previous round; determining the perturbation of the current round based on the first audio loss of the current round and the perturbation of the previous round, so as to use the perturbation of the current round for subsequent optimization processes until an optimized transferable perturbation is obtained.
[0012] In an alternative embodiment, the original audio adversarial sample of the current round is determined through the following processing: selecting the original audio signal of the current round from the set of original audio signals, encoding the original audio signal of the current round in the latent space to obtain the second latent feature encoding of the current round; adding the perturbation of the previous round to the second latent feature encoding of the current round in the latent space to generate the perturbed second latent feature encoding of the current round; after decoding the perturbed second latent feature encoding of the current round, obtaining the original audio adversarial sample of the current round.
[0013] In an alternative embodiment, the method further includes: for the current round of optimization process, determining a first parameter value corresponding to the noise parameter and a second parameter value corresponding to the acoustic environment parameter; adding the first parameter value to the perturbed second latent feature encoding of the current round and decoding the perturbed second latent feature encoding of the current round to obtain the original audio adversarial sample of the current round; determining the second audio loss of the current round based on the original audio adversarial sample of the current round, the second parameter value corresponding to the acoustic environment parameter, and the target audio.
[0014] In an alternative embodiment, determining the second audio loss of the current round based on the original audio adversarial sample of the current round, the second parameter value corresponding to the acoustic environment parameter, and the target audio includes: determining the convolution result of the second parameter value and the original audio adversarial sample of the current round added; substituting the convolution result and the target audio into a preset loss function to determine the second audio loss of the current round.
[0015] In an alternative embodiment, decoding the perturbed current latent feature encoding to obtain the current speech adversarial sample includes: using a decoder to reconstruct the perturbed current latent feature encoding to obtain the current speech adversarial sample.
[0016] Second aspect, the embodiments of the present application further provide a voice adversarial sample generation device, and the device includes:
[0017] An encoding module, configured to receive the input current audio signal, encode the current audio signal in the latent space, and obtain the current latent feature encoding;
[0018] A perturbation adding module, configured to select a perturbation from the constructed transferable perturbation set and add it to the current latent feature encoding to obtain the perturbed current latent feature encoding, where the perturbations in the transferable perturbation set are trained in the latent space;
[0019] A decoding module, configured to decode the perturbed current latent feature encoding to obtain the current voice adversarial sample.
[0020] Third aspect, the embodiments of the present application further provide an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps of the voice adversarial sample generation method as described above are executed.
[0021] The embodiments of the present application bring the following beneficial effects:
[0022] A voice adversarial sample generation method, device, and electronic device provided by the embodiments of the present application can encode the current audio signal in the latent space, and apply perturbations to the perturbed current latent feature encoding to transfer the perturbations to the latent space, avoiding the problem of the decline in audio quality caused by introducing noise in the original audio data space. Compared with the voice adversarial sample generation methods in the prior art, the problem of the decline in audio quality during the process of adding perturbations to the voice is solved.
[0023] To make the above objects, features, and advantages of the present application more obvious and understandable, the following specific embodiments are given, and in conjunction with the accompanying drawings, the detailed description is as follows. Description of the Drawings
[0024] To more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required to be used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as a limitation of the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0025] Figure 1 Shows the flowchart of the voice adversarial sample generation method provided by the embodiments of the present application;
[0026] Figure 2 The flowchart of the single-round perturbation optimization method provided by the embodiments of the present application is shown;
[0027] Figure 3 The flowchart of the current-round perturbation determination method provided by the embodiments of the present application is shown;
[0028] Figure 4 The structural schematic diagram of the voice adversarial sample generation device provided by the embodiments of the present application is shown;
[0029] Figure 5 The structural schematic diagram of the electronic device provided by the embodiments of the present application is shown. Detailed implementation manners
[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. Usually, the components of the embodiments of the present application described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application claimed, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, every other embodiment obtained by those skilled in the art without creative efforts belongs to the scope of protection of the present application.
[0031] It should be noted that before the present application was proposed, with the rapid development of artificial intelligence technology, an Automatic Speech Recognition (ASR) system could efficiently convert speech into text using deep learning technology. This technology has been widely applied in industries such as financial services, healthcare, and customer service, improving work efficiency and accuracy. For example, in the financial services industry, automatic speech recognition technology is used for speech monitoring, aiming to enhance the reliability and efficiency of compliance management. However, this extensive application also brings new challenges, especially in terms of privacy protection. Indiscriminate large-scale speech monitoring has raised public concerns about personal privacy leakage, especially globally, in the context of extensive monitoring of personal phone and Internet data. Therefore, how to effectively protect personal privacy while ensuring the quality of voice communication has become an urgent problem to be solved.
[0032] Existing speech processing methods attempt to protect privacy by adding general adversarial perturbations to the audio waveform, making it impossible for automatic speech recognition systems to accurately recognize the speech content. However, this method often leads to a serious decline in audio quality and introduces a large amount of noise, thus affecting normal speech communication. At the same time, existing technologies usually require iterative optimization for specific speech inputs, and different speeches need to be iteratively optimized differently, which cannot meet the requirements of real-time speech communication.
[0033] In addition, since users cannot know the specific monitoring means used by the listener and the specific ASR model is also unknown to users, it is required that the adversarial perturbation must be effective for different ASR models. However, existing technologies are usually designed for a single model and are difficult to adapt to different ASR systems.
[0034] Based on this, the embodiments of the present application provide a method for generating speech adversarial samples to improve the audio quality during the process of adding perturbations to speech and the adaptability to different ASR systems.
[0035] Please refer to Figure 1 , Figure 1 which is a flowchart of a method for generating speech adversarial samples provided by the embodiments of the present application. As Figure 1 shown, the method for generating speech adversarial samples provided by the embodiments of the present application includes:
[0036] Step S101: Receive the input current audio signal, encode the current audio signal in the latent space to obtain the current latent feature encoding;
[0037] Step S102: Select a perturbation from the constructed transferable perturbation set and add it to the current latent feature encoding to obtain the perturbed current latent feature encoding;
[0038] Step S103: Decode the perturbed current latent feature encoding to obtain the current speech adversarial sample.
[0039] The method for generating speech adversarial samples provided by the embodiments of the present application can encode the current audio signal in the latent space and apply perturbations to the perturbed current latent feature encoding to transfer the perturbations to the latent space, avoiding the problem of audio quality degradation caused by introducing noise in the original audio data space and solving the problem of audio quality degradation during the process of adding perturbations to speech.
[0040] For the convenience of understanding this embodiment, the following takes the application of the method for generating speech adversarial samples to a terminal device as an example to separately describe the above exemplary steps provided by the embodiments of the present application.
[0041] In step S101, the input current audio signal is received, and the current audio signal is encoded in the latent space to obtain the current latent feature encoding.
[0042] In this step, the current audio signal may refer to the audio signal currently received by the terminal device, and the current audio signal may be the audio signal input by the user through the microphone.
[0043] The latent space may refer to an abstract high-dimensional space in audio processing, where each point corresponds to a potential audio representation or feature. The latent space is not only used to reduce the data dimension but also to capture and express the essential characteristics of the audio signal.
[0044] In the embodiment of the present application, when the user inputs speech to the terminal device, the terminal device receives the current audio signal input by the user through the microphone of the terminal device, and inputs the current audio signal into a variational autoencoder or an autoencoder for latent space encoding to obtain the current latent feature encoding.
[0045] Among them, the variational autoencoder may refer to VAE (Variational Autoencoder). In the latent space, the input current audio signal is compressed into a lower-dimensional representation, that is, a point in the latent space, after being encoded by the VAE. This low-dimensional representation contains the key features of the input audio, that is, the current latent feature encoding.
[0046] In step S102, a perturbation is selected from the constructed transferable perturbation set and added to the current latent feature encoding to obtain the perturbed current latent feature encoding.
[0047] In this step, the perturbations in the transferable perturbation set are trained in the latent space, and the transferable perturbation set is a constructed set including multiple transferable perturbations.
[0048] The transferable perturbation may refer to a transferable universal adversarial perturbation. Each transferable perturbation can be added to any audio signal to achieve real-time addition of perturbations. Moreover, the perturbation is added in the latent space and does not directly act on the audio input, which can maintain the audio semantics in voice communication and significantly improve the audio quality.
[0049] In addition, each transferable perturbation is generated based on the target feature adaptation process, so that each transferable perturbation can learn the robust features of the target text, enhance its transferability, and remain effective when facing an unknown ASR model, meeting the requirement of ASR model independence.
[0050] In an embodiment of the present application, the transferable perturbation set can be stored in a database. After receiving the current audio signal input by the user, a transferable perturbation is randomly selected from the database, and the transferable perturbation is added to the current latent feature encoding corresponding to the current audio signal, so as to obtain the perturbed current latent feature encoding.
[0051] In step S103, the perturbed current latent feature encoding is decoded to obtain the current voice adversarial sample.
[0052] In this step, the decoder can be used to reconstruct the perturbed current latent feature encoding to obtain the current voice adversarial sample.
[0053] In an embodiment of the present application, the current audio adversarial sample is protected audio. After the perturbation addition process, it has strong anti-attack ability. Then, the protected audio is sent to the virtual microphone of the terminal device, and the protected audio signal is read from the virtual microphone through the set downstream communication software, so as to ensure that the downstream application receives the protected audio, thereby protecting the user's voice privacy.
[0054] In an example, before applying the voice adversarial sample generation method provided by the embodiment of the present application, it is also necessary to construct a transferable perturbation set. The transferable perturbation set is constructed based on multiple optimized transferable perturbations, and each optimized transferable perturbation corresponds to an initial perturbation. Here, multiple initial perturbations can be obtained through a random initialization method, and the iterative optimization process of each initial perturbation is the same. That is, the initial perturbation is iteratively optimized by taking the minimum audio loss as the optimization target to obtain the optimized transferable perturbation corresponding to the initial perturbation. Among them, the audio loss includes a first audio loss representing the loss between the first latent feature encoding corresponding to the target audio and the perturbation, and a second audio loss representing the speech recognition loss.
[0055] The following takes a single initial perturbation as an example to introduce in detail the process of obtaining the optimized transferable perturbation corresponding to the initial perturbation.
[0056] First, determine the target audio corresponding to the initial perturbation.
[0057] The target audio is a fixed audio segment selected through an experimental method, so that the initial perturbation can learn the latent features of the target audio through iterative optimization. Each initial perturbation corresponds to a target audio. During the process of multi-round iterative optimization for a certain initial perturbation, the target audio remains unchanged. Here, the text corresponding to the target audio is called the target text, and the target text is denoted as: t.
[0058] Encoding the target audio through VAE to obtain the first latent feature encoding corresponding to the target audio, thereby realizing the conversion of the target audio into a low-dimensional latent space representation. The first latent feature encoding is denoted as: .
[0059] Then, the initial perturbation is iteratively optimized using the first latent feature encoding to obtain the optimized transferable perturbation corresponding to the initial perturbation. Among them, the iterative optimization process is a multi-round optimization process.
[0060] Next, refer to Figure 2 to introduce the single-round iterative optimization process in detail.
[0061] Figure 2 Fig. shows the flowchart of the single-round perturbation optimization method provided by the embodiment of the present application. As Figure 2 shown, the single-round perturbation optimization method includes:
[0062] Step S201, determining the second latent feature encoding of this round corresponding to the original audio signal of this round.
[0063] In the optimization process of this round, an original audio signal is selected from the set of original audio signals (training data set) as the original audio signal used in the optimization process of this round, which is called the original audio signal of this round. The original audio signal of this round is mapped to the latent space through VAE to obtain the latent feature encoding corresponding to the original audio signal of this round, thereby realizing the conversion of the original audio signal into a low-dimensional latent space representation.
[0064] Here, the original audio signal of this round is denoted as: x, and the encoder is denoted as: , and the latent feature encoding corresponding to the original audio signal x of this round is called the second latent feature encoding of this round. The second latent feature encoding of this round is denoted as: , then there is .
[0065] Step S202, adding the perturbation and noise to the second latent feature encoding of this round to obtain the perturbed second latent feature encoding of this round.
[0066] In the latent space, the perturbation δ of the previous round and the noise parameter p are added to the second latent feature encoding of this round to generate the perturbed second latent feature encoding of this round after adding the perturbation and noise. The perturbed second latent feature encoding of this round after adding the perturbation and noise is denoted as: .
[0067] In addition, the perturbation δ of the previous round can also be added only to the second latent feature encoding of this round in the latent space to obtain the perturbed second latent feature encoding of this round, so as to decode the perturbed second latent feature encoding of this round to obtain the original audio adversarial sample of this round.
[0068] Step S203: Decode the second-round latent feature encoding after adding perturbation and noise to obtain the original audio adversarial sample of this round.
[0069] Use the decoder to decode the second-round latent feature encoding after adding perturbation and noise to obtain the original audio adversarial sample of this round. The original audio adversarial sample of this round is denoted as: , and the decoder is denoted as: , then there is .
[0070] Step S204: Determine the perturbation of this round using the original audio adversarial sample of this round.
[0071] To improve the transferability of the perturbation, the target feature adaptive method can be used for optimization to minimize the cosine similarity loss between the first latent feature encoding corresponding to the target audio and the perturbation, ensuring that the perturbation can be transferred to the target audio and enabling the initial perturbation to learn the latent features of the target audio.
[0072] Next, refer to Figure 3 to introduce the determination process of the perturbation of this round in detail.
[0073] Figure 3 shows the flowchart of the perturbation determination method provided by the embodiment of the present application. As shown in Figure 3 , the perturbation determination method of this round includes:
[0074] Step S2041: Determine the total audio loss of this round according to the total optimization objective.
[0075] In one case, when the influence of the physical environment on the perturbation is not considered, the total optimization objective can be expressed as:
[0076] ;
[0077] In the above formula, represents the expectation, represents the data distribution of natural audio in the physical world, x represents the sampling of this expectation in , represents the hyperparameter for adjusting the balance between the two loss terms, f represents the ASR model for converting audio to text, represents the noise parameter. The minimum can achieve the effect of making the perturbation learn the latent features of the target audio, and the minimum can achieve the effect of interfering with the ASR model.
[0078] In another case, to improve the robustness of perturbations in the real physical environment, the Room Impulse Response (RIR) can be integrated into the perturbation optimization process, that is, the physical environment impact is added to the optimization objective. Among them, RIR is used to simulate the environmental effects in the audio propagation process, such as signal attenuation, multipath effects, and environmental noise, which will affect the effectiveness of adversarial perturbations.
[0079] For example: adding the physical environment impact to the loss function , the impact of the physical environment on the perturbation is characterized by the noise parameter and the acoustic environment parameter. The noise parameter is denoted as: , and the acoustic environment parameter is denoted as: .
[0080] First, through the random sampling method, the first parameter value corresponding to the noise parameter p and the second parameter value corresponding to the acoustic environment parameter are determined. The first parameter value is added to the second potential feature encoding of this round after the perturbation to determine the original audio adversarial sample of this round with added noise. The original audio adversarial sample of this round with added noise is denoted as: .
[0081] Then, based on the original audio adversarial sample of this round with added noise , the acoustic environment parameter r, and the target text t corresponding to the target audio, the second audio loss of this round is determined. At this time, the transcription text corresponding to can be calculated first, and the convolution result of the second parameter value r and the original audio adversarial sample of this round with added noise is determined. Substituting the convolution result and the target audio into the preset loss function , the second audio loss of this round can be determined. At this time, the second audio loss of this round can be expressed as: .
[0082] When adding the impact of the physical environment on the perturbation, the optimization objective can be expressed as:
[0083] ;
[0084] In the above formula, R represents the RIR distribution in the physical environment, r represents the parameter value obtained by sampling the RIR distribution once in each round of iterative optimization process, represents the Gaussian distribution, ⊗ represents the convolution operation, which is used to simulate the environmental impact on the audio signal during propagation; p represents the Gaussian noise obtained by sampling the Gaussian distribution during each round of iterative optimization, and this Gaussian noise is added to the perturbation, represents the variance used to control the noise, which is used to increase the diversity of the input and avoid overfitting.
[0085] In addition, the constraint condition of the total optimization objective is , where is used to control the perturbation range, which is a set value. At this time, by integrating RIR into the perturbation optimization process, the generated optimized transferable perturbation can have stronger robustness in the physical environment, so as to effectively improve the audio privacy protection ability in different environments.
[0086] Here, the total audio loss of this round includes the first audio loss of this round and the second audio loss of this round. At this time, it is necessary to calculate the first audio loss of this round and the second audio loss of this round respectively.
[0087] For the first audio loss of this round, regardless of whether the influence of the physical environment on the perturbation is added, the calculation method of the first audio loss of this round is the same. When calculating the first audio loss of this round, it can be based on the first latent feature encoding and the perturbation of the previous round to determine the first audio loss of this round. For example: calculate the cosine similarity loss between the first latent feature encoding corresponding to the target audio and the perturbation of the previous round, and determine this cosine similarity loss as the first audio loss of this round. The first audio loss of this round is denoted as: .
[0088] For the second audio loss of this round, when calculating the second audio loss of this round without adding the influence of the physical environment on the perturbation, it can be based on the original audio adversarial sample of this round, the target audio and the perturbation of the previous round to determine the second audio loss of this round. For example: first determine the original audio adversarial sample of this round corresponding transcription text , and substitute the transcription text and the target text t corresponding to the target audio into the CTC loss function or the cross-entropy loss function to obtain the calculation result of the loss function, and determine this calculation result of the loss function as the second audio loss of this round. The second audio loss of this round is denoted as: .
[0089] At this time, according to the total optimization objective, the total audio loss of this round can be determined. The total audio loss of this round is denoted as: , .
[0090] For the second audio loss of this round, when calculating the second audio loss of this round with the influence of the physical environment on the perturbation added, it can be based on the original audio adversarial sample of this round, the target audio and the perturbation of the previous round to determine the second audio loss of this round. For example: first determine the original audio adversarial sample of this round , then determine the convolution result of the second parameter value and the original audio adversarial sample of this round , and use the convolution result The corresponding transcribed text Substitute the target text t corresponding to the target audio into the CTC loss function or the cross - entropy loss function to obtain the calculation result of the loss function, and determine the calculation result of this loss function as the second audio loss of this round. The second audio loss of this round is denoted as: .
[0091] At this time, according to the overall optimization objective, the overall audio loss of this round can be determined. The overall audio loss of this round is denoted as: , .
[0092] Step S2042: Determine the perturbation of this round according to the overall audio loss of this round.
[0093] The perturbation of this round can be determined based on the first audio loss of this round and the second audio loss of this round, so as to use the perturbation of this round for subsequent optimization processes.
[0094] For example: Use the Projected Gradient Descent (PGD) method to iteratively optimize the perturbation of this round. The iterative formula of δ is as follows:
[0095] ;
[0096] In the above formula, represents the learning rate.
[0097] Through the above formula, the perturbation of this round can be determined. Then, use the perturbation of this round for the next - round optimization process. In the next - round optimization process, a new original audio signal will be selected to continue the iterative optimization. And so on, until the iteration stop condition is met, the optimization for the initial perturbation ends, and the optimized transferable perturbation is obtained. Among them, the iteration stop condition can refer to the number of iterations reaching the set number threshold.
[0098] In this way, through multiple - round iterative optimization, the optimization objective of minimizing the overall loss is achieved, and the optimized transferable perturbation can be made universal and effective for different audios.
[0099] Through the above method, an optimized transferable perturbation corresponding to a perturbation can be determined. By performing iterative optimization for multiple initial perturbations respectively, multiple optimized transferable perturbations can be obtained. A set of transferable perturbations is composed of multiple optimized transferable perturbations. When actually used, a random one is selected from this set of transferable perturbations and added to the current audio signal.
[0100] Based on the same inventive concept, embodiments of the present application also provide a voice adversarial sample generation device corresponding to the voice adversarial sample generation method. Since the principle of problem-solving of the device in the embodiments of the present application is similar to that of the above voice adversarial sample generation method in the embodiments of the present application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be elaborated.
[0101] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of a voice adversarial sample generation device provided by an embodiment of the present application. As Figure 4 shown in
[0102] The encoding module 301 is configured to receive the input current audio signal, encode the current audio signal in the latent space, and obtain the current latent feature encoding;
[0103] The perturbation addition module 302 is configured to select a perturbation from the constructed transferable perturbation set and add it to the current latent feature encoding to obtain the perturbed current latent feature encoding, and the perturbations in the transferable perturbation set are trained in the latent space;
[0104] The decoding module 303 is configured to decode the perturbed current latent feature encoding to obtain the current voice adversarial sample.
[0105] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 5 shown in
[0106] The memory 420 stores machine-readable instructions executable by the processor 410. When the electronic device 400 runs, the processor 410 communicates with the memory 420 through the bus 430. When the machine-readable instructions are executed by the processor 410, the steps of the voice adversarial sample generation method in the method embodiment as shown in the above Figure 1 can be executed. The specific implementation manner can refer to the method embodiment and will not be elaborated here.
[0107] Embodiments of the present application also provide a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is run by a processor, the steps of the voice adversarial sample generation method in the method embodiment as shown in the above Figure 1 can be executed. The specific implementation manner can refer to the method embodiment and will not be elaborated here.
[0108] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0109] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some communication interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0110] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0111] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0112] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0113] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present application, used to illustrate the technical solutions of the present application, rather than limiting it. The protection scope of the present application is not limited thereto. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the technical field of the present application can still modify the technical solutions recorded in the foregoing embodiments or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A method for generating speech adversarial samples, characterized in that: The method comprises: Receiving an input current audio signal, encoding the current audio signal in a latent space, and obtaining a current latent feature code; Selecting a perturbation from the constructed transferable perturbation set and adding it to the current potential feature code to obtain the perturbed current potential feature code, wherein the perturbations in the transferable perturbation set are obtained by training in the latent space, and the transferable perturbation is a transferable universal adversarial perturbation; Decoding the current potential feature code after the disturbance to obtain a current speech adversarial sample; Before receiving the input current audio signal, the method further includes: For each initial disturbance, iteratively optimize the initial disturbance with minimum audio loss as the optimization goal to obtain an optimized transferable disturbance, wherein the audio loss includes a first audio loss representing the loss between the first latent feature code corresponding to the target audio and the disturbance and a second audio loss representing the speech recognition loss; Based on multiple optimized transferable perturbations, a transferable perturbation set is constructed.
2. The method according to claim 1, characterized in that: The encoding of the current audio signal in the latent space to obtain the current latent feature encoding includes: The current audio signal is input into a variational autoencoder for latent space encoding to obtain a current latent feature code.
3. The method according to claim 1, characterized in that The iterative optimization of the initial disturbance with the minimum audio loss as the optimization objective to obtain the optimized transferable disturbance includes: For this round of optimization process, based on the first potential feature encoding and the previous round of disturbance, determining the first audio loss of this round; Determine the loss of the second audio of this round based on the original audio adversarial sample of this round, the target audio and the disturbance of the previous round; Based on the first audio loss of the current round, the second audio loss of the current round and the disturbance of the previous round, the disturbance of the current round is determined, so as to use the disturbance of the current round to perform a subsequent optimization process until an optimized transferable disturbance is obtained.
4. The method according to claim 3, characterized in that The original audio adversarial sample of this round is determined by the following processing: Selecting a current round of original audio signals from the original audio signal set, and encoding the current round of original audio signals in a latent space to obtain a current round of second latent feature encoding; Adding the previous round of disturbance to the second potential feature code of this round in the latent space to generate the perturbed second potential feature code of this round; After decoding the perturbed second latent feature code of this round, the original audio adversarial sample of this round is obtained.
5. The method according to claim 4, characterized in that The method further comprises: For this round of optimization process, determine a first parameter value corresponding to the noise parameter and a second parameter value corresponding to the acoustic environment parameter; Adding the first parameter value to the second potential feature code of this round after the disturbance, and decoding the second potential feature code of this round after the disturbance is added to obtain the original audio adversarial sample of this round; The loss of the second audio of this round is determined based on the original audio adversarial sample of this round, the second parameter value corresponding to the acoustic environment parameter and the target audio.
6. The method according to claim 5, characterized in that The determining the loss of the second audio of the current round based on the original audio adversarial sample of the current round, the second parameter value corresponding to the acoustic environment parameter and the target audio includes: Determine a convolution result of the second parameter value and the current round of original audio adversarial sample; Substitute the convolution result and the target audio into a preset loss function to determine the second audio loss of this round.
7. The method according to claim 1, characterized in that The decoding of the current potential feature code after the disturbance to obtain the current speech adversarial sample includes: The decoder is used to reconstruct the current latent feature encoding after the disturbance to obtain the current speech adversarial sample.
8. A speech adversarial sample generation device, characterized in that: include: An encoding module, used for receiving an input current audio signal, encoding the current audio signal in a latent space, and obtaining a current latent feature code; A disturbance adding module, used for selecting a disturbance from the constructed transferable disturbance set and adding it to the current potential feature code to obtain the current potential feature code after disturbance, wherein the disturbance in the transferable disturbance set is obtained by training in the latent space, and the transferable disturbance is a transferable universal adversarial disturbance; A decoding module, used for decoding the current potential feature code after the disturbance to obtain a current speech adversarial sample; The encoding module is also used to: For each initial disturbance, iteratively optimize the initial disturbance with minimum audio loss as the optimization goal to obtain an optimized transferable disturbance, wherein the audio loss includes a first audio loss representing the loss between the first latent feature code corresponding to the target audio and the disturbance and a second audio loss representing the speech recognition loss; Based on multiple optimized transferable perturbations, a transferable perturbation set is constructed.
9. An electronic device, characterized in that: include: A processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the memory via the bus, and the processor executes the machine-readable instructions to perform the steps of the method for generating speech adversarial samples as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Voice confrontation sample generation method and system, and storage medium
CN117636857A