Adversarial sample generation method and device, electronic equipment and storage medium
By generating a text set, filtering reference audio, constructing a target loss function, and determining the perturbation amount, this method solves the problem of generating adversarial examples in Chinese speech recognition. It achieves adversarial example generation with high robustness, high transferability, and good sound quality, and is applicable to various speech recognition models.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-08
- Publication Date
- 2026-03-27
AI Technical Summary
Existing adversarial example generation methods cannot generate Chinese speech recognition adversarial examples with high concealment and high transferability, especially for Chinese speech recognition models. They are difficult to overcome the interference of different tones and homophones, and the sound quality is poor.
By generating a text set, filtering reference audio, analyzing the decoding results of the speech recognition model, constructing a target loss function, and determining the target perturbation amount, adversarial examples are generated based on the original audio and the loss value of the target loss function. Natural language processing and speech synthesis technologies are used to enrich the target command, a target loss function is designed to enhance the target pronunciation features, and the perturbation amount is searched through the gradient descent algorithm.
It generates highly robust, highly transferable, and high-quality Chinese speech recognition adversarial examples that can mislead intelligent speech control models to perform specific speech control behaviors in a way that is difficult for humans to detect. It is applicable to both classic and novel speech recognition models.
Smart Images

Figure CN115641836B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the fields of deep learning, information security, artificial intelligence, and the like, and in particular to an adversarial sample generation method and device, an electronic device, and a storage medium. BACKGROUND
[0002] Intelligent voice systems empowered by Chinese speech recognition technology are widely used, and existing adversarial sample generation methods mainly involve English voice commands, and cannot generate Chinese speech recognition adversarial samples with high concealment and high transferability. Therefore, in order to better support the academic and industrial communities to propose more reliable speech recognition algorithms, and to meet the demand for serving the reliability and security of intelligent speech recognition systems, how to generate Chinese speech recognition adversarial samples with high robustness, high transferability, and good sound quality has become a technical problem to be solved. SUMMARY
[0003] The present application provides an adversarial sample generation method, device, electronic device, and storage medium.
[0004] According to a first aspect of the present application, an adversarial sample generation method is provided, comprising:
[0005] generating a text set for a target result misleading a speech recognition model output;
[0006] generating a sound file set based on the text set, the sound file set including a plurality of candidate audios;
[0007] selecting a reference audio of the target result from the plurality of candidate audios;
[0008] analyzing a decoding result of the speech recognition model on the reference audio to obtain a target pronunciation feature of the target result;
[0009] constructing a target loss function based on the target pronunciation feature;
[0010] determining a target perturbation amount based on the original audio and a loss value of the target loss function;
[0011] generating an adversarial sample based on the original audio and the target perturbation amount.
[0012] According to a second aspect of the present application, an adversarial sample generation device is provided, comprising:
[0013] a first generation unit configured to generate a text set for a target result misleading a speech recognition model output;
[0014] a second generation unit configured to generate a sound file set based on the text set, the sound file set including a plurality of candidate audios;
[0015] The screening unit is configured to screen a reference audio of a target result from a plurality of candidate audios;
[0016] The analysis unit is configured to analyze a decoding result of the reference audio by the speech recognition model to obtain a target pronunciation feature of the target result.
[0017] The construction unit is configured to construct a target loss function based on the target pronunciation feature.
[0018] The determination unit is configured to determine a target perturbation amount based on the original audio and a loss value of the target loss function.
[0019] The third generation unit is configured to generate an adversarial sample based on the original audio and the target perturbation amount.
[0020] According to a third aspect of the present application, an electronic device is provided, comprising:
[0021] at least one processor; and
[0022] a memory in communication with the at least one processor; wherein
[0023] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method provided by any one of the embodiments of the present application.
[0024] According to a fourth aspect of the present application, a non-transitory computer-readable storage medium storing computer instructions is provided, the computer instructions being used to cause the computer to perform the method provided by any one of the embodiments of the present application.
[0025] According to a fifth aspect of the present application, a computer program product is provided, comprising computer instructions, which, when executed by a processor, implement the method provided by any one of the embodiments of the present application.
[0026] By using the present application, the determined target perturbation amount is more consistent with the target pronunciation feature of the target result, thereby helping to generate a Chinese speech recognition adversarial sample with high robustness, high transferability and good concealment.
[0027] It should be understood that the contents described in this part are not intended to identify key or important features of the embodiments of the present application, nor to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0028] The accompanying drawings serve to better understand the present application and do not constitute limitations thereof. Among them:
[0029] Figure 1 is a flowchart of an adversarial sample generation method according to an embodiment of the present application;
[0030] Figure 2 is a flowchart of selecting a target command and reference audio according to an embodiment of the present application;
[0031] Figure 3 is a flowchart of designing a target loss function according to an embodiment of the present application;
[0032] Figure 4 is a flowchart of iteratively searching for a perturbation according to an embodiment of the present application;
[0033] Figure 5 is a flowchart of adversarial sample misleading model testing according to an embodiment of the present application;
[0034] Figure 6 is an application example of adversarial sample testing in a digital world according to an embodiment of the present application;
[0035] Figure 7 is an application example of adversarial sample testing in a physical world according to an embodiment of the present application;
[0036] Figure 8 is an architecture diagram of an adversarial sample generation system according to an embodiment of the present application;
[0037] Figure 9 is a component structure diagram of an adversarial sample generation device according to an embodiment of the present application;
[0038] Figure 10 is a block diagram of an electronic device for implementing an adversarial sample generation method according to an embodiment of the present application. DETAILED DESCRIPTION
[0039] Exemplary embodiments of the present application are described herein with reference to the accompanying drawings, which are meant to be exemplary. It should be understood that various changes and modifications to the embodiments described herein will be apparent to those skilled in the art without departing from the scope and spirit of the application. Similarly, it should be understood that the drawings are not intended to be to scale and that, where appropriate, certain features have been exaggerated or omitted in order to more clearly illustrate and describe the present embodiments.
[0040] The term "and / or", used in the present document, only describes the association relationship of the associated objects, and indicates that there can be three relationships, for example, A and / or B can represent three cases of A alone, A and B together, and B alone. The term "at least one" in the present document means any one of a plurality or any combination of at least two of a plurality, for example, at least one of A, B and C includes any one or more elements selected from the set consisting of A, B and C. The terms "first", "second" in the present document mean to refer to a plurality of similar technical terms and to distinguish them, and do not mean to limit the order or mean to limit only two, for example, the first feature and the second feature refer to two categories / two features, the first feature can be one or more, and the second feature can also be one or more.
[0041] Before introducing the technical solutions of the embodiments of the present application, the technical background of the present application is further described.
[0042] According to the degree of understanding of the model by the adversarial sample generator, the model can be divided into "white box model" and "black box model". Among them, the white box model refers to that the parameters and model structure are known; the black box model refers to that people can only get the final output result of the model. At present, the adversarial samples for speech recognition mainly involve new end-to-end (End-to-End) white box model, classic hybrid (Hybrid) white box model and commercial black box model. The key of the white box model adversarial sample generation algorithm is to design a special target loss function, and the main purpose of the target loss function is to strengthen the probability of the acoustic features of the target text in the audio. If an audio file can make the target loss function converge, the final decoding result of the model for the audio is the target text, that is, the audio misleads the machine to recognize the target result. If a sample can mislead 2 or more than 2 models at the same time, it is said that the sample has transferability on these models. Limited by the information of the black box model not being disclosed, the adversarial samples for the black box model mainly use the transferability of the adversarial samples.
[0043] The adversarial sample for the black box model is mainly based on the transferability of the adversarial sample. Most samples with robustness and transferability can achieve the effect that people cannot hear the target command, but the sample may be abnormal in hearing, and therefore the sound quality of the adversarial sample still needs to be improved. In addition, in the related art, the main target is the English voice control command, that is, the adversarial sample is recognized as special English text by misleading the machine, and the voice recognition adversarial sample with the target result as Chinese text is unknown. Compared with English pronunciation, Chinese voice recognition is more complex. For example, there are one, two, three, four and light tone in the pronunciation of Mandarin, such as: eight (bā), pull (bá), put (bǎ), dad (bà), and stop (ba); and there are many homophonic words in Chinese vocabulary, such as "duty" and "plant", "money" and "forearm", "telephone" and "electrification". It can be seen that generating a Chinese voice recognition model not only needs to embed the acoustic features of the target pronunciation, but also needs to overcome the interference of different tones and homophonic words. Therefore, it is very difficult to generate a Chinese voice recognition adversarial sample with high robustness, high transferability and good concealment.
[0044] According to an embodiment of the present application, an adversarial sample generation method is provided, Figure 1 is a flowchart of an adversarial sample generation method according to an embodiment of the present application. The method can be applied to an adversarial sample generation device, for example, the device can be deployed in an electronic device, and can generate an adversarial sample. Wherein, the electronic device includes but is not limited to fixed equipment and / or mobile equipment. For example, the fixed equipment includes but is not limited to a server, which can be a cloud server or a general server. For example, the mobile equipment includes but is not limited to: mobile phone, tablet computer, vehicle terminal. In some possible implementation ways, the adversarial sample generation method can also be realized by the way of calling the computer readable instructions stored in the memory by the processor. As Figure 1 shown, the adversarial sample generation method includes:
[0045] S101: generating a text set for a target result output by a misleading voice recognition model;
[0046] S102: generating a sound file set based on the text set, the sound file set including a plurality of candidate audios;
[0047] S103: selecting a reference audio of the target result from the plurality of candidate audios;
[0048] S104: analyzing a decoding result of the voice recognition model on the reference audio to obtain a target pronunciation feature of the target result;
[0049] S105: constructing a target loss function based on the target pronunciation feature;
[0050] S106: determining a target perturbation amount based on the original audio and a loss value of the target loss function;
[0051] S107: generating an adversarial sample based on the original audio and the target perturbation amount.
[0052] In the embodiments of the present application, the original audio is a carrier for generating an adversarial sample. For example, the original audio can be a piece of music, or a piece of natural noise. The present disclosure does not limit the source of the original audio.
[0053] In the embodiments of the present application, the speech recognition model can be a classic Hybrid speech recognition model, or a new end-to-end speech recognition model. The present disclosure does not limit the type of speech recognition model.
[0054] In the embodiments of the present application, the target command is used to mislead the speech recognition model to output a target result. For example, the target command is "call mom", and after the speech recognition model receives the target command, it dials the mom's phone and outputs the target result "calling mom's phone". For another example, the target command is "turn on the air conditioner", and after the speech recognition model receives the target command, it sends an open instruction to the air conditioner and outputs the target result "air conditioner turned on". For another example, the target command is "open the security door", and after the speech recognition model receives the target command, it outputs an open instruction to the security door and outputs the target result "security door opened".
[0055] In the embodiments of the present application, the sound file set is generated based on the text set, which includes: synthesizing sound files by using a speech synthesis technology (Text-to-Speech, TTS). Further, when synthesizing the sound files, appropriate speed conversion can be performed, and appropriate noise can also be added. In this way, it helps to improve the complexity and authenticity of the synthesized sound files.
[0056] In the embodiments of the present application, the generated adversarial sample can mislead the intelligent speech control model to recognize the adversarial sample audio file as a target command text, and further execute a specific speech control behavior, but a human cannot hear the target command.
[0057] The method for generating an adversarial sample provided in the embodiments of the present application can enrich the number of target commands based on the text set of the target result generated by the target command; the accuracy of the reference audio can be improved based on the text set to generate a sound file set to filter the reference audio of the target result from multiple candidate audios; the target pronunciation feature of the target result is obtained by analyzing the decoding result of the reference audio by the speech recognition model; the target loss function is constructed based on the target pronunciation feature; the target perturbation quantity is determined based on the original audio and the loss value of the target loss function, and the adversarial sample is generated based on the original audio and the target perturbation quantity, which can make the determined target perturbation quantity more consistent with the target pronunciation feature, thereby helping to generate a Chinese speech recognition adversarial sample with high robustness, high transferability and good concealment.
[0058] In some embodiments, the text set includes a target command corresponding to a target result and an expanded command obtained based on the target command. S101 can include:
[0059] S1011: Extracting a plurality of keywords of the target command;
[0060] S1012: Determining a grade corresponding to each of the plurality of keywords, the grade being used to represent an importance;
[0061] S1013: Determining candidate keywords and non-candidate keywords based on the grades corresponding to the plurality of keywords respectively;
[0062] S1014: Processing the non-candidate keywords to obtain expandable expanded words;
[0063] S1015: Obtaining at least one expanded command of the target command based on the candidate keywords and the expanded words;
[0064] S1016: Generating a text set based on the target command and the at least one expanded command.
[0065] As shown in Table 1, first, based on the sentence "call mom" (denoted as TA0), the sentence is divided by using natural language processing technology (NLP), i.e., the keywords "give", "mom", "call", and "phone" are obtained. Among them, the verb "call", the noun "phone" and "mom" are candidate keywords, "give" is a non-candidate keyword, and "I" and "now" are expanded words. Then, under the premise of retaining important words or replacing their synonyms, other words are processed by deletion, expansion, synonym replacement, etc., and M new phrases or sentences are formed after transformation, which form a text set TA together with TA0, i.e., including TA0, TA1, TA2, TA3... TAM.
[0066]
[0067] Table 1
[0068] In this way, the text set of target commands can be enriched, which helps to improve the diversity of the sound file set generated based on the text set, thereby helping to improve the diversity of the Chinese speech recognition adversarial sample.
[0069] In some embodiments, S1016 can include:
[0070] S10161: converting each augmented command into a foreign language version of the augmented command by using a first translation software;
[0071] S10162: converting the foreign language version of the augmented command into a Chinese version of the augmented command by using a second translation software;
[0072] S10163: determining the semantic similarity between each augmented command and the corresponding converted Chinese version of the augmented command;
[0073] S10164: determining the augmented command with a semantic similarity greater than a first threshold value as an available augmented command;
[0074] S10165: generating a text set based on the target command and the available augmented command.
[0075] In the embodiments of the present disclosure, the first translation software is software for translating Chinese into a foreign language.
[0076] In the embodiments of the present disclosure, the second translation software is software for translating a foreign language into Chinese.
[0077] As shown in Table 2, the TA is translated into English by using the translation software, denoted as TAE; and then the English is translated back into Chinese, denoted as TAEC; the semantic similarity of each pair of texts before and after translation in TA and TAEC is compared, and if the semantic similarity of a certain pair of TAi and TAECi is greater than a first threshold value (Simi_Th), it is considered that the semantic augmentation manner processing for TAi is reasonable. In this process, the intelligibility of the sentences in the set TA is detected by using the English-Chinese mutual translation mode.
[0078]
[0079]
[0080] Table 2
[0081] In some embodiments, S103 can include:
[0082] S1031: if the speech recognition model can obtain a target result based on the played candidate audio, determining the candidate audio as the reference audio; or
[0083] S1032: If the voice recognition model can obtain the target result based on the played candidate audio, and the transformation degree of the candidate audio with respect to the target audio corresponding to the target result is greater than a preset threshold, the candidate audio is determined as the reference audio.
[0084] It should be noted that S1031 and S1032 are in parallel relationship, and two options are executed.
[0085] In the embodiments of the present disclosure, the preset threshold can be set or adjusted according to requirements.
[0086] In actual application, when selecting the reference audio, the sound file is input into the voice recognition model, and the text with a large transformation degree and which can be correctly recognized is screened out. Since the model has strong recognition ability for them, these texts are taken as the target command, and the audio files corresponding to them are taken as the reference audio of the target result.
[0087] In this way, the accuracy of the determined reference audio can be improved, thereby helping to improve the accuracy of the determined target pronunciation feature.
[0088] In order to make the adversarial sample have good migration and robustness, the pronunciation feature of the target result embedded in the adversarial sample needs to be as similar as possible to the feature of normal speech. Therefore, the audio with better pronunciation quality and close to daily language is taken as the reference. Figure 2 The flowchart of the selection of the target command and the reference audio is shown as Figure 2 As shown in the figure, first, the NLP technology is used to process the target command to obtain a text set; the voice synthesis tool is used to synthesize the audio of the target command selected from the text set, denoted as TAi_TTS.wav; then, the audio is processed by speed change and noise addition and played to the intelligent voice assistant for recognition. If the played sound can make the intelligent voice assistant execute the action of the target command with a high success rate, it is considered that the transformed TAi is a voice command sensitive to the model; finally, TAi is taken as the target command, and TAi_TTS.wav is taken as the reference audio thereof.
[0089] In some embodiments, S104 can include:
[0090] S1041: Extracting syllables or characters decoded by the voice recognition model for each frame of the reference audio;
[0091] S1042: Determining the pronunciation duration and the corresponding probability density function index of each frame of the decoded syllables or characters;
[0092] S1043: Forming a target sequence by all syllables or characters included in the reference audio;
[0093] S1044: Obtain the target pronunciation feature of the target result based on the pronunciation duration of each syllable or character in the target sequence and the corresponding probability density function index.
[0094] In the embodiments of the present disclosure, the target pronunciation feature comprises at least one of the following: a syllable and a pronunciation duration corresponding to the syllable; a character and a pronunciation duration corresponding to the character; and a probability density function of the syllable / character.
[0095] For example, assuming that the reference audio of the target result is N frames, the syllables or characters obtained by decoding each frame of the reference audio by using the speech recognition model are extracted. The classic speech recognition model can extract syllables, and the new end-to-end speech recognition model can extract characters. If the syllables or characters of each frame are connected in sequence and combined to obtain the target text S, wherein S is composed of syllables or characters S1S2...S n ...S N The result of decoding the reference audio can be labeled frame by frame as the target sequence S1, S2,..., S n ...S N , and the corresponding syllable or character duration is T0~T1, T1~T2,..., T n-1 ~T n ...T N-1 ~T N .
[0096] In this way, the target pronunciation feature with strong relevance to the target command can be quickly and accurately determined, thereby helping to generate a Chinese speech recognition adversarial sample with high robustness, high transferability, good hiding property, and good sound quality.
[0097] In some embodiments, S106 can comprise:
[0098] Obtaining the to-be-tested audio determined according to the original audio and the perturbation amount;
[0099] Designing a target loss function based on the real probability value of the to-be-tested audio decoded by the speech recognition model as the target syllable or character, and the expected probability value of the target pronunciation feature to the target syllable or character;
[0100] Determining the target perturbation amount based on the target loss function and the original audio.
[0101] Here, the target loss function needs to strengthen the features of the target pronunciation, and the probability of the target syllable or character is expected to be the maximum, so the probability of the target syllable or character is set to be much larger than other values in the target loss function.
[0102] Here, the to-be-tested audio x'(t)=x0(t)+δ(t), wherein x'(t) is the to-be-tested audio, x0(t) is the original audio, and δ(t) is the perturbation amount.
[0103] The adversarial disturbance quantity can be calculated by a gradient descent algorithm, and the target is to make x'(t) promote the convergence of the target loss function, and x'(t) still converges after being superimposed with the random noise μ(t).
[0104] The smallest unit in the classical Chinese speech recognition model is a syllable (i.e., an initial and a final of a word pronunciation), and the smallest unit in the new end-to-end Chinese speech recognition model is a character (i.e., a single Chinese character), and the purpose of designing the target loss function is to expect that the acoustic features of the target command embedded in the adversarial sample, the syllable and the character are different, but the design idea of the target loss function is consistent, and both are to make the syllable or character sequence decoded by the modified audio consistent with the reference audio.
[0105] In this way, the features of the target pronunciation are strengthened by the target loss function, which can improve the accuracy of the target disturbance quantity, thereby helping to improve the success rate of the generated adversarial sample.
[0106] In some embodiments, designing the target loss function can further include at least one of the following:
[0107] using a band-pass filter to limit the frequency domain of the disturbance quantity δ(t);
[0108] using a preset norm to limit the time domain of the disturbance quantity δ(t) after the frequency domain limitation;
[0109] using a preset norm to limit the time domain of the random noise μ(t), which is noise simulating the influence of the physical environment and used to add noise to the to-be-tested audio x'(t).
[0110] In the embodiments of the present disclosure, the preset norm can be an Lp norm.
[0111] Figure 3 The flowchart of designing the target loss function is shown as follows. Figure 3 As shown in the figure, first, the decoding result of the speech recognition model is analyzed, assuming that the reference audio of the target result is N frames, and the syllable or character decoded by the model for each frame of the reference audio is extracted, wherein the classical speech recognition model extracts syllables, and the new end-to-end speech recognition model extracts characters. Specifically, the result of decoding the reference audio TAi_TTS.wav by the white-box model is analyzed, since it can be recognized as the target text S by the model, then the syllable or character decoded by the model for each frame of TAi_TTS.wav can be extracted; secondly, each frame of the extracted syllable / character is regarded as a string sequence, i.e., S is regarded as a string sequence composed of syllables / characters S1S2...S n ...S Ncomposition, in which the same syllable / character can appear several times (i.e. the same pronunciation state lasts for a period of time), therefore, the duration of each syllable / character is annotated as t0~t1, t1~t2,..., t n-1 ~t n ,..., t N-1 ~t N . The target loss function expects x'(t) to be decoded as the probability of the target syllable / character as large as possible, where pdf represents the probability density function for calculating the probability of a phoneme in the model, and pdf-S n represents the probability density function number (also the probability density function index number) for calculating the probability of this syllable / character (S n ) in the model. The time-domain signal y(t n-1 , t n ) in the t n-1 ~t n period is input into the speech recognition model, represents the probability value calculated by the probability density function with index number , max1P[y(t n-1 , t n )] represents the maximum value of the probability calculated by all index probability density functions, and max2P[y(t n-1 , t n )] represents the second maximum value of the probability calculated by all index probability density functions. To strengthen the probability of the target pronunciation, the target loss function can be made to maximize the probability calculated by the probability density function with index pdf-S n . Therefore, the value of should be as large as possible in the target loss function.
[0112] The calculated perturbation γ(t) is input into a band-pass filter to obtain the perturbation δ(t) after removing high-frequency noise, and finally the perturbation δ(t) is added to the original audio x0(t) to obtain the perturbed audio x'(t). In addition, to improve the robustness of the sample, random noise μ(t) is introduced in each iteration, so that x'(t) can overcome the interference of μ(t), and the combination of the two can still be decoded by the model as the target sequence. At the same time, this scheme can use the Lp norm to constrain the time-domain sample values of δ(t) and μ(t) to the ranges of a and b, respectively, where a and b are positive numbers.
[0113] The target loss function L can be represented as:
[0114]
[0115] where str(S) = str(S1) +... + str(S n ), β ≥ 0.
[0116] wherein y(t n-1 ,t n ) = x'(t n-1 ,t n ) + μ(t n-1 ,t n ) ;
[0117] wherein x'(t n-1 ,t n ) = x0(t n-1 ,t n ) + δ(t n-1 ,t n ) ;
[0118] wherein δ(t) = bandpass[γ(t),f l ,f h ], ||δ(t)|| < a, ||μ(t)|| < b; p p
[0119] wherein, if α ≥ 0, max{α, 0} = α;
[0120] wherein, if α < 0, max{α, 0} = 0.
[0121] In this way, the rationality of the design of the target loss function can be improved, thereby helping to improve the concealment, robustness and transferability of the generated adversarial samples.
[0122] The perturbation design is iteratively searched, and the adversarial perturbation amount can be calculated by feedback of the gradient descent algorithm. The goal is to make the to-be-tested audio x'(t) = x(t) + δ(t) promote the convergence of the target loss function, and x'(t) still converges after being superimposed with the random noise μ(t). The search method can use various iterative search algorithms. In some embodiments, S107 can include: based on the target loss function, searching for the perturbation amount by using an iterative search algorithm, and stopping the iterative search when an iterative stopping condition is reached; determining the perturbation amount obtained when the iterative search is stopped as a target perturbation amount; adding the target perturbation amount to the original audio to generate a candidate sample; and in a case where the candidate sample can be decoded by the speech recognition model as the target result, determining the candidate sample as the adversarial sample.
[0123] wherein, in the iterative process, further comprising: setting an iterative stopping condition according to a change in a loss value (loss) in the search process.
[0124] wherein the iterative stopping condition comprises one of the following:
[0125] a maximum number of iterations is reached;
[0126] The loss value does not continue to decrease within a certain period of time in the iterative search;
[0127] The maximum number of iterations is not reached, but the loss value has decreased to a second threshold value.
[0128] In practical applications, the iteration stopping condition is set according to the change of the loss value in the search process, such as: reaching the maximum number of iterations; the loss value does not continue to decrease within a certain period of time in the iterative search; although the maximum number of iterations is not reached, the loss value has decreased to a smaller threshold, indicating that the target loss function converges quickly, and the iterative search can be stopped.
[0129] Exemplarily, given an original audio x0(t) and a designed target loss function, the adversarial perturbation amount δ(t) can be searched by using various gradient descent algorithms, so that x'(t) = x0(t) + δ(t) converging to the target loss function will be decoded as the target text with a high probability. In the process of iteratively calculating the perturbation amount, multiple important parameters need to be set or optimized. Suggestions are as follows:
[0130] 1) The upper sideband cutoff frequency f l and the lower sideband cutoff frequency f h of the band-pass filter, since the human ear is sensitive to high-frequency noise, and some devices have a cutoff for sound signals below 100 Hz, it is recommended to set 100 Hz ≤ f l ≤ 4000 Hz. h
[0131] 2) To further reduce the perturbation noise, it is recommended to limit the perturbation amount δ(t) and μ(t) to 20 / 1 and 10 / 1 of the maximum value of the original audio sampling value, respectively.
[0132] 3) Set the maximum iteration threshold to max_iter = 1000, and stop the search when the number of iterations reaches max_iter.
[0133] 4) If the loss has converged to a smaller threshold LTh within max_iter times, the search can also be stopped.
[0134] 5) Due to unreasonable parameter settings or the carrier being unsuitable for embedding the target command, the loss may not continue to decrease in the iteration process, therefore, the loss change trend is counted during the search process, if the loss does not continue to decrease within max_invalid_iter (such as max_invalid_iter = 100), the search is stopped. Then, the original audio is abandoned as a carrier of the target command, or the limit amplitude of δ(t) and μ(t) is enlarged and then the original audio is reattempted to modify.
[0135] Thus, the optimal target disturbance amount can be determined, so that the Chinese speech recognition adversarial sample with high robustness, high transferability and good concealment can be generated.
[0136] Figure 4 The flowchart of iterative search of adversarial disturbance is shown as follows: Figure 4 The specific process is as follows:
[0137] 1) Given the original audio x0(t), the disturbance amount γ(t) is calculated by using the gradient descent algorithm, which is frequency domain limited by the bandpass filter BandpassFilter(γ(t),f l ,f h ), and then time domain limited by the Lp norm, so that the spectrum of the new disturbance amount δ(t) is limited in the range of f l ~f h , and the Lp norm of the time domain sample value is limited in the range of a, that is:
[0138] δ(t)=BandpassFilter(γ(t),f l ,f h ),||δ(t)||≤a.
[0139] 2) In the gradient calculation process, y(t) is obtained after adding random noise μ(t) to δ(t), wherein the Lp norm of μ(t) is limited in the range of b, that is:
[0140] y(t)=δ(t)+μ(t),||μ(t)||≤b.
[0141] 3) Since the target syllable or character is divided into multiple segments, the target loss function L is divided into N segments, wherein the target loss function of each segment is set to l n , which mainly calculates the loss of the time segment, that is:
[0142]
[0143] 4) y(t) is input into the speech recognition model, and the probability value of the syllable / character in each frame of audio decoding result is extracted, which is represented as P[y(t n-1 ~t n )], which represents a vector composed of the probability values of the n-th frame of audio signal decoded by all probability density functions in the model, wherein the dimension K of the vector is determined by the type of the probability density function in the model, max1P[y(t n-1 ~t n )] represents the maximum value in the K-dimensional vector, max2P[y(t n-1 ~t n )] represents the second maximum value in the K-dimensional vector, P pdf-Sn represents the index of the pdf-Sn The probability value calculated by the probability density function of P n-1 ~t n )], max2P[y(t n-1 ~t n )] and the index P pdf-Sn The probability value calculated by the probability density function of P
[0144] In some embodiments, the method for generating an adversarial sample can further include:
[0145] inputting the adversarial sample into a white-box type speech recognition model to obtain a first test result;
[0146] If the white-box type speech recognition model achieves the target result, inputting the adversarial sample into a black-box type speech recognition model to obtain a second test result;
[0147] determining the transferability of the adversarial sample based on the first test result and the second test result.
[0148] In some embodiments, the method for generating an adversarial sample can further include:
[0149] using a preset parameter index to measure the perturbation scale of the adversarial sample relative to the original audio.
[0150] Here, the preset parameter index can be set or adjusted according to requirements.
[0151] In this way, the good and bad of the adversarial sample can be better verified.
[0152] In some embodiments, the method for generating an adversarial sample can further include:
[0153] using survey data obtained from a questionnaire survey to evaluate the concealment of the adversarial sample relative to the original audio.
[0154] In this way, the good and bad of the adversarial sample can be better evaluated.
[0155] In order to test the transferability and robustness of the generated adversarial sample, white-box testing and black-box testing can be performed in sequence. The flowchart of the adversarial sample misleading model testing is as follows: Figure 5As shown: if the loss of the adversarial sample x'(t) is less than a certain test threshold, it is considered that better audio features have been embedded, and then it is input into the white-box speech recognition model to test the recognition result. If the white-box model can successfully recognize the target text (target command), the adversarial sample is input into the black-box speech recognition model for testing, otherwise the adversarial sample is deleted. If the black-box speech recognition model can successfully recognize the target text (target command), it is determined that the adversarial sample has transferability, otherwise the adversarial sample is deleted.
[0156] The adversarial sample is tested in the digital world and applied, which is divided into:
[0157] White-box speech recognition model digital world test: for the iteratively modified audio, if the target loss function value is less than a certain threshold, it is considered that the target loss function converges, and it is input into the white-box speech recognition model for testing, and the audio file that can successfully mislead the white-box speech recognition model is screened out.
[0158] Black-box speech recognition model digital world test: upload the successful adversarial sample of the white-box speech recognition model to the black-box application programming interface (Application Programming Interface, API) to test the transferability of the adversarial sample.
[0159] An example of testing and applying the adversarial sample in the digital world is shown in Figure 6 As shown, read or upload an audio file (adversarial sample), and the speech recognition model interface identifies the adversarial sample. If the target command is identified, the corresponding operation of the target command can be performed. For example, the target command is malicious text content, such as misleading reading certain content, converting speech to text, or displaying certain content on a machine subtitle. Here, the malicious text content refers to a command not issued by the user.
[0160] In the case where the white-box speech recognition model can successfully recognize the target text (target command) or the black-box speech recognition model can successfully recognize the target text (target command), the adversarial sample can also be repeatedly played in the physical world to test the robustness of the adversarial sample. In the physical world, if the success rate of the speech recognition model in recognizing the target text (target command) is greater than a certain threshold, the current adversarial sample is determined to be a robust sample; otherwise, the current adversarial sample is determined to be a non-robust sample. Here, a robust sample refers to a sample with strong robustness, and a non-robust sample refers to a sample with weak robustness. The strength of robustness can be determined by a preset index representing the strength of robustness. The preset index can be that the sample is tested X times and succeeds Y times, and the robustness index is X / Y*100%.
[0161] The adversarial sample is tested in the physical world and applied, which is divided into:
[0162] White-box speech recognition model simulation physical world test: play the adversarial sample with the device and record the played sound, convert the recorded file into a format that can be decoded by the white-box speech recognition model, input the model and decode, test the robustness of the adversarial sample in the simulated physical world.
[0163] Black-box speech recognition model physical world test: for smart voice assistants or smart speakers with actual application scenarios, samples in the black-box digital world can be played directly, and the recognition effect of smart voice assistants and smart speakers can be detected to evaluate the effect of misleading the device in the real physical world. Repeat the test to evaluate the robustness of the sample.
[0164] The application example of the adversarial sample in the physical world test is shown in Figure 7 The device plays the adversarial sample, and the smart voice assistant recognizes the target command based on the played adversarial sample, and then executes the operation corresponding to the target command. For example, the smart voice assistant recognizes malicious text content, such as controlling bound social information or smart devices, and then executes the operation of controlling the bound social information or smart devices. Here, the malicious text content refers to a command not issued by the user.
[0165] The present application provides a Chinese speech recognition adversarial sample generation method based on feature enhancement. For the target result of misleading the model, a diversified target command text set is designed using natural language processing; according to the sensitivity of the model, a voice command that is easy to succeed is selected; the target pronunciation feature is strengthened in the target loss function, and the frequency domain and time domain of the disturbance are limited; based on the frequency spectrum of the original audio, the gradient descent algorithm is used to search for the local minimum disturbance amount, so that the target loss function converges to a certain threshold; the adversarial samples of the white-box model and the black-box model are obtained. The beneficial effects include at least:
[0166] 1) Special adversarial sample generation for high sensitivity target command: the sensitivity of the model to special pronunciation features can be used by using diversified voice commands, so that the model can be misled to recognize the target command with a small modification to the original audio, and the result of misleading the model can be achieved with the smallest disturbance amount, to improve the success rate of generating adversarial samples.
[0167] 2) In the target loss function, the disturbance amount is limited to low frequency using a band-pass filter to reduce human perception noise caused by high-frequency disturbance and improve the sound quality of the adversarial sample.
[0168] 3) This method is not only suitable for classical speech recognition models trained based on syllables as the smallest unit, but also suitable for end-to-end speech recognition models based on characters as the smallest unit, and can generate adversarial samples with high success rate, high transferability and high concealment.
[0169] The present application provides an architecture diagram of an adversarial sample generation system, as shown inFigure 8 As shown, the system mainly includes four modules of command setting and reference preparation, algorithm design and perturbation search, sample testing on the model and adversarial sample concealment evaluation. Specifically, it includes six parts: 1) target command setting; 2) reference audio selection; 3) target loss function design; 4) adversarial perturbation search; 5) adversarial sample testing; 6) adversarial sample concealment evaluation, etc. The main goal of the present application is to design the description text of the voice command (i.e. the target command) by using the sensitivity of the model, and to generate the adversarial sample by strengthening the acoustic features of the target command in the perturbation noise.
[0170] Firstly, the target text set is designed by using NLP to mislead the expected results of the model; then, the easy-to-success text is selected as the target command and the reference audio according to the sensitivity of the model to pronunciation; secondly, the target pronunciation sequence is designed according to the decoding results of each frame of the reference audio by the model, the acoustic features of the target pronunciation are strengthened in the target loss function design, and the random noise model is used to simulate the interference of the physical environment, so as to improve the success rate of the sample to mislead the model; the perturbation amount is limited by using the band-pass filter and Lp norm, which is used to improve the concealment of the added perturbation; thirdly, in the process of multiple iteration calculation by using gradient descent, the maximum number of iterations, the number of invalid search and the loss threshold of stopping search are used to control the number of iteration search, so as to improve the efficiency of adversarial sample generation; finally, the recognition success rate, transferability and human perception effect of the modified audio on the model are tested.
[0171] The scheme of the present application is applicable to the classical Hybrid speech recognition model and the new type of end-to-end speech recognition model, can generate Chinese voice command adversarial samples for the white box model of Chinese speech recognition, has high transferability on the black box model, and the concealment of the sample is good.
[0172] The present application solves the blank of the current Chinese speech recognition command adversarial sample generation method. Among them, the natural language understanding technology is used to automatically generate a set of voice commands with similar functions, which can well utilize the sensitivity of the model to special commands to generate adversarial samples, thereby improving the success rate of misleading the speech recognition model; the target loss function design proposes a feature enhancement-based adversarial sample generation method, and random noise is used to improve the robustness of the adversarial sample, and the perturbation amount is limited by using the band-pass filter and Lp norm; the present application is applicable to traditional speech recognition algorithms (phoneme-level recognition) and end-to-end speech recognition algorithms (character-level recognition), and can generate robust, transferable and well-concealed adversarial samples. The generated adversarial samples can support the vulnerability mining and repair of the speech recognition algorithm, so as to protect the safe application of the intelligent speech recognition system.
[0173] The method mainly analyzes the principle of Chinese speech recognition, proposes an adversarial sample generation algorithm based on feature enhancement, realizes the generation of Chinese speech recognition adversarial samples with high concealment and high migration, and supports the academic and industrial fields to propose more reliable speech recognition algorithms to serve the reliability requirements of intelligent speech recognition systems.
[0174] The scheme described in the application can be applied to evaluate and support the security of automatic speech recognition technology in practical application scenarios, such as intelligent voice assistants, smart speakers, voice control systems for autonomous driving, and voice transcription applications.
[0175] It should be understood that Figures 2 to 8 The schematic diagram shown is merely exemplary and not limiting, and it is extensible, and those skilled in the art can make various obvious changes and / or replacements based on Figures 2 to 8 The resulting technical solutions still belong to the disclosure range of the embodiments of the application.
[0176] The embodiments of the application provide an adversarial sample generation device, as shown in the figure Figure 9 The adversarial sample generation device can include: a first generation unit 901 for generating a text set for a target result output by a misleading speech recognition model; a second generation unit 902 for generating a sound file set based on the text set, the sound file set including a plurality of candidate audios; a screening unit 903 for screening a reference audio of the target result from the plurality of candidate audios; an analysis unit 904 for analyzing a decoding result of the speech recognition model on the reference audio to obtain a target pronunciation feature of the target result; a construction unit 905 for constructing a target loss function based on the target pronunciation feature; a determination unit 906 for determining a target perturbation amount based on an original audio and a loss value of the target loss function; and a third generation unit 907 for generating an adversarial sample based on the original audio and the target perturbation amount.
[0177] In a possible implementation, the text set includes a target command corresponding to the target result and an expanded command obtained based on the target command. The first generation unit 901 is specifically configured to: extract a plurality of keywords of the target command; determine a grade corresponding to each of the plurality of keywords, the grade being used to represent an importance; determine a candidate keyword and a non-candidate keyword based on the grade corresponding to each of the plurality of keywords; process the non-candidate keyword to obtain an expandable expansion word; obtain at least one expanded command of the target command based on the candidate keyword and the expansion word; and generate the text set based on the target command and the at least one expanded command.
[0178] In a possible implementation, the first generating unit 901 is further configured to: convert each extended command into a foreign language version of the extended command by using first translation software; convert the foreign language version of the extended command into a Chinese language version of the extended command by using second translation software; determine semantic similarity between each extended command and the corresponding converted Chinese language version of the extended command; determine the extended command with a semantic similarity greater than a first threshold value as an available extended command; and generate the text set based on the target command and the available extended command.
[0179] In a possible implementation, the screening unit 903 is configured to: determine the candidate audio as the reference audio if the voice recognition model can obtain the target result based on the played candidate audio.
[0180] In a possible implementation, the screening unit 903 is configured to: determine the candidate audio as the reference audio if the voice recognition model can obtain the target result based on the played candidate audio, and a transformation degree of the candidate audio relative to the target audio corresponding to the target result is greater than a preset threshold value.
[0181] In a possible implementation, the analysis unit 904 is configured to: extract a syllable or a character decoded by the voice recognition model from each frame of the reference audio; determine a pronunciation duration of the syllable or the character decoded from each frame, and determine a probability density function of the syllable or the character; form a target sequence by using the syllable or the character decoded from each frame of the reference audio; and obtain a target pronunciation feature of the target result based on the pronunciation duration of each syllable or character in the target sequence and the corresponding probability density function index.
[0182] In a possible implementation, the determining unit 906 is configured to: obtain the to-be-tested audio determined according to the original audio and the perturbation amount; design a target loss function based on a real probability value of the to-be-tested audio being decoded into a target syllable or character by the voice recognition model, and an expected probability value of the target pronunciation feature on the target syllable or character; and determine the target perturbation amount based on the target loss function.
[0183] In a possible implementation, the determining unit 906 is further configured to: perform frequency domain limitation on the perturbation amount by using a band-pass filter; perform time domain limitation on the perturbation amount subjected to the frequency domain limitation by using a preset norm; and perform time domain limitation on random noise by using the preset norm, the random noise being noise simulating a physical environment.
[0184] In a possible implementation, the third generation unit 907 is specifically configured to: search for the perturbation amount based on the target loss function by using an iterative search algorithm, and stop the iterative search when an iterative stop condition is reached; determine the perturbation amount obtained when the iterative search is stopped as the target perturbation amount; and add the target perturbation amount to the original audio to generate a candidate sample; and determine the candidate sample as the adversarial sample in a case where the candidate sample can be decoded into the target result by the speech recognition model.
[0185] In a possible implementation, the third generation unit 907 is further specifically configured to: set the iterative stop condition according to a change in the loss value of the target loss function in the search process during the iteration; and the iterative stop condition includes one of the following: a maximum number of iterations is reached; the loss value does not continuously decrease within a period of time of the iterative search; the maximum number of iterations is not reached, but the loss value has decreased to a second threshold value.
[0186] In a possible implementation, the adversarial sample generation apparatus can further include a test unit configured to: input the adversarial sample into a white-box type speech recognition model to obtain a first test result; input the adversarial sample into a black-box type speech recognition model to obtain a second test result, if the white-box type speech recognition model implements the target result as represented by the first test result; and determine the transferability of the adversarial sample based on the first test result and the second test result.
[0187] In a possible implementation, the adversarial sample generation apparatus can further include a measurement unit configured to: measure a perturbation scale of the adversarial sample relative to the original audio by using a preset parameter index.
[0188] In a possible implementation, the adversarial sample generation apparatus can further include an evaluation unit configured to: evaluate the concealment of the adversarial sample relative to the original audio by using survey data obtained from a questionnaire survey.
[0189] Those skilled in the art shall understand that the functions of the processing units in the adversarial sample generation apparatus of the embodiments of the present application can be understood with reference to the foregoing related descriptions of the adversarial sample generation method, and the processing units in the adversarial sample generation apparatus of the embodiments of the present application can be implemented by an analog circuit that implements the functions described in the embodiments of the present application, or can be implemented by the running of software that implements the functions described in the embodiments of the present application on an electronic device.
[0190] The adversarial sample generation apparatus of the embodiments of the present application can generate a Chinese speech recognition adversarial sample with high robustness, high transferability and good concealment.
[0191] In the technical solutions of the present application, the acquisition, storage and application of user personal information comply with relevant laws and regulations and do not violate public order and good customs.
[0192] According to embodiments of the present application, the present application also provides an electronic device, a readable storage medium and a computer program product.
[0193] Figure 10 A schematic block diagram of an example electronic device 1000 that can be used to implement embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present application described and / or claimed in this document.
[0194] As shown in Figure 10 The device 1000 includes a computing unit 1001 that can perform various appropriate actions and processes in accordance with a computer program stored in a Read-Only Memory (ROM) 1002 or a computer program loaded into a Random Access Memory (RAM) 1003 from a storage unit 1008. Various programs and data required for the operation of the device 1000 can also be stored in the RAM 1003. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other through a bus 1004. An Input / Output (I / O) interface 1005 is also connected to the bus 1004.
[0195] Various components in the device 1000 are connected to the I / O interface 1005, including an input unit 1006, such as a keyboard, a mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; the storage unit 1008, such as magnetic disks, optical disks, etc.; and a communication unit 1009, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1009 allows the device 1000 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0196] The computing unit 1001 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), various specialized Artificial Intelligence (AI) computing chips, various computing units running machine learning model algorithms, a Digital Signal Processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs various methods and processes described above, such as the adversarial sample generation method. For example, in some embodiments, the adversarial sample generation method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded onto the RAM 1003 and executed by the computing unit 1001, one or more steps of the adversarial sample generation method described above can be performed. Alternatively, in other embodiments, the computing unit 1001 can be configured to perform the adversarial sample generation method by any other suitable means, such as by means of firmware.
[0197] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a Field Programmable Gate Array (FPGA), an Application Specific Integrated Circuit (ASIC), an Application-Specific Standard Products (ASSP), a System on Chip (SOC), a Complex Programmable Logic Device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0198] Program code for carrying out methods of the present application can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces a means for implementing the functions / acts specified in the flowcharts and / or block diagrams. The program code can be retrieved from a machine-readable medium or device and executed by a processor to produce a machine for implementing the functions / acts specified in the flowcharts and / or block diagrams. The program code can be executed as a stand-alone program, or in combination with other program codes, on the machine to produce a machine that can implement the functions / acts specified in the flowcharts and / or block diagrams.
[0199] In the context of the present application, a machine-readable medium can be a tangible medium that can contain or store program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include one or more lines of electrical wire, portable computer diskette, hard disk, random access memory, read-only memory, erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination of the foregoing.
[0200] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a Cathode Ray Tube (CRT) or Liquid Crystal Display (LCD) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0201] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0202] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server combined with a blockchain. It should be understood that the various forms of flow shown above can be re-ordered, added to, or deleted from without departing from the scope of the present disclosure. For example, the steps recited in the present disclosure can be executed in parallel, executed in sequence, or executed in different orders, as long as the desired results of the present disclosure are achieved. The present disclosure is not limited in this regard.
[0203] The specific embodiments described above are not intended to limit the scope of the present application. Those skilled in the art will understand that various modifications, combinations, sub-combinations, and alternatives can be made to the specific embodiments without departing from the application. Any modifications, equivalent substitutions, improvements, and the like, made within the principles of the present application are intended to be included in the scope of the application.
Claims
1. A method for generating adversarial examples, characterized in that, The method includes: For the target result output by the misleading speech recognition model, a text set is generated; wherein the text set includes the target command corresponding to the target result, and at least one expanded command obtained based on the target command through keyword extraction, level division and word expansion processing; A set of audio files is generated based on the text set, and the set of audio files includes multiple candidate audio files; The reference audio for the target result is selected from the plurality of candidate audios; wherein the selection includes: testing the ability of each candidate audio to be recognized as the target result by the speech recognition model, and determining the candidate audio that is successfully recognized as the reference audio; The decoding results of the speech recognition model on the reference audio are analyzed to obtain the target pronunciation features of the target result; wherein the target pronunciation features include the syllables or characters decoded from each frame extracted from the reference audio, the pronunciation duration, and the corresponding probability density function index; A target loss function is constructed based on the target pronunciation features; wherein the target loss function is configured as the probability value of strengthening the target syllable or character; The target perturbation is determined based on the original audio and the loss value of the target loss function; wherein determining the target perturbation includes using a bandpass filter to limit the perturbation in the frequency domain and using a preset norm to limit it in the time domain; Adversarial examples are generated based on the original audio and the target perturbation.
2. The method according to claim 1, characterized in that, The text set includes the target command corresponding to the target result and the extended command obtained based on the target command. Generating the text set based on the target result output by the misleading speech recognition model includes: Extract multiple keywords from the target command; Determine the levels corresponding to the multiple keywords, where the levels are used to indicate the degree of importance; Based on the respective levels of the multiple keywords, candidate keywords and non-candidate keywords are determined; The non-candidate keywords are processed to obtain expandable keywords; Based on the candidate keywords and the expanded keywords, at least one expanded command of the target command is obtained; The text set is generated based on the target command and the at least one extended command.
3. The method according to claim 2, characterized in that, Based on the target command and the at least one extended command, the text set is generated, including: Each expansion command was converted into a foreign language version using the first translation software; The foreign language version of the extended command was converted into a Chinese version using a second translation software. Determine the semantic similarity between each extended command and its corresponding converted Chinese version extended command; The expanded commands with semantic similarity greater than the first threshold value are determined as available expanded commands; The text set is generated based on the target command and the available extended commands.
4. The method according to claim 1, characterized in that, The reference audio for selecting the target result from the plurality of candidate audios includes: If the speech recognition model can obtain the target result based on the played candidate audio, then the candidate audio is determined as the reference audio; or If the speech recognition model can obtain the target result based on the played candidate audio, and the degree of change of the candidate audio relative to the target audio corresponding to the target result is greater than a preset threshold, then the candidate audio is determined as the reference audio.
5. The method according to claim 1, characterized in that, The analysis of the decoding results of the speech recognition model on the reference audio to obtain the target pronunciation features of the target result includes: Extract the syllables or characters obtained by the speech recognition model from decoding each frame of the reference audio; Determine the pronunciation duration and corresponding probability density function index of the syllable or character obtained from decoding each frame; The syllables or characters obtained from decoding each frame of the reference audio are combined into a target sequence; Based on the pronunciation duration and corresponding probability density function index of each syllable or character in the target sequence, the target pronunciation features of the target result are obtained.
6. The method according to claim 5, characterized in that, The determination of the target perturbation based on the loss value of the original audio and the target loss function includes: Obtain the audio to be tested, determined based on the original audio and the perturbation amount; Based on the true probability value of the target syllable or character being decoded by the speech recognition model from the audio to be tested, and the expected probability value of the target pronunciation feature for the target syllable or character, a target loss function is designed. The target perturbation amount is determined based on the loss value of the target loss function and the original audio.
7. The method according to claim 6, characterized in that, Designing the target loss function also includes: The disturbance is frequency-domain limited using a bandpass filter; The perturbation quantity that has been limited in the frequency domain is time-domain limited using a preset norm; The random noise is time-domain limited by a preset norm, and the random noise is noise that simulates the influence of the physical environment.
8. The method according to claim 6 or 7, characterized in that, Generate adversarial examples based on the original audio and the target perturbation, including: Based on the target loss function, an iterative search algorithm is used to search for the disturbance amount, and the iterative search is stopped when the iteration cutoff condition is reached; The perturbation obtained when the iterative search stops is determined as the target perturbation. The target perturbation amount is added to the original audio to generate candidate samples; If the candidate sample can be decoded into the target result by the speech recognition model, the candidate sample is identified as the adversarial sample.
9. The method according to claim 8, characterized in that, The iteration process also includes: Based on the change in the target loss function value during the search process, set the iteration cutoff condition; The iteration cutoff condition includes one of the following: Reaching the maximum number of iterations; The loss value of the iterative search does not decrease over a period of time; The maximum number of iterations has not been reached, but the loss value has dropped to the second threshold.
10. The method according to claim 1, characterized in that, Also includes: The adversarial sample is input into a white-box speech recognition model to obtain the first test result; If the first test result indicates that the white-box type speech recognition model achieves the target result, then the adversarial sample is input into the black-box type speech recognition model to obtain the second test result; Based on the first test results and the second test results, the transferability of the adversarial example is determined.
11. The method according to claim 1, characterized in that, Also includes: The perturbation scale of the adversarial example relative to the original audio is measured using preset parameter indicators; or The stealth of the adversarial example relative to the original audio was assessed using survey data obtained from a questionnaire survey.
12. An adversarial sample generation device, characterized in that, The device includes: The first generation unit is used to generate a text set for the target result output by the misleading speech recognition model; wherein the text set includes the target command corresponding to the target result, and at least one expanded command obtained based on the target command through keyword extraction, level division and word expansion processing; The second generation unit is used to generate a set of sound files based on the text set, the set of sound files including multiple candidate audio files; A filtering unit is used to filter reference audio for the target result from the plurality of candidate audios; wherein the filtering includes: testing the ability of each candidate audio to be recognized as the target result by the speech recognition model, and determining the successfully recognized candidate audio as the reference audio; The analysis unit is used to analyze the decoding result of the speech recognition model on the reference audio to obtain the target pronunciation features of the target result; wherein the target pronunciation features include the syllables or characters decoded from each frame extracted from the reference audio, the pronunciation duration, and the corresponding probability density function index; A construction unit is configured to construct a target loss function based on the target pronunciation features; wherein the target loss function is configured as a probability value for strengthening a target syllable or character; The determining unit is used to determine the target perturbation amount based on the original audio and the loss value of the target loss function; wherein determining the target perturbation amount includes using a bandpass filter to limit the perturbation amount in the frequency domain and using a preset norm to limit it in the time domain; The third generation unit is used to generate adversarial examples based on the original audio and the target perturbation.
13. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-11.
14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-11.
15. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-11.
Citation Information
Patent Citations
Firefly algorithm and gradient evaluation-based confrontation audio generation method and system
CN113345420A
Adversarial text generation method and system for black box text classification model and medium
CN113886559A
Voice confrontation sample generation method and device, electronic equipment and storage medium
CN114267363A