Method, apparatus, medium and device for generating speech adversarial examples of significant sequences
By positioning significant sequences in the speech sequence and adding adversarial samples to generate adversarial perturbations, the problems of low efficiency of speech adversarial sample generation and insufficient noise robustness in the prior art are solved, and more efficient and robust adversarial sample generation is achieved.
Patent Information
- Application Number
- CN202111531120.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-14
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-12-14
AI Technical Summary
Existing voice adversarial samples are susceptible to noise interference when applied in the physical world, have low generation efficiency and are not robust to noise.
By locating significant sequences from phonological sequences and adding counter-perturbation to generate adversarial samples on the significant sequences, the generation efficiency is improved and the robustness to noise is enhanced.
The generation efficiency of speech adversarial samples is improved and the robustness to noise is enhanced, making the generated adversarial samples more effective in practical applications.
Smart Images

Figure CN114283797B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of speech recognition technology, and in particular, to a method, device, medium, and equipment for generating speech adversarial samples of significant sequences. Background Art
[0002] With the continuous improvement of the accuracy of deep learning models in the field of speech recognition, some minor interference factors, such as noise interference or other speech sequence interferences, may cause errors in the output of deep learning models, thus seriously affecting the further development of deep learning in this field.
[0003] Samples adjusted by minor interference factors are adversarial samples. As a non-negligible source of errors, adversarial samples seriously hinder the development of speech recognition technology. Therefore, it is crucial to strengthen the research on adversarial samples. Existing speech adversarial samples are usually generated in a laboratory environment.
[0004] When this solution is applied in the physical world, it usually loses its effect due to the influence of noise, and the efficiency of generating speech adversarial samples is low. Summary of the Invention
[0005] Embodiments of the present application provide a method, device, medium, and equipment for generating speech adversarial samples of significant sequences, which can locate significant sequences from speech sequences and generate adversarial samples by adding adversarial perturbations to the significant sequences, thereby achieving the purpose of improving the generation efficiency of speech adversarial samples and enhancing the robustness of speech adversarial samples to noise.
[0006] In a first aspect, embodiments of the present application provide a method for generating speech adversarial samples of significant sequences, the method comprising:
[0007] Determining a significant sequence from a speech sequence based on the speech recognition rules of the speech sequence; wherein, the speech sequence includes a significant sequence and a fixed sequence other than the significant sequence;
[0008] Generating an adversarial perturbation having the same time dimension as the significant sequence through a speech recognition model;
[0009] Adding the adversarial perturbation to the significant sequence in the speech sequence to generate an adversarial sample.
[0010] In a second aspect, embodiments of the present application provide a device for generating speech adversarial samples of significant sequences, the device comprising:
[0011] A significant sequence determination module, configured to determine a significant sequence from a speech sequence based on the speech recognition rules of the speech sequence; wherein, the speech sequence includes a significant sequence and a fixed sequence other than the significant sequence;
[0012] An adversarial perturbation generation module, configured to generate an adversarial perturbation having the same time dimension as the significant sequence through a speech recognition model;
[0013] An adversarial sample generation module, configured to add the adversarial perturbation to the significant sequence in the speech sequence to generate an adversarial sample.
[0014] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method for generating a speech adversarial sample of a significant sequence as described in the embodiments of the present application.
[0015] In a fourth aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the method for generating a speech adversarial sample of a significant sequence as described in the embodiments of the present application.
[0016] The technical solution provided by the embodiments of the present application determines a significant sequence from a speech sequence based on the speech recognition rule of the speech sequence; wherein, the speech sequence includes a significant sequence and a fixed sequence other than the significant sequence; further, an adversarial perturbation having the same time dimension as the significant sequence is generated through a speech recognition model; and then the adversarial perturbation is added to the significant sequence in the speech sequence to generate an adversarial sample. By adopting the technical solution of the present application, the significant sequence can be located from the speech sequence, and an adversarial sample can be generated by adding an adversarial perturbation to the significant sequence, without any processing on the fixed part of the speech sequence, thereby improving the generation efficiency of the speech adversarial sample and enhancing the robustness of the speech adversarial sample to noise. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a flowchart of a method for generating a speech adversarial sample of a significant sequence provided in Embodiment 1 of the present application;
[0018] Figure 2 is a schematic diagram of a process for generating a speech adversarial sample provided in Embodiment 1 of the present application;
[0019] Figure 3 is a flowchart of a method for generating a speech adversarial sample of a significant sequence provided in Embodiment 2 of the present application;
[0020] Figure 4 is a schematic diagram of a speech sequence recognition process provided in Embodiment 2 of the present application;
[0021] Figure 5 is a flowchart of a method for locating a significant sequence provided in Embodiment 2 of the present application;
[0022] Figure 6It is a flowchart of a method for generating speech adversarial samples of a significant sequence provided in Embodiment 3 of the present application;
[0023] Figure 7 It is a structural block diagram of a device for generating speech adversarial samples of a significant sequence provided in Embodiment 4 of the present application;
[0024] Figure 8 It is a schematic structural diagram of an electronic device provided in Embodiment 6 of the present application. Detailed implementation manners
[0025] The present application will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present application, rather than limiting the present application. Additionally, it should be noted that for the sake of description, only parts related to the present application rather than all structures are shown in the accompanying drawings.
[0026] Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the steps as sequential processes, many of the steps can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the steps can be rearranged. The process can be terminated when its operation is completed, but it can also have additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, and so on.
[0027] Embodiment 1
[0028] Figure 1 It is a flowchart of a method for generating speech adversarial samples of a significant sequence provided in Embodiment 1 of the present application. This embodiment is applicable to the scenario of adding adversarial perturbations to the significant sequence located in the speech sequence to generate speech adversarial samples. This method can be executed by the device for generating speech adversarial samples of the significant sequence provided in the embodiments of the present application. The device can be implemented in software and / or hardware and can be integrated into an electronic device.
[0029] As Figure 1 shown, the method for generating speech adversarial samples of the significant sequence includes:
[0030] S110, determining a significant sequence from the speech sequence based on the speech recognition rules of the speech sequence; wherein, the speech sequence includes a significant sequence and a fixed sequence other than the significant sequence.
[0031] Among them, speech recognition can refer to the technology by which a machine converts speech signals into corresponding text or commands by recognizing and understanding speech sequences. The speech recognition rules can refer to the pre-set rules for performing speech recognition on speech sequences. For example, the speech recognition rules can be set according to the characters in the sequence or according to the sequence time length. Specifically, the speech sequence can include two parts: a significant sequence and a fixed sequence. The significant sequence can refer to the set of speech sequences in a speech task that are more likely to cause changes in the classification of the speech recognition model. The fixed sequence can refer to the set of the remaining speech sequences in the speech sequence except for the significant sequence. It can be understood that the significant sequence and the fixed sequence are composed of some speech sequence segments in the speech sequence.
[0032] S120, generate an adversarial perturbation with the same time dimension as the significant sequence through the speech recognition model.
[0033] Among them, the speech recognition model can refer to the network model used for performing speech recognition on speech sequences. For example, the speech recognition model can be a temporal classification model based on a neural network or a Gaussian mixture model. The adversarial perturbation can refer to a tiny interference sequence that can cause changes in the classification of the speech recognition model. For example, the adversarial perturbation can be a random noise sequence or other speech sequences.
[0034] In this embodiment, the generated adversarial perturbation has the same time dimension as the significant sequence, that is to say, they have the same time length. Exemplarily, if the significant sequence is a speech sequence with a total duration of 5 seconds, it can be determined that the adversarial perturbation is also a speech sequence with a total duration of 5 seconds.
[0035] S130, add the adversarial perturbation to the significant sequence in the speech sequence to generate an adversarial sample.
[0036] Among them, the adversarial sample can refer to a sample that has been slightly adjusted. Specifically, when the speech recognition model is attacked by a slightly adjusted sample, it may get an incorrect output, and such a slightly adjusted sample is the adversarial sample. In this embodiment, the adversarial perturbation with the same time dimension as the significant sequence can be directly superimposed on the significant sequence to generate the corresponding adversarial sample.
[0037] Figure 2 It is a schematic diagram of a process for generating a speech adversarial sample provided in the first embodiment of this application. As Figure 2As shown, the significant sequence uses a speech recognition model to recognize the speech signal as the text string "how are you". When adding a tiny adversarial perturbation with the same time length as the significant sequence, an adversarial sample is generated. At this time, the speech recognition model can recognize the adversarial sample as the text string "open the door". Thus, adversarial samples can be generated by adding adversarial perturbations to the significant sequence in the speech sequence, resulting in an output different from the speech recognition result of the significant sequence, that is, an unwanted wrong result.
[0038] The technical solution provided by the embodiments of the present application determines a significant sequence from the speech sequence based on the speech recognition rules of the speech sequence, then generates an adversarial perturbation with the same time dimension as the significant sequence through the speech recognition model, and then adds the adversarial perturbation to the significant sequence in the speech sequence to generate an adversarial sample. It can locate the significant sequence from the speech sequence and generate an adversarial sample by adding an adversarial perturbation to the significant sequence, without any processing on the fixed part of the speech sequence, thereby improving the generation efficiency of the speech adversarial sample and enhancing the robustness of the speech adversarial sample to noise.
[0039] In this embodiment, optionally, after adding the adversarial perturbation to the significant sequence in the speech sequence to generate an adversarial sample, a method for generating a speech adversarial sample of a significant sequence provided by the embodiments of the present application may further include:
[0040] Perform model iterative training on the adversarial sample and determine whether the iteration ends;
[0041] If the iteration ends, output the predicted target text character sequence.
[0042] Among them, the target text character sequence may refer to the text character sequence that the target attacker wants to obtain by adding an adversarial perturbation to the significant sequence in the speech sequence. The predicted target text character sequence may refer to the output text character sequence obtained by inputting the adversarial sample into the speech recognition model and undergoing model training. It can be understood that due to the accuracy problem of the model, there may be a certain deviation between the predicted target text character sequence output by the model and the target text character sequence desired by the attacker, and the specific deviation size depends on the accuracy of the model.
[0043] In this embodiment, after generating adversarial samples, the speech recognition model can be used to iteratively train the adversarial samples. Specifically, the adversarial samples can be input into the speech recognition model, the loss value can be calculated through the CTC loss function, and the gradient can be backpropagated to update the adversarial perturbation. Then, it is judged whether the iteration ends based on the preset loss value and the maximum number of times. If the actual number of iterations reaches the maximum number of times, the iteration stops; alternatively, when the calculated loss value is less than the preset loss value, even if the actual number of iterations does not reach the maximum number of times, the iteration can also stop. If the iteration has not ended, the updated adversarial perturbation is added to the significant sequence in the speech sequence to regenerate the adversarial sample, and the model iteration process is repeated; if the iteration ends, the predicted target text character sequence is directly output.
[0044] With such a setting, this solution can continuously update the adversarial perturbation through the iterative training of the speech recognition model to generate adversarial samples with high accuracy, thereby improving the prediction accuracy of the speech recognition model for the target text character sequence and better meeting the target requirements of the attacker.
[0045] Embodiment 2
[0046] Figure 3 FIG. is a flowchart of a method for generating a speech adversarial sample of a significant sequence provided in Embodiment 2 of this application. This embodiment is optimized based on the above embodiment. The specific optimization is as follows: determining a significant sequence from the speech sequence based on the speech recognition rule of the speech sequence, including: obtaining a predicted character sequence by using the speech recognition rule according to the obtained speech sequence; determining the different characters between the target text character sequence and the predicted character sequence according to the target text character sequence and the predicted character sequence; and determining the speech sequence segment corresponding to the different characters in the predicted character sequence as the significant sequence.
[0047] As Figure 3 shown, the method of this embodiment specifically includes the following steps:
[0048] S310, obtaining a predicted character sequence by using the speech recognition rule according to the obtained speech sequence.
[0049] Among them, the predicted character sequence may refer to the recognition result of the text character sequence without repeated characters and blank characters obtained by performing speech recognition on the speech sequence according to the speech recognition rules. Specifically, the speech sequence can be segmented into several speech sequence segments, and then, with a preset time length as a time unit, each speech sequence segment is respectively subjected to speech recognition in the time order of the speech sequence, and finally several predicted character sequences can be obtained. Among them, a speech sequence segment under one time unit can correspond to the recognition of one text character. It can be understood that the length of the speech sequence segment under one time unit is less than the length of the segmented speech sequence segment, and the length of each predicted character sequence is less than the length of the corresponding speech sequence segment.
[0050] In this embodiment, optionally, obtaining the predicted character sequence according to the acquired speech sequence by using the speech recognition rules may include the following steps:
[0051] Obtain a marker sequence having the same time dimension as the speech sequence according to the speech sequence;
[0052] Obtain the predicted character sequence by deleting the blank symbols and repeated characters in the marker sequence.
[0053] Among them, the marker sequence may refer to the text character sequence containing repeated characters and blank characters obtained during the speech recognition of the speech sequence and having the same time dimension as the speech sequence. Figure 4 It is a schematic diagram of a speech sequence recognition process provided in the second embodiment of the present application. Among them, x represents the speech sequence, y represents the predicted character sequence, and z represents the marker sequence. As Figure 4 shown, by pre-recognizing the speech sequence, a marker sequence that may contain repeated characters and blank characters can be obtained first, and then by deleting the blank symbols and repeated characters contained in the marker sequence, a predicted character sequence without blank symbols and repeated characters can be obtained.
[0054] Through such a setting in this solution, a predicted character sequence without blank symbols and repeated characters can be obtained by performing deletion operations on the blank symbols and repeated characters in the marker sequence, which can effectively avoid the interference of blank symbols and repeated characters on speech recognition and make the speech recognition result more accurate.
[0055] S320. Determine the different characters between the target text character sequence and the predicted character sequence according to the target text character sequence and the predicted character sequence.
[0056] Among them, the differentiating character may refer to the different character obtained by performing character alignment comparison between the target text character sequence and the predicted character sequence. It can be understood that the differentiating character may be located in the target text character sequence or in the predicted character sequence.
[0057] S330, determine the speech sequence segment corresponding to the differentiating character in the predicted character sequence as the significant sequence.
[0058] In this embodiment, the significant sequence can be located according to the differentiating character in the predicted character sequence. Specifically, the differentiating character in the predicted character sequence can be located in the speech sequence, so as to obtain the position of the differentiating character in the speech sequence, and then the speech sequence segment at this position is determined as the significant sequence.
[0059] In this embodiment, optionally, determining the speech sequence segment corresponding to the differentiating character in the predicted character sequence as the significant sequence may include the following steps:
[0060] Determine the sequence position of the differentiating character in the predicted character sequence in the marker sequence according to the differentiating character;
[0061] Determine the speech sequence segment in the speech sequence corresponding to the differentiating character in the predicted character sequence according to the sequence position;
[0062] Determine the speech sequence segment as the significant sequence.
[0063] Among them, the sequence position may refer to the position of the speech sequence segment in the speech sequence in the speech sequence. Figure 5 is a flowchart of a method for locating a significant sequence provided in the second embodiment of the present application. As Figure 5 shown, the differentiating character between the two can be found by comparing the target text character sequence and the predicted character sequence, then the position of the differentiating character in the predicted character sequence in the marker sequence is found, and then the position of the differentiating character in the speech sequence can be correspondingly found according to this position, and the speech sequence segment at the corresponding position is obtained, so that the significant sequence can be located from the speech sequence.
[0064] With such a setting in this solution, the position of the differentiating character in the marker sequence can be correspondingly found according to the differentiating character, and then its position in the speech sequence can be correspondingly found, so that the significant sequence can be quickly and accurately located.
[0065] S340, generate an adversarial perturbation with the same time dimension as the significant sequence through a speech recognition model.
[0066] S350, add the adversarial perturbation to the significant sequence in the speech sequence to generate an adversarial sample.
[0067] Among them, the specific implementation processes of S340 and S350 can refer to the detailed descriptions in S120 and S130.
[0068] The technical solution provided by the embodiment of the present application obtains a predicted character sequence without blank characters and repeated characters by deleting the blank characters and repeated characters in the marked sequence, determines the significant sequence from the speech sequence according to the different characters between the target text character sequence and the predicted character sequence, then generates an adversarial perturbation with the same time dimension as the significant sequence through a speech recognition model, and adds the adversarial perturbation to the significant sequence in the speech sequence to generate an adversarial sample, effectively avoiding the interference of blank symbols and repeated characters on speech recognition, improving the speech recognition accuracy, quickly and accurately locating the significant sequence from the speech sequence through different characters, then adding an adversarial perturbation to the significant sequence to generate an adversarial sample, without making any processing on the fixed part of the speech sequence, improving the generation efficiency of the speech adversarial sample, and at the same time enhancing the robustness of the speech adversarial sample to noise.
[0069] Embodiment III
[0070] This embodiment is a preferred embodiment provided on the basis of Embodiment II. Figure 6 It is a flowchart of a method for generating a speech adversarial sample of a significant sequence provided by Embodiment III of the present application. The method of this embodiment specifically includes the following steps:
[0071] S610, according to the obtained speech sequence, obtain a predicted character sequence by using a speech recognition rule.
[0072] S620, compare the target text character sequence and the predicted character sequence to obtain a short-length sequence.
[0073] Among them, the short-length sequence may refer to the sequence with a shorter length in the target text character sequence and the predicted character sequence. It can be understood that when the target text character sequence and the predicted character sequence are known, it is very easy to determine the sequence with a shorter length between the two.
[0074] S630, insert spaces in the short-length sequence to align the common characters of the target text character sequence and the predicted character sequence.
[0075] Among them, the common characters can refer to the same characters in the target text character sequence and the predicted character sequence. In this embodiment, since the lengths of the target text character sequence and the predicted character sequence may not be the same, in order to compare the target text character sequence and the predicted character sequence character by character, spaces can be inserted into the shorter sequence to make the lengths of the two sequences the same. Specifically, the target text character sequence and the predicted character sequence with the same length can be regarded as two strings arranged in two rows, and spaces can be inserted at any position of the strings to make the two strings achieve the maximum alignment, while requiring that there cannot be space characters at the same time above and below. Among them, the maximum alignment can refer to the highest degree of correspondence between the corresponding characters of the two strings.
[0076] S640, determine the characters other than the common characters as the distinguishing characters.
[0077] In this embodiment, when the alignment degree of the two strings is the highest, that is, when the common string achieves the maximum alignment, the set of common characters can be regarded as non-significant characters, and the remaining uncorresponding characters can be regarded as distinguishing characters, that is, the set of significant characters corresponding to the significant sequence.
[0078] S650, determine the speech sequence segment corresponding to the distinguishing character in the predicted character sequence as the significant sequence.
[0079] S660, generate an adversarial perturbation with the same time dimension as the significant sequence through the speech recognition model.
[0080] In this embodiment, optionally, generating an adversarial perturbation with the same time dimension as the significant sequence through the speech recognition model may include the following steps:
[0081] Generate a perturbed speech sequence segment according to the distinguishing character in the target text character sequence;
[0082] Determine the perturbed speech sequence segment as the adversarial perturbation.
[0083] Among them, the perturbed speech sequence segment can refer to the speech sequence segment corresponding to the adversarial perturbation. In this embodiment, each distinguishing character in the target text character sequence can be respectively corresponding to generate a perturbed speech sequence segment, and the set of each perturbed speech sequence segment is the adversarial perturbation.
[0084] Through such a setting, this solution can generate an adversarial perturbation according to the distinguishing character in the target text character sequence, can specifically apply the perturbation only to the significant sequence, and avoids perturbing the fixed sequence in the speech sequence, thereby reducing the interference of external noise.
[0085] S670, add the adversarial perturbation to the significant sequence in the speech sequence to generate an adversarial sample.
[0086] The technical solution provided by the embodiments of the present application maximally aligns the target text character sequence and the predicted character sequence by inserting spaces, determines the characters other than the common characters as the distinguishing characters, then determines the significant sequence from the speech sequence according to the distinguishing characters between the target text character sequence and the predicted character sequence, and then generates an adversarial perturbation with the same time dimension as the significant sequence through a speech recognition model, and adds the adversarial perturbation to the significant sequence in the speech sequence to generate an adversarial sample. It can quickly locate the significant sequence from the speech sequence through the distinguishing characters, and then add the adversarial perturbation to the significant sequence to generate an adversarial sample, without any processing on the fixed part of the speech sequence, improving the generation efficiency of the speech adversarial sample and enhancing the robustness of the speech adversarial sample to noise.
[0087] Embodiment 4
[0088] Figure 7 is a structural block diagram of a device for generating a speech adversarial sample of a significant sequence provided by Embodiment 4 of the present application. This device can execute the method for generating a speech adversarial sample of a significant sequence provided by any embodiment of the present application, and has the corresponding functional modules and beneficial effects for executing the method. As Figure 7 shown, the device may include:
[0089] A significant sequence determination module 710, configured to determine a significant sequence from a speech sequence based on the speech recognition rule of the speech sequence; wherein, the speech sequence includes a significant sequence and a fixed sequence other than the significant sequence;
[0090] An adversarial perturbation generation module 720, configured to generate an adversarial perturbation with the same time dimension as the significant sequence through a speech recognition model;
[0091] An adversarial sample generation module 730, configured to add the adversarial perturbation to the significant sequence in the speech sequence to generate an adversarial sample.
[0092] Based on the above embodiments, optionally, the significant sequence determination module 710 is specifically configured to:
[0093] Obtain a predicted character sequence by using the speech recognition rule according to the obtained speech sequence;
[0094] Determine the distinguishing characters between the target text character sequence and the predicted character sequence according to the target text character sequence and the predicted character sequence;
[0095] Determine the speech sequence segment corresponding to the distinguishing characters in the predicted character sequence as the significant sequence.
[0096] Based on the above embodiments, optionally, according to the obtained speech sequence, obtaining a predicted character sequence by using a speech recognition rule includes:
[0097] According to the speech sequence, obtaining a marker sequence having the same time dimension as the speech sequence;
[0098] By deleting blank symbols and repeated characters in the marker sequence, obtaining a predicted character sequence.
[0099] Based on the above embodiments, optionally, determining the different characters between the target text character sequence and the predicted character sequence according to the target text character sequence and the predicted character sequence includes:
[0100] Comparing the target text character sequence and the predicted character sequence to obtain a short-length sequence;
[0101] Inserting spaces in the short-length sequence to align the common characters of the target text character sequence and the predicted character sequence;
[0102] Determining the characters other than the common characters as different characters.
[0103] Based on the above embodiments, optionally,
[0104] Determining the speech sequence segment corresponding to the different character in the predicted character sequence as a significant sequence includes:
[0105] According to the different character, determining the sequence position of the different character in the predicted character sequence in the marker sequence;
[0106] According to the sequence position, determining the speech sequence segment in the speech sequence corresponding to the different character in the predicted character sequence;
[0107] Determining the speech sequence segment as a significant sequence.
[0108] Based on the above embodiments, optionally, generating an adversarial perturbation having the same time dimension as the significant sequence through a speech recognition model includes:
[0109] Generating a perturbed speech sequence segment according to the different character in the target text character sequence;
[0110] Determining the perturbed speech sequence segment as an adversarial perturbation.
[0111] Based on the above embodiments, optionally, a speech adversarial sample generation device for a significant sequence provided by an embodiment of the present application further includes:
[0112] A model training module for iteratively training a model with the adversarial samples and determining whether the iteration ends;
[0113] A character sequence output module for outputting a predicted target text character sequence if the iteration ends.
[0114] The above product can execute the method for generating speech adversarial samples of the significant sequence provided in the embodiments of the present application, and has functional modules and beneficial effects corresponding to the execution of the method.
[0115] Embodiment Five
[0116] Embodiment Five of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method for generating speech adversarial samples of the significant sequence provided in all inventive embodiments of the present application:
[0117] Determining a significant sequence from a speech sequence based on the speech recognition rule of the speech sequence; wherein the speech sequence includes a significant sequence and a fixed sequence other than the significant sequence;
[0118] Generating an adversarial perturbation having the same time dimension as the significant sequence through a speech recognition model;
[0119] Adding the adversarial perturbation to the significant sequence in the speech sequence to generate an adversarial sample.
[0120] Any combination of one or more computer-readable media can be used. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system, apparatus, or device.
[0121] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including - but not limited to - electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0122] The program code contained on a computer-readable medium may be transmitted using any appropriate medium, including - but not limited to - wireless, wire, optical fiber cable, RF, etc., or any suitable combination of the above.
[0123] The computer program code for performing the operations of this application may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages - such as Java, Smalltalk, C++, and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or, alternatively, may be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).
[0124] Example Six
[0125] Example Six of this application provides an electronic device. Figure 8 It is a schematic structural diagram of an electronic device provided by Example Six of this application. As Figure 8 shown, this embodiment provides an electronic device 800, which includes: one or more processors 820; a storage device 810 for storing one or more programs, and when the one or more programs are executed by the one or more processors 820, the one or more processors 820 implement the method for generating speech adversarial samples of the significant sequence provided by the embodiments of this application. The method includes:
[0126] Determining a significant sequence from a speech sequence based on the speech recognition rules of the speech sequence; wherein, the speech sequence includes a significant sequence and a fixed sequence other than the significant sequence;
[0127] Generating an adversarial perturbation having the same time dimension as the significant sequence through a speech recognition model;
[0128] Add the adversarial perturbation to the significant sequence in the speech sequence to generate an adversarial sample.
[0129] Of course, those skilled in the art can understand that the processor 820 also implements the technical solution of the method for generating a speech adversarial sample of the significant sequence provided in any embodiment of the present application.
[0130] Figure 8 The displayed electronic device 800 is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.
[0131] As Figure 8 shown, the electronic device 800 includes a processor 820, a storage device 810, an input device 830, and an output device 840; the number of processors 820 in the electronic device can be one or more, Figure 8 and one processor 820 is taken as an example herein; the processor 820, the storage device 810, the input device 830, and the output device 840 in the electronic device can be connected through a bus or other means, Figure 8 and taking the connection through the bus 880 as an example herein.
[0132] The storage device 810, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and module units, such as the program instructions corresponding to the method for generating a speech adversarial sample of the significant sequence in the embodiments of the present application.
[0133] The storage device 810 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the terminal, etc. In addition, the storage device 810 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some instances, the storage device 810 can further include a memory remotely set relative to the processor 820, and these remote memories can be connected through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0134] The input device 830 can be used to receive input digital, character information, or voice information, and generate key signal inputs related to the user settings and function controls of the electronic device. The output device 840 can include electronic devices such as a display screen and a speaker.
[0135] The voice adversarial sample generation device, medium, and electronic device for the significant sequence provided in the above embodiments can execute the voice adversarial sample generation method for the significant sequence provided in any embodiment of the present application, and have the corresponding functional modules and beneficial effects for executing the method. For the technical details not described in detail in the above embodiments, reference can be made to the voice adversarial sample generation method for the significant sequence provided in any embodiment of the present application.
[0136] Note that the above is only the preferred embodiment of the present application and the technical principles applied. Those skilled in the art will understand that the present application is not limited to the specific embodiments described herein, and various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present application. Therefore, although the present application has been described in more detail through the above embodiments, the present application is not limited to the above embodiments. Without departing from the concept of the present application, more other equivalent embodiments can be included, and the scope of the present application is determined by the scope of the appended claims.
Claims
1. A method for generating speech adversarial examples of a significant sequence, characterized in that, The method includes: Determining a significant sequence from a speech sequence based on a speech recognition rule of the speech sequence; wherein, the speech sequence includes a significant sequence and a fixed sequence other than the significant sequence; Generating an adversarial perturbation having the same time dimension as the significant sequence through a speech recognition model; Adding the adversarial perturbation to the significant sequence in the speech sequence to generate an adversarial sample; Wherein, determining a significant sequence from a speech sequence based on a speech recognition rule of the speech sequence includes: According to the obtained speech sequence, obtaining a predicted character sequence by using a speech recognition rule; Determining a different character between the target text character sequence and the predicted character sequence according to the target text character sequence and the predicted character sequence; Determining a speech sequence segment corresponding to the different character in the predicted character sequence as the significant sequence.
2. The method according to claim 1, characterized in that, According to the obtained speech sequence, obtaining a predicted character sequence by using a speech recognition rule, includes: According to the speech sequence, obtaining a marker sequence having the same time dimension as the speech sequence; Obtaining a predicted character sequence by deleting blank symbols and repeated characters in the marker sequence.
3. The method according to claim 1, characterized in that, Determining a different character between the target text character sequence and the predicted character sequence according to the target text character sequence and the predicted character sequence, includes: Comparing the target text character sequence and the predicted character sequence to obtain a short-length sequence; Inserting spaces in the short-length sequence to align common characters of the target text character sequence and the predicted character sequence; Determining characters other than the common characters as different characters.
4. The method according to claim 1, characterized in that, Determining a speech sequence segment corresponding to the different character in the predicted character sequence as the significant sequence, includes: According to the different character, determining a sequence position of the different character in the predicted character sequence in the marker sequence; According to the sequence position, determining a speech sequence segment in the speech sequence corresponding to the different character in the predicted character sequence; Determining the speech sequence segment as the significant sequence.
5. The method according to claim 1, characterized in that Generating an adversarial perturbation having the same time dimension as the significant sequence through a speech recognition model, includes: Generating a perturbed speech sequence segment according to the different character in the target text character sequence; Determining the perturbed speech sequence segment as the adversarial perturbation.
6. The method according to claim 1, characterized in that, After adding the adversarial perturbation to the significant sequence in the speech sequence to generate an adversarial sample, the method further includes: Performing model iterative training on the adversarial sample and determining whether the iteration ends; If the iteration ends, outputting a predicted target text character sequence.
7. A voice adversarial sample generation device for a significant sequence, characterized in that The apparatus includes: A significant sequence determination module, configured to determine a significant sequence from a speech sequence based on a speech recognition rule of the speech sequence; wherein, the speech sequence includes a significant sequence and a fixed sequence other than the significant sequence; An adversarial perturbation generation module, configured to generate an adversarial perturbation having the same time dimension as the significant sequence through a speech recognition model; An adversarial sample generation module, configured to add the adversarial perturbation to the significant sequence in the speech sequence to generate an adversarial sample; Wherein, the significant sequence determination module is specifically configured to: According to the obtained speech sequence, a predicted character sequence is obtained by using speech recognition rules; According to the target text character sequence and the predicted character sequence, the different characters between the target text character sequence and the predicted character sequence are determined; The speech sequence segment corresponding to the different characters in the predicted character sequence is determined as the significant sequence.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the method for generating speech adversarial samples of the significant sequence according to any one of claims 1-6.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method for generating speech adversarial samples of the significant sequence according to any one of claims 1-6.