A method and device for confusing speech samples
Through blocking and multiple obfuscation methods, the problem of insufficient confusion of speech samples in the prior art is solved, and sufficient confusion of massive speech samples is achieved, and the stability of model training is improved.
Patent Information
- Application Number
- CN202211137162.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-19
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-09-19
AI Technical Summary
The prior art can only achieve local confusion in speech sample obfuscation, and cannot fully confuse massive speech samples, affecting the stability of model training.
By obtaining sample indexes of multiple voice data sources, blocking them according to the preset block size, then selecting the block index from the obfuscated block index, dividing it into multiple batch indexes, and then obfuscating the batch index, achieving sufficient and global obfuscation of massive voice samples.
Granular confusion of multiple voice data sources and batch voice samples is achieved, and massive voice samples are fully confused, which improves the stability of the model training process.
Smart Images

Figure CN115497464B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a speech sample obfuscation method and device. Background Art
[0002] In speech-related training tasks such as speech recognition, voice wake-up, and voiceprint recognition, the model needs to process massive speech samples for training. Massive speech samples are generally stored in multiple speech data sources. Sufficiently confusing the speech samples is conducive to stable convergence of the model.
[0003] The method of speech sample obfuscation in the prior art is: first, a partial speech sample of a speech data source is loaded, then the partial speech sample is divided into a plurality of batch speech samples, and finally, the plurality of batch speech samples are obfuscated for training.
[0004] However, after research, it was found that the speech sample confusion performed using the existing technology can only load part of the speech samples of a speech data source at a time to form multiple batches of speech sample confusions, and can only achieve local confusion of the speech samples. The insufficient confusion affects the stability of the model training process. Summary of the invention
[0005] In view of this, an embodiment of the present application provides a speech sample obfuscation method and device, aiming to achieve sufficient obfuscation of speech samples and improve the stability of the model training process.
[0006] In a first aspect, an embodiment of the present application provides a speech sample obfuscation method, the method comprising:
[0007] Acquire multiple sample indexes corresponding to multiple voice samples, where the multiple voice samples belong to multiple voice data sources;
[0008] Divide the multiple sample indexes into blocks according to a preset block size to obtain multiple block indexes, each block index including at least two sample indexes;
[0009] Obfuscating the multiple block indexes to obtain multiple obfuscated block indexes;
[0010] Selecting a preset number of block indexes from the multiple obfuscated block indexes to obtain multiple block indexes to be processed;
[0011] Divide the sample indexes included in the multiple block indexes to be processed into multiple batch indexes, each batch index includes at least two sample indexes, and the number of the multiple batch indexes is greater than the number of the multiple block indexes to be processed;
[0012] Obfuscating the multiple batch indexes to obtain multiple obfuscated batch indexes;
[0013] According to the sample indexes included in the multiple confused batch indexes, multiple speech samples to be processed are obtained.
[0014] Optionally, obtaining a plurality of sample indexes corresponding to a plurality of speech samples includes:
[0015] Acquire multiple description files corresponding to the multiple voice data sources;
[0016] According to the multiple description files, multiple sample indexes corresponding to the multiple voice samples are obtained, and the description files include the multiple sample indexes corresponding to the multiple voice samples.
[0017] Optionally, each description file also includes:
[0018] Each voice data source includes the number of voice samples and the sample frame lengths of the multiple voice samples corresponding to the multiple sample indexes.
[0019] Optionally, dividing the multiple sample indexes into blocks according to a preset block size to obtain multiple block indexes includes:
[0020] In the process of dividing the multiple sample indexes into blocks according to a preset block size, if the number of sample indexes remaining in the blocks in the voice data source is less than the preset block size, the sample indexes remaining in the blocks are separately divided into blocks to obtain the block index, or the sample indexes remaining in the blocks are repeated to supplement the preset block size to obtain the block index, or the sample indexes remaining in the blocks are discarded.
[0021] Optionally, the method further comprises:
[0022] Updating the preset quantity to obtain an updated preset quantity;
[0023] The step of selecting a preset number of block indexes from the obfuscated multiple block indexes to obtain multiple block indexes to be processed is specifically as follows:
[0024] An updated preset number of block indexes are selected from the multiple obfuscated block indexes to obtain the multiple block indexes to be processed.
[0025] Optionally, the method further comprises:
[0026] The multiple speech samples to be processed are input into the speech model for training.
[0027] Optionally, dividing the sample indexes included in the multiple to-be-processed block indexes into multiple batch indexes includes:
[0028] Obtaining a sample frame length of a speech sample corresponding to a sample index included in the plurality of block indexes to be processed;
[0029] Sorting the sample indexes included in the plurality of to-be-processed block indexes according to the sample frame length to obtain sorted sample indexes;
[0030] The sorted sample indexes are spliced into the multiple batch indexes according to a preset splicing method.
[0031] Optionally, the step of sorting the sample indexes included in the plurality of to-be-processed block indexes according to the sample frame length to obtain sorted sample indexes is specifically as follows:
[0032] The sample indexes included in the multiple to-be-processed block indexes are sorted in descending order or ascending order according to the sample frame length to obtain the sorted sample indexes.
[0033] In a second aspect, an embodiment of the present application provides a speech sample obfuscation device, the device comprising:
[0034] A first acquisition module, configured to acquire a plurality of sample indexes corresponding to a plurality of voice samples, wherein the plurality of voice samples belong to a plurality of voice data sources;
[0035] A block division module, used for dividing the multiple sample indexes into blocks according to a preset block size to obtain multiple block indexes, each block index includes at least two sample indexes;
[0036] A first obfuscation module, configured to obfuscate the plurality of block indexes to obtain a plurality of obfuscated block indexes;
[0037] A selection module, used for selecting a preset number of block indexes from the multiple obfuscated block indexes to obtain multiple block indexes to be processed;
[0038] A division module, used to divide the sample indexes included in the multiple block indexes to be processed into multiple batch indexes, each batch index includes at least two sample indexes, and the number of the multiple batch indexes is greater than the number of the multiple block indexes to be processed;
[0039] A second obfuscation module is used to obfuscate the multiple batch indexes to obtain multiple obfuscated batch indexes;
[0040] The second acquisition module is used to acquire multiple speech samples to be processed according to the sample indexes included in the multiple confused batch indexes.
[0041] In a third aspect, an embodiment of the present application provides a voice data obfuscation device, the device comprising:
[0042] Memory for storing computer programs;
[0043] The processor is used to execute the computer program so that the device performs the speech sample obfuscation method described in the first aspect.
[0044] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed, a device executing the computer program implements the speech sample confusion method described in the first aspect.
[0045] Compared with the prior art, the embodiments of the present application have the following beneficial effects:
[0046] The embodiment of the present application provides a method for confusing speech samples, which includes obtaining multiple sample indexes corresponding to multiple speech samples belonging to multiple speech data sources; dividing the multiple sample indexes into blocks according to a preset block size to obtain multiple block indexes; confusing the multiple block indexes to obtain multiple confused block indexes, and realizing speech sample confusion at the granularity of multiple speech data sources; selecting a preset number of block indexes from the multiple confused block indexes as multiple block indexes to be processed; dividing the sample indexes included in the multiple block indexes to be processed into multiple batch indexes; confusing the multiple batch indexes to obtain multiple confused batch indexes, and realizing speech sample confusion at the granularity of multiple batch speech samples; obtaining multiple speech samples to be processed through the sample indexes included in the multiple confused batch indexes. It can be seen that the method can not only realize speech sample confusion at the granularity of multiple speech data sources, but also realize speech sample confusion at the granularity of multiple batch speech samples, thereby realizing full and global confusion of massive speech samples, and subsequently inputting multiple speech samples to be processed into the speech model for training, which can improve the stability of the model training process. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solutions in this embodiment or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0048] Figure 1 An application scenario of a speech sample obfuscation method provided in an embodiment of the present application;
[0049] Figure 2 A flowchart of a speech sample obfuscation method provided in an embodiment of the present application;
[0050] Figure 3 A flow chart of a method for dividing a block index into multiple batch indexes provided in an embodiment of the present application;
[0051] Figure 4 A schematic diagram of a specific speech sample confusion provided in an embodiment of the present application;
[0052] Figure 5 A schematic diagram of a speech sample obfuscation device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0053] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0054] At present, the existing speech sample obfuscation method is: first, load part of the speech samples of a speech data source, then divide the part of the speech samples into multiple batches of speech samples, and finally confuse the multiple batches of speech samples for training. However, after research, it was found that the speech sample obfuscation performed by the existing technology can only load part of the speech samples of a speech data source at a time to form multiple batches of speech sample obfuscation, which can only achieve partial obfuscation of speech samples, and the insufficient obfuscation affects the stability of the model training process.
[0055] Based on this, in order to solve the above problems and improve the stability of the model training process, the embodiment of the present application provides a method and device for confusing speech samples, which obtains multiple sample indexes corresponding to multiple speech samples belonging to multiple speech data sources; divides the multiple sample indexes into blocks according to a preset block size to obtain multiple block indexes; confuses the multiple block indexes to obtain multiple confused block indexes, and realizes the confusion of speech samples at the granularity of multiple speech data sources; selects a preset number of block indexes from the multiple confused block indexes as multiple block indexes to be processed; divides the sample indexes included in the multiple block indexes to be processed into multiple batch indexes; confuses the multiple batch indexes to obtain multiple confused batch indexes, and realizes the confusion of speech samples at the granularity of multiple batch speech samples; obtains multiple speech samples to be processed through the sample indexes included in the multiple confused batch indexes. It can be seen that the method can not only realize the confusion of speech samples at the granularity of multiple speech data sources, but also realize the confusion of speech samples at the granularity of multiple batch speech samples, thereby realizing full and global confusion of massive speech samples, and subsequently inputting multiple speech samples to be processed into the speech model for training, which can improve the stability of the model training process.
[0056] For example, one of the scenarios of the embodiment of the present application may be applied to Figure 1 In the scenario shown, the scenario includes a database 101 and a server 102, wherein the database 101 includes multiple voice data sources, and the server 102 uses the implementation method provided in the embodiment of the present application to obtain multiple voice samples to be processed from the database 101.
[0057] First, in the above application scenario, although the action description of the implementation method provided in the embodiment of the present application is executed by the server 102; however, the embodiment of the present application is not limited in terms of the execution subject, as long as the actions disclosed in the implementation method provided in the embodiment of the present application are executed.
[0058] Secondly, the above scenario is only an example scenario provided by the embodiment of the present application, and the embodiment of the present application is not limited to this scenario.
[0059] The specific implementation of the speech sample obfuscation method and device in the embodiment of the present application is described in detail below with reference to the accompanying drawings.
[0060] See also Figure 2 , which is a flow chart of a speech sample obfuscation method provided in an embodiment of the present application, combined with Figure 2 As shown, it may specifically include:
[0061] S201: Acquire multiple sample indexes corresponding to multiple voice samples, where the multiple voice samples belong to multiple voice data sources.
[0062] Generally, there are multiple voice data sources, each of which includes multiple voice samples. Corresponding multiple sample indexes can be set in advance for multiple voice samples belonging to the multiple voice data sources, where the sample index is used to quickly find the corresponding voice sample, and subsequent steps are performed based on the obtained sample index. There is no need to read specific voice sample information, thereby reducing the occupied memory.
[0063] The embodiment of the present application may not specifically limit the process of obtaining multiple sample indexes. For ease of understanding, a possible implementation is described below.
[0064] In a possible implementation, on the basis of pre-setting corresponding multiple sample indexes for multiple voice samples belonging to multiple voice data sources, corresponding multiple description files can also be pre-set for the multiple voice data sources, and the description files corresponding to the voice data sources can be first obtained, and then the sample indexes of the multiple voice samples can be obtained according to the description files. Therefore, S201 can specifically include: obtaining multiple description files corresponding to multiple voice data sources; obtaining multiple sample indexes corresponding to multiple voice samples according to the multiple description files, and the description files include multiple sample indexes corresponding to the multiple voice samples. Among them, each description file also includes the number of voice samples included in each voice data source and the sample frame lengths of the multiple voice samples corresponding to the multiple sample indexes, and the sample frame lengths can be used to sort the voice samples.
[0065] For example, multiple voice data sources are recorded as File i , i∈{1,…,N F},NF Indicates the number of speech data sources; a speech data source contains multiple speech samples, and the multiple sample indexes corresponding to the multiple speech samples are recorded as S ij , i∈{1,…,N F}, j∈{1,…,N i}, i represents the i-th voice data source File i , j represents the jth speech sample of a speech data source, then S ij Indicates the i-th voice data source File i The jth speech sample, N i Indicates the i-th voice data source File i The number of speech samples included. i The corresponding description file is recorded as Desc i , description file Desc i The number of speech samples N included in the i-th speech data source i , the sample index S corresponding to the speech sample of the i-th speech data source ij , and the sample frame length of the speech sample of the i-th speech data source, etc., and the sample frame length is recorded as f ij Therefore, multiple sample indexes corresponding to multiple speech samples can be obtained according to the description file.
[0066] Among them, the description file Desc i The description format can include json format, specifically:
[0067]
[0068] The "source" item indicates the i-th voice data source File i The specific storage path of the "lengths" item stores an integer list, which is used to represent the i-th voice data source File i The sample frame length of the speech sample (the jth number represents the i-th speech data source File i The jth speech sample S ij The sample frame length f ij ), the "size" item indicates the i-th voice data source File i Of course, other methods may also be used without affecting the implementation of the embodiments of the present application.
[0069] S202: Divide the multiple sample indexes into blocks according to a preset block size to obtain multiple block indexes, each block index including at least two sample indexes.
[0070] According to the preset number of specific sample indexes included in each block, the obtained multiple sample indexes are divided into blocks to obtain multiple divided block indexes, each block index includes at least two sample indexes. The preset block size refers to the preset number of sample indexes included in each block index.
[0071] Among them, the embodiment of the present application may not specifically limit the process of dividing multiple sample indexes into blocks according to a preset block size. For ease of understanding, it is described below in conjunction with a possible implementation method.
[0072] Generally, the number of speech samples in a speech data source is uncertain. When the sample indexes corresponding to the speech samples are divided into blocks according to a preset block size, the number of sample indexes remaining after the division may be insufficient to be divided into an index block. At this time, the sample indexes remaining after the division may be processed as follows: the sample indexes remaining after the division may be divided into a separate block to obtain a block index, or the sample indexes remaining after repeated division may be supplemented to a preset block size to obtain a block index, or the sample indexes remaining after the division may be discarded.
[0073] Therefore, in a possible implementation, S202 may specifically include: in the process of dividing the multiple sample indexes into blocks according to the preset block size, if the number of sample indexes remaining in the block in the voice data source is less than the preset block size, the sample indexes remaining in the block are separately divided into blocks to obtain block indexes, or the sample indexes remaining in the repeated block are supplemented to the preset block size to obtain block indexes, or the sample indexes remaining in the block are discarded. Of course, other methods may also be used, which does not affect the implementation of the embodiments of the present application.
[0074] S203: Obfuscate the multiple block indexes to obtain multiple obfuscated block indexes.
[0075] The multiple divided block indexes are confused and their original order is disrupted, so that the block indexes included in the same voice data source are dispersed, and the voice sample confusion at the granularity of multiple voice data sources is realized, so that when some block indexes are selected later, the sample indexes of voice samples belonging to different voice data sources can be included, so that the subsequent voice samples can be confused more fully.
[0076] S204: Select a preset number of block indexes from the multiple obfuscated block indexes to obtain multiple block indexes to be processed.
[0077] The obtained multiple to-be-processed block indexes include sample indexes of speech samples belonging to different speech data sources.
[0078] In addition, it may be necessary to select block indexes from the multiple block indexes after confusion multiple times to confuse the speech samples multiple times, and the predetermined number of block indexes selected each time may change according to the number of speech samples actually required for speech model training. Therefore, before each block index is selected, the selected preset number can be updated. Therefore, in an optional embodiment of the present application, the method may also include S1: updating the preset number to obtain an updated preset number; accordingly, S204 may specifically be: selecting an updated preset number of block indexes from the multiple block indexes after confusion to obtain multiple block indexes to be processed. Among them, the number of block indexes selected each time can be updated to achieve subsequent dynamic control of batch index confusion. Each time a block index is selected from the multiple block indexes after confusion according to the updated preset number, dynamic control of the range of speech sample confusion numbers can be achieved to meet the actual needs of speech model training.
[0079] S205: Divide the sample indexes included in the multiple block indexes to be processed into multiple batch indexes, each batch index includes at least two sample indexes, and the number of the multiple batch indexes is greater than the number of the multiple block indexes to be processed.
[0080] The sample indexes included in the multiple block indexes to be processed are further divided into multiple batch indexes, so the number of the divided multiple batch indexes is greater than the number of the multiple block indexes to be processed. The batch index includes a batch of sample indexes, and a batch of sample indexes is at least two sample indexes.
[0081] The specific implementation method of dividing the block index into multiple batch indexes is not specifically limited in the present application embodiment. For ease of understanding, a possible implementation method is described below. For technical details, please refer to the introduction below.
[0082] S206: Obfuscate the multiple batch indexes to obtain multiple obfuscated batch indexes.
[0083] The multiple divided batch indexes are confused. On the basis of confusing the block indexes, both the speech sample confusion at the granularity of multiple speech data sources and the speech sample confusion at the granularity of multiple batch speech samples are realized, thereby achieving full and global confusion of massive speech samples.
[0084] S207: Acquire multiple speech samples to be processed according to the sample indexes included in the multiple batch indexes after confusion.
[0085] According to the sample indexes included in the multiple batch indexes after obfuscation, multiple voice samples can be obtained as multiple voice samples to be processed for model training. In the voice sample obfuscation process, sample indexes, block indexes and batch indexes are used, and the reading of specific voice sample information is not involved. Therefore, the memory space occupied is small, which can easily meet the obfuscation needs of massive data.
[0086] In addition, in an optional embodiment of the present application, the method may further include S2: inputting a plurality of speech samples to be processed into a speech model for training. In which, the plurality of speech samples to be processed that have been fully and globally confused are input into the speech model for training, thereby improving the stability of the model training process.
[0087] Based on the above S201-S207, it can be known that in the embodiment of the present application, multiple sample indexes corresponding to multiple speech samples belonging to multiple speech data sources are obtained; multiple sample indexes are divided into blocks according to a preset block size to obtain multiple block indexes; multiple block indexes are confused to obtain multiple confused block indexes, so as to achieve speech sample confusion at the granularity of multiple speech data sources; a preset number of block indexes are selected from the multiple confused block indexes as multiple block indexes to be processed; sample indexes included in the multiple block indexes to be processed are divided into multiple batch indexes; multiple batch indexes are confused to obtain multiple confused batch indexes, so as to achieve speech sample confusion at the granularity of multiple batch speech samples; multiple speech samples to be processed are obtained through the sample indexes included in the multiple confused batch indexes. It can be seen that the method can not only achieve speech sample confusion at the granularity of multiple speech data sources, but also achieve speech sample confusion at the granularity of multiple batch speech samples, thereby achieving full and global confusion of massive speech samples, and subsequently inputting multiple speech samples to be processed into the speech model for training, which can improve the stability of the model training process.
[0088] See also Figure 3 , which is a flow chart of a method for dividing a block index into multiple batch indexes provided by an embodiment of the present application, combined with Figure 3 As shown, it may specifically include:
[0089] S301: Obtain sample frame lengths of speech samples corresponding to sample indexes included in a plurality of to-be-processed block indexes.
[0090] According to the sample index in the description file corresponding to the voice data source to which the voice sample belongs, the corresponding sample frame length is obtained to sort the sample index.
[0091] S302: Sort the sample indexes included in the plurality of to-be-processed block indexes according to the sample frame length to obtain sorted sample indexes.
[0092] According to the size of the sample frame length, the sample indexes included in the multiple block indexes to be processed are sorted according to a preset sorting method. There may be many sorting methods. For example, it may include sorting the sample indexes included in the multiple block indexes to be processed in descending order or ascending order according to the sample frame length to obtain the sorted sample indexes.
[0093] S303: splicing the sorted sample indexes into multiple batch indexes according to a preset splicing method.
[0094] The preset splicing method refers to a pre-set method for splicing the sorted sample indexes. For example, the sum of the sample frame lengths corresponding to the multiple sample indexes contained in each batch index can be set in advance to be a fixed value, or the number of multiple sample indexes contained in each batch index is the same. Of course, other splicing methods can also be used, which does not affect the implementation of the embodiments of the present application.
[0095] The model can usually only process batches of speech samples of equal length. When the sample frame lengths of the same batch of speech samples differ greatly, zero padding is required. At this time, no matter what concatenation method is used, the sample frame lengths of the multiple sample indexes included in the divided batch index are arranged in descending or ascending order, and the difference in their sample frame lengths is small. The multiple speech samples to be processed obtained according to the sample indexes included in the multiple batch indexes after confusion are input into the speech model for training, which reduces the zero padding operation.
[0096] Based on the above S301-S303 related contents, it can be known that in the embodiment of the present application, the sample frame lengths of the multiple sample indexes included in the batch indexes divided in this way are relatively small. Before the multiple speech samples to be processed are obtained according to the sample indexes included in the multiple batch indexes after confusion and input into the speech model, the zero-padding operation is reduced, the training efficiency is further improved, and the model converges stably.
[0097] For ease of understanding, the voice sample obfuscation method provided in the embodiment of the present application is described in detail below. Figure 4 , which is a schematic diagram of a specific speech sample confusion provided in an embodiment of the present application. In this embodiment, there are a total of 3 speech data sources, File 1 Contains 20,000 voice samples, 20,000 voice samples correspond to 20,000 sample indexes, File 2 Contains 37695 sample indexes, 37695 voice samples correspond to 37695 sample indexes, File 3 Contains 19228 sample indexes, and 19228 voice samples correspond to 19228 sample indexes. The specific voice sample obfuscation steps may include:
[0098] S401: According to the three description files corresponding to the three voice data sources, 76923 sample indexes corresponding to the 76923 voice samples are obtained. 1 Corresponding description file Desc 1 Contains 20,000 sample indexes corresponding to 20,000 speech samples, File 2 Corresponding description file Desc 2 Including 37695 sample indexes corresponding to 37695 speech samples, File 3 Corresponding description file Desc 3 Includes 19228 sample indexes corresponding to 19228 speech samples.
[0099] S402: Divide the 76923 sample indexes corresponding to the 76923 speech samples belonging to the three speech data sources into blocks containing 10,000 sample indexes, wherein File 1 The 20,000 sample indexes are divided into two blocks, and two block indexes are obtained. 2 The 30,000 sample indexes are divided into 3 blocks, and the remaining 7,695 sample indexes can be grouped into a single block, resulting in 4 block indexes. 3 The 10,000 sample indexes are divided into 1 block, and the remaining 9,228 sample indexes can be divided into a single block to obtain 2 block indexes, and finally a total of 9 block indexes are obtained.
[0100] S403: 9 block indexes are obfuscated and their order is disrupted so that the block indexes belonging to the same voice data source are dispersed to obtain 9 obfuscated block indexes.
[0101] S404: Select 3 block indexes from the 9 obfuscated block indexes to obtain 3 to-be-processed block indexes respectively belonging to different voice data sources.
[0102] S405: According to the three description files corresponding to the three speech data sources to which the three block indexes to be processed belong, the sample frame lengths of the speech samples corresponding to the sample indexes included in the three block indexes to be processed are obtained.
[0103] S406: Sort the sample indexes included in the three to-be-processed block indexes in ascending order according to the sample frame lengths to obtain sorted sample indexes.
[0104] S407: In a manner in which the sum of sample frame lengths of speech samples corresponding to the sample indexes included in each batch index is a fixed frame length, the sorted sample indexes are spliced into five batch indexes.
[0105] S408: Obfuscate the five batch indexes to obtain five obfuscated batch indexes;
[0106] S409: Acquire multiple speech samples to be processed according to the sample indexes included in the five confused batch indexes.
[0107] S410: Input a plurality of speech samples to be processed into a speech model for training.
[0108] Based on the relevant contents of S401-S410 above, it can be known that in the embodiment of the present application, it is possible to achieve both speech sample confusion at the granularity of multiple speech data sources and speech sample confusion at the granularity of multiple batch speech samples, thereby achieving full and global confusion of massive speech samples. Subsequently, multiple speech samples to be processed are input into the speech model for training, thereby improving the stability of the model training process.
[0109] The above are some specific implementations of the speech sample obfuscation method provided in the embodiment of the present application. Based on this, the present application also provides a corresponding device. The device provided in the embodiment of the present application will be introduced from the perspective of functional modularization.
[0110] See also Figure 5 , which is a schematic diagram of a speech sample obfuscation device provided in an embodiment of the present application, and the device may include:
[0111] A first acquisition module 501 is used to acquire multiple sample indexes corresponding to multiple voice samples, where the multiple voice samples belong to multiple voice data sources;
[0112] A block division module 502, configured to divide the plurality of sample indexes into blocks according to a preset block size to obtain a plurality of block indexes, each block index including at least two sample indexes;
[0113] A first obfuscation module 503, configured to obfuscate multiple block indexes to obtain multiple obfuscated block indexes;
[0114] A selection module 504 is used to select a preset number of block indexes from the multiple obfuscated block indexes to obtain multiple block indexes to be processed;
[0115] A division module 505, configured to divide the sample indexes included in the plurality of block indexes to be processed into a plurality of batch indexes, each batch index including at least two sample indexes, and the number of the plurality of batch indexes is greater than the number of the plurality of block indexes to be processed;
[0116] A second obfuscation module 506 is used to obfuscate the multiple batch indexes to obtain the multiple batch indexes after obfuscation;
[0117] The second acquisition module 507 is used to acquire a plurality of speech samples to be processed according to the sample indexes included in the plurality of batch indexes after confusion.
[0118] In the embodiment of the present application, the speech samples are confused by the cooperation of the first acquisition module 501, the blocking module 502, the first confusion module 503, the selection module 504, the division module 505, the second confusion module 506, and the second acquisition module 507. It can realize the confusion of speech samples at the granularity of multiple speech data sources and the confusion of speech samples at the granularity of multiple batch speech samples, thereby realizing full and global confusion of massive speech samples. Subsequently, multiple speech samples to be processed are input into the speech model for training, which improves the stability of the model training process.
[0119] As an implementation manner, the first acquisition module 501 may specifically include:
[0120] A first acquisition unit, used to acquire multiple description files corresponding to multiple voice data sources;
[0121] The second acquisition unit is used to acquire a plurality of sample indexes corresponding to a plurality of speech samples according to a plurality of description files, wherein the description files include a plurality of sample indexes corresponding to the plurality of speech samples.
[0122] As an implementation manner, each description file of the first acquisition unit further includes the number of speech samples included in each speech data source and the sample frame lengths of the multiple speech samples corresponding to the multiple sample indexes.
[0123] As an implementation manner, the block module 502 may be specifically used for:
[0124] In the process of dividing multiple sample indexes into blocks according to a preset block size, if the number of sample indexes remaining in the speech data source is less than the preset block size, the sample indexes remaining in the blocks are separately divided into blocks to obtain block indexes, or the sample indexes remaining in repeated blocks are used to supplement the preset block size to obtain block indexes, or the sample indexes remaining in the blocks are discarded.
[0125] As an implementation manner, the speech sample obfuscation device may further include:
[0126] The updating module is used to update the preset quantity and obtain the updated preset quantity.
[0127] Accordingly, the selection module 504 may be specifically used for:
[0128] An updated preset number of block indexes are selected from the multiple obfuscated block indexes to obtain multiple block indexes to be processed.
[0129] As an implementation manner, the speech sample obfuscation device may further include:
[0130] The input module is used to input multiple speech samples to be processed into the speech model for training.
[0131] As an implementation manner, the division module 505 may specifically include:
[0132] A third acquisition unit, used to acquire a sample frame length of a speech sample corresponding to a sample index included in a plurality of block indexes to be processed;
[0133] A sorting unit, used to sort the sample indexes included in the plurality of to-be-processed block indexes according to the sample frame length to obtain sorted sample indexes;
[0134] The splicing unit is used to splice the sorted sample indexes into multiple batch indexes according to a preset splicing method.
[0135] As an implementation mode, the sorting unit may be specifically used for:
[0136] The sample indexes included in the plurality of to-be-processed block indexes are sorted in descending order or ascending order according to the sample frame length to obtain sorted sample indexes.
[0137] The embodiments of the present application also provide corresponding devices and computer storage media for implementing the solutions provided by the embodiments of the present application.
[0138] The device includes a memory and a processor, the memory is used to store a computer program, and the processor is used to execute the computer program, so that the device executes the speech sample confusion method described in any embodiment of the present application.
[0139] The computer storage medium stores codes, and when the codes are executed, the device executing the codes implements the speech sample obfuscation method described in any embodiment of the present application.
[0140] The "first" and "second" in the names such as "first" and "second" (if any) mentioned in the embodiments of the present application are only used as name identifiers and do not represent the first or second in order.
[0141] Through the description of the above implementation methods, it can be known that those skilled in the art can clearly understand that all or part of the steps in the above-mentioned embodiment method can be implemented by means of software plus a general hardware platform. Based on such an understanding, the technical solution of the present application can be embodied in the form of a software product, which can be stored in a storage medium, such as a read-only memory (ROM) / RAM, a magnetic disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network communication device such as a router) to execute the methods described in each embodiment of the present application or some parts of the embodiments.
[0142] It should be noted that each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described above is merely schematic, in which the unit described as a separate component may or may not be physically separated, and the component prompted as a unit may or may not be a physical unit, that is, it may be located in one place, or it may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative work.
[0143] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A speech sample obfuscation method, It is characterized in that The method comprises: Acquire multiple sample indexes corresponding to multiple voice samples, where the multiple voice samples belong to multiple voice data sources; Divide the multiple sample indexes into blocks according to a preset block size to obtain multiple block indexes, each block index including at least two sample indexes; Obfuscating the multiple block indexes to obtain multiple obfuscated block indexes; Selecting a preset number of block indexes from the multiple obfuscated block indexes to obtain multiple block indexes to be processed; Divide the sample indexes included in the multiple block indexes to be processed into multiple batch indexes, each batch index includes at least two sample indexes, and the number of the multiple batch indexes is greater than the number of the multiple block indexes to be processed; Obfuscating the multiple batch indexes to obtain multiple obfuscated batch indexes; According to the sample indexes included in the multiple confused batch indexes, multiple speech samples to be processed are obtained.
2. The method according to claim 1, It is characterized in that The obtaining of multiple sample indexes corresponding to the multiple voice samples includes: Acquire multiple description files corresponding to the multiple voice data sources; According to the multiple description files, multiple sample indexes corresponding to the multiple voice samples are obtained, and the description files include the multiple sample indexes corresponding to the multiple voice samples.
3. The method according to claim 2, It is characterized in that Each description file also includes the number of speech samples included in each speech data source and the sample frame lengths of the multiple speech samples corresponding to the multiple sample indexes.
4. The method according to claim 1, It is characterized in that The step of dividing the plurality of sample indexes into blocks according to a preset block size to obtain a plurality of block indexes includes: In the process of dividing the multiple sample indexes into blocks according to a preset block size, if the number of sample indexes remaining in the blocks in the voice data source is less than the preset block size, the sample indexes remaining in the blocks are separately divided into blocks to obtain the block index, or the sample indexes remaining in the blocks are repeated to supplement the preset block size to obtain the block index, or the sample indexes remaining in the blocks are discarded.
5. The method according to claim 1, It is characterized in that The method further comprises: Updating the preset quantity to obtain an updated preset quantity; The step of selecting a preset number of block indexes from the obfuscated multiple block indexes to obtain multiple block indexes to be processed is specifically as follows: An updated preset number of block indexes are selected from the multiple obfuscated block indexes to obtain the multiple block indexes to be processed.
6. The method according to any one of claims 1 to 5, It is characterized in that The method further comprises: The multiple speech samples to be processed are input into the speech model for training.
7. The method according to claim 1, It is characterized in that The step of dividing the sample indexes included in the plurality of to-be-processed block indexes into a plurality of batch indexes comprises: Obtaining a sample frame length of a speech sample corresponding to a sample index included in the plurality of block indexes to be processed; Sorting the sample indexes included in the plurality of to-be-processed block indexes according to the sample frame length to obtain sorted sample indexes; The sorted sample indexes are spliced into the multiple batch indexes according to a preset splicing method.
8. The method according to claim 7, It is characterized in that The step of sorting the sample indexes included in the plurality of to-be-processed block indexes according to the sample frame length to obtain the sorted sample indexes is specifically as follows: The sample indexes included in the multiple to-be-processed block indexes are sorted in descending order or ascending order according to the sample frame length to obtain the sorted sample indexes.
9. A speech sample obfuscation device, It is characterized in that The device comprises: A first acquisition module, configured to acquire a plurality of sample indexes corresponding to a plurality of voice samples, wherein the plurality of voice samples belong to a plurality of voice data sources; A block division module, used for dividing the multiple sample indexes into blocks according to a preset block size to obtain multiple block indexes, each block index includes at least two sample indexes; A first obfuscation module, configured to obfuscate the plurality of block indexes to obtain a plurality of obfuscated block indexes; A selection module, used for selecting a preset number of block indexes from the multiple obfuscated block indexes to obtain multiple block indexes to be processed; A division module, used to divide the sample indexes included in the multiple block indexes to be processed into multiple batch indexes, each batch index includes at least two sample indexes, and the number of the multiple batch indexes is greater than the number of the multiple block indexes to be processed; A second obfuscation module is used to obfuscate the multiple batch indexes to obtain multiple obfuscated batch indexes; The second acquisition module is used to acquire multiple speech samples to be processed according to the sample indexes included in the multiple confused batch indexes.
10. A speech sample obfuscation device, It is characterized in that The device comprises: Memory for storing computer programs; A processor is used to execute the computer program so that the device performs the speech sample obfuscation method according to any one of claims 1 to 8.