Speech dataset screening and processing method, screening and processing device, and storage medium
By performing time-frequency feature analysis and multiple label prediction processing on the speech data set, combined with the screening method of predicting error-free times threshold, the problem of low data preprocessing accuracy in the prior art is solved, more efficient and accurate cleaning of dirty data is achieved, and the accuracy of the speech keyword detection model is improved.
Patent Information
- Application Number
- CN202211730524.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-30
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-12-30
AI Technical Summary
In the training process of the speech keyword detection model, the prior art relies on the detection effect of the keyword detection model, resulting in low accuracy of data preprocessing, which may lead to the correct speech data of the label being cleaned and data information loss.
By performing time-frequency feature analysis on the speech data set, multiple label prediction processing is performed using the keyword detection model, average prediction error sequence is calculated, and the prediction error error threshold is determined based on the sequence, and the speech data is misaligned to be screened to ensure the quality of the clean speech data set.
It improves the accuracy and robustness of dirty data cleaning, avoids the clean voice data being erroneously cleaned, improves the data cleaning efficiency of the voice keyword data set, and helps to build a more accurate keyword detection model.
Smart Images

Figure CN115966201B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data communication technologies, and more specifically, to a method and apparatus for screening and processing a voice data set, and a storage medium. Background Art
[0002] With the development of technologies, as an important basis for realizing human-machine voice interaction and intelligent device control, the voice keyword detection technology has been widely applied in fields such as intelligent vehicles, smart homes, and robots. This technology mainly trains an offline recognition model based on the supervised learning method in machine learning and applies it to the online recognition process. However, when the keyword recognition technology is in the offline R & D process, the effectiveness of the sample labels in the data set often affects the effect of the final recognition model in online keyword recognition. Specifically, if the data set label is incorrect, that is, there is dirty data in the training data, the model will incorrectly learn the feature information of the training data.
[0003] To solve the above problems, the industry uses the method of inputting the keyword voice data to be preprocessed into a trained keyword detection model, and determines whether it is low-quality data according to whether the keyword classification result output by the model matches the actual label, and performs a preprocessing operation on the training data set according to the judgment result. However, the method of screening based on the keyword detection model still has some problems: the effect of data preprocessing completely depends on the detection effect of the keyword detection model. If the accuracy of the keyword detection model is low, some voice data with correct labels will be cleaned, resulting in a certain degree of loss of data information, which is not conducive to the subsequent construction of an accurate voice keyword detection model.
[0004] Content of the Application
[0005] The purpose of the embodiments of this application is to provide a method and apparatus for screening and processing a voice data set, and a storage medium, which can effectively improve the cleaning efficiency, cleaning accuracy, and robustness of dirty data, and can avoid the cleaning of clean voice data, so as to facilitate the subsequent construction of an accurate keyword detection model.
[0006] According to the first aspect of the present application, a method for screening and processing a speech data set is provided. The screened and processed speech data set is used for training a keyword detection model, and includes the following steps. By a processor: obtaining an original speech data set to be screened and processed, where each piece of speech data includes speech signal data and its keyword label; determining a valid speech data set based on the original speech data set; calculating time-frequency features for each piece of speech data in the valid speech data set; for each piece of speech data, based on the time-frequency features, using the keyword detection model to perform multiple label prediction processes to determine a sequence of prediction misalignment times for each time. Each element of the sequence of prediction misalignment times is arranged in the order of speech data and represents the prediction misalignment times of the corresponding speech data at that time; averaging the sequences of prediction misalignment times for multiple times to obtain an average sequence of prediction misalignment times. Each element of the average sequence of prediction misalignment times is arranged in the order of speech data and represents the average prediction misalignment times of the corresponding speech data for multiple times; determining a prediction misalignment times threshold based on the average sequence of prediction misalignment times; performing the following misalignment screening process for each piece of speech data to obtain a clean speech data set for training the keyword detection model: obtaining the average prediction misalignment times of the piece of speech data in the average sequence of prediction misalignment times, comparing it with the prediction misalignment times threshold, if it is greater than the latter, it is determined as dirty speech data and deleted, otherwise it is retained and stored in the clean speech data set.
[0007] According to the second aspect of the present application, a device for screening and processing a speech data set is provided, including: an interface configured to obtain an original speech data set to be screened and processed, where each piece of speech data includes speech signal data and its keyword label; and a processor configured to execute the method for screening and processing a speech data set according to various embodiments of the present application.
[0008] According to the third aspect of the present application, a non-transitory computer storage medium is provided, on which executable instructions are stored. When executed by a processor, the method for screening and processing a speech data set according to various embodiments of the present application is implemented.
[0009] The beneficial effects of the embodiments of the present application are as follows:
[0010] The screening and processing method provided by the embodiments of the present application performs multiple label prediction processes on each piece of speech data based on time-frequency features using a keyword detection model, and averages the multiple prediction misalignment times sequences to obtain the average prediction misalignment times. The average prediction misalignment times can accurately, objectively and truly reflect the misalignment situation after multiple prediction misalignment analyses of each piece of speech data. Screening for dirty data based on the average prediction misalignment times has higher reliability. In addition, a prediction misalignment times threshold is determined based on the average prediction misalignment times sequence, and the prediction misalignment times threshold can effectively reflect the reasonable threshold situation where clean speech data is located. By comparing the size of the average prediction misalignment times and the prediction misalignment times threshold, when the average prediction misalignment times is greater than the prediction misalignment times threshold, it is determined that this piece of data is dirty speech data. In this way, problems such as screening errors caused by relying on the keyword detection model can be avoided, which is beneficial to ensuring the accuracy of dirty data cleaning, minimizing the loss of clean speech data, and greatly improving the efficiency of speech keyword dataset data cleaning. At the same time, it helps to improve the accuracy of keyword recognition modeling based on the cleaned speech dataset subsequently. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In the drawings which are not necessarily drawn to scale, the same reference numerals may describe similar components in different views. The same reference numerals with alphabetical suffixes or different alphabetical suffixes may represent different instances of similar components. The drawings generally illustrate various embodiments by way of example and not limitation, and are used together with the description and the claims to explain the disclosed embodiments. Where appropriate, the same reference numerals are used throughout the drawings to refer to the same or similar parts. Such embodiments are illustrative and are not intended to be an exhaustive or exclusive embodiment of the apparatus or method.
[0012] Figure 1 The flowchart showing the screening and processing method of the speech dataset according to the embodiments of the present application is shown.
[0013] FIG. 2(a) shows the flowchart of performing each label prediction process on each piece of speech data according to the embodiments of the present application.
[0014] FIG. 2(b) shows the method flowchart for determining the prediction misalignment times increment according to the embodiments of the present application.
[0015] Figure 3 The screening flowchart of the speech dataset according to a specific embodiment of the present application is shown.
[0016] Figure 4 The schematic diagram of the screening and processing device of the speech dataset according to the embodiments of the present application is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] To enable those skilled in the art to better understand the technical solutions of the present application, the present application will be described in detail below in conjunction with the accompanying drawings and specific embodiments. The embodiments of the present application will be further described in detail below in conjunction with the accompanying drawings and specific examples, but this is not a limitation on the present application.
[0018] The "first", "second" and similar terms used in the present application do not denote any order, quantity or importance, but are only used for distinction. Terms such as "including" or "comprising" mean that the elements before this term cover the elements listed after this term, and do not exclude the possibility of also covering other elements. In the present application, the arrows shown in the figures for each step are only examples of the execution order and do not limit. The technical solution of the present application is not limited to the execution order described in the embodiments. Each step in the execution order can be executed jointly, can be decomposed, and can be reordered, as long as the logical relationship of the execution content is not affected.
[0019] All terms used in the present application (including technical terms or scientific terms) have the same meaning as understood by those of ordinary skill in the art to which the present application pertains, unless otherwise specifically defined. It should also be understood that terms defined in a general dictionary, for example, should be interpreted as having a meaning consistent with their meaning in the context of the relevant art, and should not be interpreted in an idealized or overly formal sense, unless specifically defined as such here. Technologies, methods and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the said technologies, methods and devices should be regarded as part of the specification.
[0020] Figure 1 A flowchart showing the screening processing method of the speech data set according to an embodiment of the present application. The screened speech data set is used for the training of a keyword detection model, as Figure 1 shown, the method includes the following steps:
[0021] In step S101, a processor obtains an original speech data set to be screened. Each piece of speech data includes speech signal data and its keyword label. Each piece of speech data in the original speech data set to be screened has its own keyword label. The data in this original speech data set may include dirty speech data with incorrect labels or invalid features, and needs to be screened to delete the dirty speech data and retain the clean speech data. In some embodiments, the keyword speech includes instruction speech issued by a user to a portable intelligent device. For example, instructions such as "Please play music!" issued by the user to the portable intelligent device by voice.
[0022] In step S102, based on the original speech dataset, an effective speech dataset is determined. Specifically, invalid speech data (such as blank speech data or speech data with particularly strong noise) in the original speech dataset can be removed, so as to retain the effective speech dataset. Among them, the original speech dataset may also not be processed, but directly used as the effective speech dataset. In addition, other embodiments for determining the effective speech dataset are not excluded.
[0023] In step S103, time-frequency features are calculated for each piece of speech data in the effective speech dataset. Among them, the time-frequency features may include one of MFCC features, Fbank features (features based on Mel filter banks) and their variants, and Mel spectrograms. Correspondingly, methods already disclosed in the prior art can be used to calculate their time-frequency features. MFCC is the Mel cepstral coefficient. The extraction process of MFCC features includes steps such as pre-emphasis filtering, framing, windowing, FFT, filtering by Mel filter banks, logarithmic operation, discrete cosine transform (DCT), and dynamic feature (differential feature) extraction. The Mel filter bank simulates the auditory mechanism of the cochlear hair acoustic sensors of the human body, has high low-frequency resolution and low high-frequency resolution, and the corresponding relationship with linear frequency is approximately logarithmic. Details are not described here.
[0024] In step S104, for each piece of speech data, based on the time-frequency features, multiple label prediction processes are performed using a keyword detection model to determine the sequence of prediction misalignment times for each time. Each element of this sequence of prediction misalignment times is arranged in the order of speech data and represents the prediction misalignment time of the corresponding speech data for that time. Specifically, for example, 4 label prediction processes are performed on the 1st, 2nd... Nth pieces of speech data. Among them, when the 1st, 2nd... Nth pieces of speech data are subjected to the 1st label prediction process, the sequence of prediction misalignment times for the 1st time is obtained. The 1st element in this sequence of prediction misalignment times corresponds to the prediction misalignment time of the 1st piece of speech data, the 2nd element corresponds to the prediction misalignment time of the 2nd piece of speech data,... the Nth element corresponds to the prediction misalignment time of the Nth piece of speech data. That is to say, the elements in the sequence of prediction misalignment times for each time are arranged in the order of speech data, and correspond to the prediction misalignment time of the corresponding speech data for that time. In this specific embodiment, since 4 label prediction processes are performed, the sequences of prediction misalignment times corresponding to the 4 label prediction processes can be obtained.
[0025] In step S105, the average of the multiple prediction misalignment number sequences is calculated to obtain an average prediction misalignment number sequence. Each element of the average prediction misalignment number sequence is arranged in the order of the voice data and represents the average prediction misalignment number of the corresponding voice data for multiple times. As described in the above embodiment, the average of the 4 prediction misalignment number sequences can be calculated to obtain the average prediction misalignment number sequence. For example, if the first voice data is subjected to 4 label prediction processes and the prediction result shows a misalignment of 2 times each time, then the first element in the average prediction misalignment number sequence is 2, that is, the average prediction misalignment number of the first voice data is 2. By analogy, the average prediction misalignment number of each voice data for multiple times can be obtained. The average prediction misalignment number can accurately, objectively and truly reflect the misalignment situation after multiple prediction misalignment analyses of each voice data, and can avoid the adverse effects caused by inaccurate single prediction results. Screening for dirty data based on the average prediction misalignment number has higher reliability and robustness.
[0026] In step S106, based on the average prediction misalignment number sequence, a prediction misalignment number threshold is determined. The specific value of the prediction misalignment number threshold is not limited and can be set by itself based on statistical data. For example, for a voice data set where most of the voice data is clean data and the proportion of dirty data is low, through statistical calculation, a threshold can be obtained that can efficiently distinguish between the misalignment conditions of dirty data and clean data, and this threshold has both robustness and sensitivity. The prediction misalignment number threshold can generally reflect the overall situation of the voice data set that meets the psychological expectations. This is only an illustrative example, and other methods for determining the prediction misalignment number threshold are not excluded.
[0027] In step S107, the following misalignment screening process is performed on each voice data to obtain a clean voice data set for training the keyword detection model: Obtain the average prediction misalignment number of this voice data in the average prediction misalignment number sequence, and compare it with the prediction misalignment number threshold. If it is greater than the latter, it is determined as dirty voice data and deleted; otherwise, it is retained and stored in the clean voice data set. As described above, the average prediction misalignment number of each voice data can be obtained based on the average prediction misalignment number sequence, and the average prediction misalignment number can overall reflect the misalignment situation of this voice data and has high robustness. Moreover, the prediction misalignment number threshold can also better reflect the misalignment situation of a better voice data set. When the average prediction misalignment number of this voice data is greater than the prediction misalignment number threshold, it is judged that the stability of this voice data is poor and the misalignment situation is relatively serious, then this voice data can be judged as dirty voice data and deleted. If the average prediction misalignment number of this voice data is less than or equal to the prediction misalignment number threshold, it means that the misalignment situation of this voice data meets the expectations and can be retained as clean voice data in the voice data set.
[0028] Compared with simply determining whether a piece of speech data is dirty speech data by judging whether the output result of the keyword detection model matches the true label, the method provided by the embodiments of the present application has higher reliability, accuracy, and efficiency in screening dirty speech data, can avoid cleaning up speech data with correct labels, greatly improves the efficiency of data cleaning for the speech keyword data set, and at the same time helps to improve the accuracy of keyword recognition modeling based on the cleaned speech data set.
[0029] In some embodiments of the present application, for each piece of speech data, performing multiple label prediction processes based on time-frequency features specifically includes performing each label prediction process on each piece of speech data through the steps shown in FIG. 2(a). In step S201, initialize the parameters of the keyword detection model and the sequence of prediction misalignment times for this time. The keyword detection model is constructed using a learning network. Among them, the learning network includes an LSTM learning network or a GRU neural network, and the present application does not make specific limitations in this regard. Usually, in the case of limited computing power (single-core) and storage space such as small chips, a 2-4 layer GRU neural network can be used to save computing power and storage space. By using an LSTM neural network, while considering the interaction between adjacent points of time-frequency speech features in the time domain and frequency domain, the influence of points that are far away in the time domain and frequency domain can be forgotten.
[0030] In step S202, perform the training and parameter tuning of the keyword detection model step by step, and determine the increment of the prediction misalignment times for each step of this piece of speech data. For each label prediction process performed on each piece of speech data, the calculation of the increment of the prediction misalignment times for multiple steps needs to be performed. And the specific steps for determining the increment of the prediction misalignment times for each step are shown in FIG. 2(b). In step S204, extract the time-frequency features of a group of speech data in the effective speech data set. Among them, the group of speech data is randomly extracted from the effective speech data set and is used to perform the training of the keyword detection model. By randomly extracting multiple pieces of speech data to form a group of speech data for the training of the keyword detection model, the risk of overfitting of the keyword detection model can be reduced, and the misalignment situation of real speech data can be better reflected.
[0031] In step S205, based on the time-frequency features of a set of extracted speech data and the labels of their keywords, the backpropagation algorithm is executed to adjust the parameters of the keyword detection model, thereby obtaining the keyword detection model with adjusted parameters. The principle of the backpropagation algorithm is to use the chain rule of differentiation to calculate the partial derivatives of the loss function between the actual output result and the ideal result with respect to each weight parameter or bias term, and then update the weights or bias terms layer by layer in reverse according to the optimization algorithm. It adopts a forward-backward propagation training method. By continuously adjusting the parameters in the model, the loss function is made to converge, thereby constructing an accurate model structure. The keyword detection model with adjusted parameters can be understood as having been trained using the time-frequency features of a randomly extracted set of speech data and the labels of their keywords. During the training process of the keyword detection model in each step, only a randomly extracted set of speech data is used for training, rather than all the speech data in all valid speech data sets. This greatly reduces the training load of the keyword detection model and makes the training process easier and more efficient.
[0032] Further, in some embodiments, the training and parameter adjustment of the keyword detection model in each step are both executed based on the initial parameters of the keyword detection model and the initialized prediction misalignment count sequence. By performing multiple initializations for training, the influence of randomly initializing network parameters on the elimination of invalid samples is eliminated, thereby enabling efficient and accurate preprocessing of training data.
[0033] In step S206, based on the time-frequency features of this piece of speech data, the keyword detection model with adjusted parameters is used to predict the label. After obtaining the keyword detection model with adjusted parameters, the label can be predicted based on the time-frequency features of each piece of speech data. It should be noted that step 206 performs the prediction of the label for each piece of speech data in each step of each label prediction process. For example, if each label prediction process includes five steps, then a set of speech data needs to be randomly extracted and the keyword detection model needs to be trained in each step to obtain the keyword detection model with adjusted parameters for each step. Then, the keyword detection model with adjusted parameters for each step is used to predict the label for each piece of speech data.
[0034] In step S207, the predicted label is compared with the keyword label of the corresponding voice data to determine whether the label prediction in this step is accurate. The judgment method is as in step S208. If the label prediction in this step is incorrect and the label prediction in the previous step is correct, the prediction misalignment count increment is 1; otherwise, the prediction misalignment count increment is 0. Specifically, for example, each label prediction process includes five steps. If the predicted label of this voice data in the first step is correct through comparison with the keyword label of this voice data, and the label prediction of this voice data in the second step is incorrect, it can be understood that the output result of the keyword detection model trained using a set of voice data extracted in the second step deteriorates, indicating that this voice data has a probability of being unstable and misaligned. Then, the prediction misalignment count increment is set to 1. And so on. If the label prediction in the third step is incorrect, the prediction misalignment count increment is 0. If the label prediction in the fourth step is correct, the prediction misalignment count increment is still 0. And if the label prediction in the fifth step is incorrect (while the label prediction in the fourth step is correct), the prediction misalignment count increment is incremented by 1 again. Then, the cumulative prediction misalignment count increment of these 5 steps of this voice data in this label prediction process is 2 (this embodiment is only for illustrative explanation), that is, the prediction misalignment count of this voice data is 2. As shown in step S203, by accumulating the prediction misalignment count increments of each step for this voice data, the prediction misalignment count of this voice data in the prediction misalignment count sequence of this time is obtained.
[0035] Further, taking Figure 3 the screening of the voice data set of a specific embodiment shown as an example for specific illustration. First, record the original voice data set as X train , which includes the clean voice data set X train_clean , the voice data set X train_aug generated by data augmentation, including multi-noise scenarios, multi-signal-to-noise ratios, and multi-spectrum features, and the blank voice data X train_empty . Each voice frequency data has a corresponding keyword label L train . The goal of the data preprocessing task is to detect the data and label φ train = (x dirty , l d ) in X d that affect the accuracy of the voice keyword detection model and perform deletion processing to determine the effective voice data set.
[0036] Further, determining the effective voice data set specifically includes determining the voice energy of each voice data in the original voice data set X train , and taking the voice data in X train and the keyword label as the sample pair φ i = (x i , l i), i = 1, 2, 3, …… N, perform steps S301 - S302, read the i-th speech (x i , l i ), calculate the speech energy E(x i ), calculate the speech energy E(x i ) according to the following formula (1):
[0037]
[0038] where, N s represents the number of speech sampling points of this speech.
[0039] Furthermore, compare the speech energy E(x i ) of each determined speech data with the representative speech energy E t of the blank speech data. If the former is greater than the latter, classify this speech data into the valid speech data set. Specifically, in step S303, determine whether E(x i ) is greater than E t . If the judgment result is no, perform step S305 to delete the blank speech. If the judgment result is yes, perform step S304 to store it in the non - blank speech data set (i.e., the valid speech data set). Repeat steps S302 - S305 until all blank speeches in the original speech data set X train are deleted, obtain the processed valid speech data set X no_silence , and perform step S306 to calculate the time - frequency feature f j corresponding to each speech x j .
[0040] In step S308, perform the k - th initialization of the keyword detection model (w, b) and Q k , that is, randomly initialize the parameters of the neural network model (w, b) for the speech keyword detection task and the vector Q k of the prediction misalignment times for this k - th time (which is a form of the sequence of prediction misalignment times for multiple times), where k represents the number of training experiments. The neural network model can minimize the loss function F loss , train the model by backpropagation of gradients, continuously learn the accurate conditional probability distribution between the input training samples and the labels, perform the t - th step of backpropagation gradient training of the model (step S309). In the t - th step of training, randomly extract η speech data features (f j , l j ) as training data and input them into the training model to obtain the updated model parameter w t(Step S310). That is, a set of voice data is randomly selected from the valid voice data set, and the number of voice data included accounts for 5%-20% of the total number of the valid voice data set. Under the updated model parameter w after parameter tuning t calculate the predicted label of each voice x j (Step S311). (Step S311).
[0041] In some embodiments, for the training and parameter tuning of the keyword detection model in each step: after predicting the label using the keyword detection model after parameter tuning, delete or update the current parameters of the keyword detection model; compare the predicted label with the keyword label of the corresponding voice data to determine whether the label prediction in this step is accurate, and then save the result of whether the label prediction in this step is accurate. In this step, after predicting the label using the keyword detection model after parameter tuning, after the predicted label, delete the current parameters of the keyword detection model after parameter tuning that have been used, which can avoid a large amount of parameters occupying a large amount of storage space. It is also possible to perform training based on the current parameters of the previous step in each step. After training to obtain new parameters, update the old current parameters of the previous step to the new current parameters. In addition, after comparing the predicted label with the keyword label of the corresponding voice data to determine whether the label prediction in this step is accurate, save the result of whether the label prediction in this step is accurate for subsequent processing based on this result.
[0042] Specifically, perform a prediction accuracy judgment on the predicted label of each voice x j and the keyword label l and define a prediction accuracy vector j to represent the state of whether the classification of voice x is accurate or not when the number of model training steps is t, that is, as in Step S312, calculate the prediction accuracy of each voice label j If the classification result (i.e., label prediction) of voice x at the t-th step j is equal to l (i.e., the label prediction is accurate), then p j = 1, otherwise p j = 0. j (Step S313).
[0043] Repeat Steps S307 - S312 until the training step t reaches the custom-defined t max and, according to the prediction accuracy vector P t count the vector of the number of times of inaccurate prediction of each voice data during the training process The counting rule is as in Step S313, and determine whether it is satisfied It can also be written as That is, for voice x jThe label prediction in the (t - 1)-th training step is accurate, but the label prediction in the t-th training step is incorrect. Then, the judgment result of step S313 is yes, and step S314 is executed. The number of times q that each voice prediction is inaccurate j ++, that is, the increment of the number of inaccurate predictions
[0044] If the judgment result of step S313 is no, then step S315 is executed. The number of times q that each voice prediction is inaccurate j = q j , then the increment of the number of inaccurate predictions is 0. Repeat steps S307 - S315 until the number of training experiments reaches the custom k max . In step S316, count the number vector Q of inaccurate predictions of the k-th model prediction k (which is a form of the sequence of the number of inaccurate predictions for multiple times). Then, execute step S317 to count the average number vector of inaccurate predictions of the model (which is a form of the sequence of average inaccurate prediction times), where Then, execute step S318. To determine the inaccurate prediction times threshold based on the average inaccurate prediction times sequence specifically includes determining the inaccurate prediction times threshold according to formula (2):
[0045] T = μ q + 3×σ q Formula (2),
[0046] where T represents the inaccurate prediction times threshold, and μ q represents the mean of each element of the average inaccurate prediction times sequence, and σ q represents the variance of each element of the average inaccurate prediction times sequence.
[0047] Based on the average number vector of inaccurate predictions obtain the average number of inaccurate predictions of the input voice (step S319). Then, execute step S320 to judge whether it is less than or equal to the inaccurate prediction times threshold T. If the judgment result is no, then execute step S321 to delete the dirty data voice. If the judgment result of step S320 is yes, then execute step S322 to store the voice data set after cleaning (x clean , l clean) Eventually, a preprocessed speech dataset is obtained (step S323), which can be used as the input for the subsequent speech keyword detection model, helping to improve the detection accuracy of the model. Starting from the data perspective, this application screens and deletes the dirty data in the original dataset used for model training due to improper data augmentation methods, incorrect data label annotation, etc., which can help the neural network model better learn the features of valid data, thereby improving the accuracy of the speech keyword detection task.
[0048] Furthermore, by performing secondary manual detection on the dirty data deleted by the screening and processing method of this application's embodiment, it is found that all the screened dirty data can indeed be determined as dirty data (blank speech, speech with obvious missing spectrogram features, speech with incorrect corresponding labels, etc.). Therefore, using this application can effectively improve the label credibility of the keyword detection dataset. In addition, based on the screening and processing method provided in this application's embodiment, the accuracy of keyword detection of the multi-keyword detection model and the efficiency during model training can be indirectly improved. By using the screening and processing method provided in this application's embodiment to process the original dataset, the data quality of the input model becomes higher. With the same model structure, it can quickly learn the correct keyword data features. The processed data removes the difficult-to-detect noise data, and the detection accuracy of the model for a single keyword and the overall accuracy have both been greatly improved.
[0049] Figure 4 The schematic diagram of the screening and processing device for the speech dataset according to the embodiment of this application is shown. Among them, the screening and processing device 400 includes an interface 401 and a processor 402. The interface 401 is configured to obtain the original speech dataset to be screened and processed, where each piece of speech data includes speech signal data and its keyword label. The interface 401 can include but is not limited to network adapters, cable connectors, serial connectors, USB connectors, parallel connectors, high-speed data transfer adapters, etc., such as optical fibers, USB 3.0, Thunderbolt interfaces, etc., and wireless network adapters, such as WiFi adapters, telecommunications (3G, 4G / LTE, etc.) adapters, etc.
[0050] In some embodiments, the interface 401 can be a network interface, and the device 400 can be connected to the network through the interface 401, such as but not limited to a local area network or the Internet.
[0051] The processor 402 is configured to execute the screening processing method of the voice data set described in various embodiments of the present application. The processor 402 can be a dedicated processor or a general-purpose processor. The processor 402 can include one or more known processing devices, such as microprocessors from the PentiumTM, CoreTM, XeonTM, or Itanium series manufactured by IntelTM, etc. Additionally, the processor 402 can include more than one processor, for example, a multi-core design or multiple processors, each with a multi-core design.
[0052] In some embodiments, the processor 402 can be a processing device including one or more general-purpose processing devices, such as a processing device of one or more general-purpose processing devices such as a microprocessor, a central processing unit (CPU), a graphics processing unit (GPU), etc. More specifically, the processor 602 can be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a processor running other instruction sets, or a processor running a combination of instruction sets. The processor 602 can also be one or more dedicated processing devices such as an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), a system-on-chip (SoC), etc.
[0053] The present application also provides a computer-readable storage medium, on which executable instructions are stored. When executed by a processor, the screening processing method of the voice data set described in various embodiments of the present application is implemented.
[0054] It can be understood that the computer-readable storage medium is such as but not limited to a read-only memory (ROM), a random access memory (RAM), a phase change random access memory (PRAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), an electrically erasable programmable read-only memory (EEPROM), other types of random access memories (RAM), a flash drive or other forms of flash memory, a cache, a register, a static memory, a compact disc read-only memory (CD-ROM), a digital versatile disc (DVD) or other optical storage, a cassette tape or other magnetic storage device, or any other non-transitory medium for storing information or instructions that can be accessed by a computer device, etc.
[0055] This document describes various operations or functions, which can be implemented or defined as software code or instructions. Such content can be in the form of directly executable ("object" or "executable") source code or differential code ("incremental" or "patch" code). The software implementation of the embodiments described herein can be provided via an article of manufacture storing the code or instructions or via a method of operating a communication interface to send data via the communication interface. A machine or computer-readable storage medium can cause a machine to perform the described functions or operations and includes any mechanism for storing information in a form accessible by a machine (e.g., a computing device, an electronic system, etc.), such as a recordable / non-recordable medium (e.g., read-only memory (ROM), random access memory (RAM), magnetic disk storage media, optical storage media, flash devices, etc.). A communication interface includes any mechanism for engaging with any one of a hardwired, wireless, optical, etc. medium to communicate with another device, such as a memory bus interface, a processor bus interface, an Internet connection, a disk controller, etc. The communication interface can be configured to be ready to provide a data signal describing the software content by providing configuration parameters and / or sending signals. The communication interface can be accessed via one or more commands or signals sent to the communication interface.
[0056] The components in this document can be implemented by an SOC (System on Chip). For example, various RISC (Reduced Instruction Set Computer) processor IPs purchased from companies such as ARM can be used as the processor of the SOC to perform corresponding functions, and it can be implemented as an embedded system. Specifically, there are many modules on the modules (IPs) available in the market, such as but not limited to memory, various communication modules, codecs, buffers, etc. Others such as antennas and speakers can be externally connected to the chip. Users can build an ASIC (Application Specific Integrated Circuit) based on the purchased IP or self-developed modules to implement various communication modules, codecs, etc., in order to reduce power consumption and cost. For example, users can also use an FPGA (Field Programmable Gate Array) to implement various communication modules, etc., which can be used to verify the stability of the hardware design. For various communication modules, etc., buffers are usually also provided to temporarily store the data generated during the processing.
[0057] In addition, although exemplary embodiments have been described herein, the scope includes any and all embodiments based on the present application having equivalent elements, modifications, omissions, combinations (e.g., schemes that cross various embodiments), adaptations, or alterations. The elements in the claims will be construed broadly based on the language employed in the claims and are not limited to the examples described in this specification or during the implementation of the present application, and the examples will be construed as non-exclusive. Thus, the present specification and examples are intended to be considered only as examples, and the true scope and spirit are indicated by the full scope of the claims and their equivalents.
[0058] The foregoing description is intended to be illustrative and not restrictive. For example, the above examples (or one or more of them) can be used in combination with each other. For example, those of ordinary skill in the art can use other embodiments when reading the above description. Additionally, in the above detailed description, various features can be grouped together to simplify the present application. This should not be construed as an intention that a feature of the application that is not claimed is necessary for any claim. On the contrary, the subject matter of the present application can be less than all the features of a particular embodiment of the application. Thus, the claims are incorporated herein as examples or embodiments into the detailed description, where each claim independently serves as a separate embodiment, and it is contemplated that these embodiments can be combined with each other in various combinations or permutations. The scope of the present application should be determined with reference to the appended claims and the full scope of the equivalent forms empowered by these claims.
[0059] The above embodiments are only exemplary embodiments of the present application and are not used to limit the present application. The protection scope of the present application is defined by the claims. Those skilled in the art can make various modifications or equivalent replacements within the essence and protection scope of the present application, and such modifications or equivalent replacements should also be regarded as falling within the protection scope of the present application.
Claims
1. A screening and processing method for a speech data set, where the screened and processed speech data set is used for training a keyword detection model, and is characterized in that, including the following steps by a processor: obtaining an original speech data set to be screened and processed, where each piece of speech data includes speech signal data and its keyword label; determining a valid speech data set based on the original speech data set; calculating time-frequency features for each piece of speech data in the valid speech data set; performing training and parameter tuning of a keyword detection model step by step, and determining the prediction misalignment count increment for each step of this piece of speech data. Specifically, for each step: extracting the time-frequency features of a group of speech data from the valid speech data set; based on the time-frequency features of the extracted group of speech data and the labels of their keywords, performing a backpropagation algorithm to adjust the parameters of the keyword detection model, thereby obtaining a keyword detection model with adjusted parameters; predicting labels using the keyword detection model with adjusted parameters based on the time-frequency features of this piece of speech data; comparing the predicted labels with the keyword labels of the corresponding speech data to determine whether the label prediction for this step is accurate; if the label prediction for this step is incorrect and the label prediction for the previous step is correct, the prediction misalignment count increment is 1, otherwise the prediction misalignment count increment is 0; by accumulating the prediction misalignment count increments for each step of this piece of speech data, obtaining the prediction misalignment count of this piece of speech data in the prediction misalignment count sequence for this time, to determine the prediction misalignment count sequences for each time. The elements of the prediction misalignment count sequence are arranged in the order of speech data and represent the prediction misalignment count of the corresponding speech data for this time; averaging the prediction misalignment count sequences for multiple times to obtain an average prediction misalignment count sequence. The elements of the average prediction misalignment count sequence are arranged in the order of speech data and represent the average prediction misalignment count of the corresponding speech data for multiple times; determining a prediction misalignment count threshold based on the average prediction misalignment count sequence; performing the following misalignment screening process for each piece of speech data to obtain a clean speech data set for training the keyword detection model: obtaining the average prediction misalignment count of this piece of speech data in the average prediction misalignment count sequence, comparing it with the prediction misalignment count threshold, if it is greater than the latter, it is determined as dirty speech data and deleted, otherwise it is retained and stored in the clean speech data set.
2. The screening and processing method according to claim 1, characterized in that, Before performing the training and parameter tuning of the keyword detection model step by step, the screening processing method further includes: initializing the parameters of the keyword detection model and the prediction misalignment count sequence for this time. The keyword detection model is constructed using a learning network.
3. The screening and processing method according to claim 1, characterized in that, Determining the valid speech data set based on the original speech data set specifically includes: determining the speech energy of each piece of speech data in the original speech data set; comparing the determined speech energy of each piece of speech data with the representative speech energy of blank speech data. If the former is greater than the latter, this piece of speech data is classified into the valid speech data set.
4. The screening and processing method according to claim 1, characterized in that, The group of speech data is randomly selected from the valid speech data set, and the number of speech data included accounts for 5%-20% of the total number of speech data in the valid speech data set.
5. The screening and processing method according to claim 2, characterized in that, The learning network includes an LSTM learning network or a GRU neural network, and the time-frequency features include at least one of MFCC features, Fbank features and their variants, and mel spectrograms.
6. The screening and processing method according to claim 1, characterized in that, Determining the prediction misalignment count threshold based on the average prediction misalignment count sequence specifically includes determining the prediction misalignment count threshold according to formula (2): T = μ q + 3×σ q Equation (2) where T represents the prediction misalignment count threshold, and μ q represents the mean of each element of the average prediction misalignment count sequence, and σ q represents the variance of each element of the average prediction misalignment count sequence.
7. The screening and processing method according to claim 1, characterized in that, Step-by-step execution of the training and parameter tuning of the keyword detection model specifically includes: the training and parameter tuning of the keyword detection model in each step are both executed based on the initialization parameters of the keyword detection model and the initialized prediction misalignment count sequence.
8. The screening and processing method according to claim 1, characterized in that, It further includes, for the training and parameter tuning of the keyword detection model in each step: after predicting the label using the keyword detection model with tuned parameters, deleting or updating the current parameters of the keyword detection model; comparing the predicted label with the keyword label of the corresponding speech data to determine whether the label prediction in this step is accurate, and then saving the result of whether the label prediction in this step is accurate.
9. A screening and processing device for a speech data set, characterized in that, Includes: An interface configured to obtain an original speech data set to be screened and processed, where each piece of speech data includes speech signal data and its keyword label; And A processor configured to execute the screening and processing method of the speech data set according to any one of claims 1-8.
10. A non-transitory computer storage medium, characterized in that, Stored thereon are executable instructions that, when executed by the processor, implement the screening and processing method of the speech data set according to any one of claims 1-8, and implement the screening and processing method of the speech data set according to any one of claims 1-7.
Citation Information
Patent Citations
Misannotation data screening method and device and computer storage medium
CN111931863A