Voice Detection Method, Device, Electronic Device and Computer Readable Storage Medium
Through the combination of feature extraction network and deep residual convolutional network, the extraction and dimensionality reduction speech features are solved, and the problem of low short-term speech detection accuracy in the prior art is achieved, and a higher speech classification detection accuracy is achieved.
Patent Information
- Application Number
- CN202111443833.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-30
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-11-30
AI Technical Summary
When the effective speech duration is short, the accuracy of speech detection will be seriously reduced.
A feature extraction network is used to extract the detected speech to be featured, and an encoded feature vector is obtained, and a deep residual convolutional network is used to reduce the dimensions to obtain a representation vector containing the distinction information of speech category. Finally, the speech category is determined based on the similarity between the representation vector and the target vector.
By retaining more time-frequency information and capturing the speech category distinction information of time-frequency and frequency domains in short-term voice, the accuracy of short-term voice classification detection is improved.
Smart Images

Figure CN114333771B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technologies, and in particular, to a voice detection method, apparatus, electronic device, and computer-readable storage medium. Background Art
[0002] With the development of computer technologies, the detection of voice types has become increasingly important. However, current voice detection methods mainly rely on the Total Variability system. By calculating the posterior probability and first- and second-order statistics frame by frame, and establishing a model through factor analysis, voice category detection is achieved. However, when the effective voice duration is short, the accuracy of voice detection will seriously decline. Summary of the Invention
[0003] The main technical problem to be solved by this application is to provide a voice detection method, apparatus, electronic device, and computer-readable storage medium that can improve the accuracy of voice detection.
[0004] To solve the above technical problem, in the first aspect of this application, a voice detection method is provided. The method includes: using a feature extraction network to extract features from the voice to be detected to obtain an encoded feature vector of the voice to be detected; using a deep residual convolutional network to perform dimensionality reduction processing on the encoded feature vector of the voice to be detected to obtain a representation vector of the voice to be detected, where the representation vector contains voice category discrimination information of the voice to be detected; and determining the voice category of the voice to be detected according to the similarity between the representation vector and the target vector.
[0005] To solve the above technical problem, in the second aspect of this application, a voice detection apparatus is provided. The apparatus includes: a feature extraction module, configured to use a feature extraction network to extract features from the voice to be detected to obtain an encoded feature vector of the voice to be detected; a voice detection module, configured to use a deep residual convolutional network to perform dimensionality reduction processing on the encoded feature vector of the voice to be detected to obtain a representation vector of the voice to be detected, where the representation vector contains voice category discrimination information of the voice to be detected; and a determination module, configured to determine the voice category of the voice to be detected according to the similarity between the representation vector and the target vector.
[0006] To solve the above technical problem, in the third aspect of this application, an electronic device is provided. The electronic device includes a memory and a processor coupled to each other. The memory is used to store program data, and the processor is used to execute the program data to implement the foregoing method.
[0007] To solve the above technical problem, in the fourth aspect of this application, a computer-readable storage medium is provided. The computer-readable storage medium stores program data, and when the program data is executed by a processor, it is used to implement the foregoing method.
[0008] The beneficial effects of the present application are as follows: Different from the prior art, the present application first uses a feature extraction network to extract features from the speech to be detected to obtain the encoded feature vector of the speech to be detected, and then uses a deep residual convolutional network to perform dimensionality reduction processing on the encoded feature vector of the speech to be detected to obtain the representation vector of the speech to be detected. Among them, the representation vector contains the speech category discrimination information of the speech to be detected. Then, according to the similarity between the representation vector and the target vector, the speech category of the speech to be detected is determined. Through the above, the feature extraction network can retain more time-frequency information in the short speech, and the deep residual convolutional network can capture the speech category discrimination information in the time-frequency and frequency domains of the short speech, thereby improving the problem of low accuracy of short speech classification detection. Brief Description of the Drawings
[0009] To more clearly illustrate the technical solutions in the present application, the following will briefly introduce the drawings required in the description of the embodiments. Obviously, the following described drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings. Among them:
[0010] Figure 1 is a schematic flowchart of an embodiment of the speech detection method of the present application;
[0011] Figure 2 is a schematic flowchart of another embodiment of the speech detection method of the present application;
[0012] Figure 3 is a schematic flowchart of an embodiment of the training method of the feature extraction network of the present application;
[0013] Figure 4 is Figure 3 a schematic flowchart of another implementation manner of step S33 in
[0014] Figure 5 is a schematic flowchart of an embodiment of the training method of the deep residual convolutional network of the present application;
[0015] Figure 6 is a schematic diagram of the distribution of the supervised training speech set and the weight vector of the present application;
[0016] Figure 7 is a schematic block diagram of an embodiment of the speech detection device of the present application;
[0017] Figure 8 is a schematic block diagram of an embodiment of the electronic device of the present application;
[0018] Figure 9 is a schematic block diagram of an embodiment of the computer-readable storage medium of the present application. Detailed implementation manners
[0019] In this application, the mention of "embodiment" means that the specific features, structures or characteristics described in combination with the embodiment may be included in at least one embodiment of this application. The appearance of this phrase at various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein may be combined with other embodiments.
[0020] The terms "first" and "second" in this application are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of this application, the meaning of "a plurality of" is at least two, such as two, three, etc., unless otherwise specifically defined. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.
[0021] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by this application.
[0022] Before introducing the voice detection method of this application, a simple introduction to the currently adopted full-variable factor analysis model is given first:
[0023] (1) Extract spectral features reflecting frequency characteristics from a segment of voice, such as Mel Frequency Cepstrum Coefficient (MFCC), Perceptual Linear Predictive (PLP), etc.; specifically, the spectral features can be extracted by Fourier transform.
[0024] (2) Calculate the posterior occupancy of each frame of speech data in each Gaussian component of the Gaussian Mixture Model (GMM) in chronological order through the Baum-Welch algorithm. Through the Total Variability (TV) model pre-trained with unlabeled true and false voice data, linearly project the speech features to obtain the corresponding vector model (Ivector) of this segment of speech.
[0025] (3) Use two types of labeled data (such as true and false voice data) to train the Linear Discriminant Analysis (LDA) backend, further reduce the dimension of the Ivector to obtain a lower-dimensional vector L-Ivector, and average all the speeches with the same label to obtain a characterization vector of true voice and a characterization vector of synthetic voice.
[0026] (4) During testing, the similarity between the L-Ivector of the speech to be detected and the above two characterization vectors can be calculated respectively to determine the category of the speech to be detected.
[0027] Different from the above method, the present application provides a speech detection technology based on an end-to-end deep neural network. Among them, more time-frequency information in short-time speech can be retained through the feature extraction network, and the deep residual convolutional network can capture the speech category discrimination information in the time-frequency and frequency domains of short-time speech, so as to improve the problem of low accuracy in short-time speech classification and detection.
[0028] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of an embodiment of the speech detection method of the present application. Among them, the execution subject of the present application is an electronic device such as a mobile phone, a computer, a wearable device, etc.
[0029] The method may include the following steps:
[0030] Step S11: Use the feature extraction network to extract features from the speech to be detected to obtain the encoded feature vector of the speech to be detected.
[0031] Specifically, a segment of speech to be detected can be input into a pre-trained feature extraction network, and then the feature extraction network extracts features from the speech to be detected, so as to obtain the encoded feature vector of the speech to be detected. Compared with the spectral features extracted by the traditional total-variability factor analysis method, the encoded feature vector extracted by the feature extraction network in this embodiment can retain more time-frequency information, which is beneficial for the backend to judge the speech category. For the structure and training process of the feature extraction network, refer to the subsequent embodiments.
[0032] Step S12: Use a deep residual convolutional network to perform dimensionality reduction on the encoded feature vector of the speech to be detected, so as to obtain a representation vector of the speech to be detected, where the representation vector contains speech category discrimination information of the speech to be detected.
[0033] Specifically, the encoded feature vector can be input into the deep residual convolutional network, and the deep residual convolutional network extracts the speech category discrimination information of the speech to be detected. Among them, the deep residual convolutional network adopts a structure of multiple layers of convolution + residual, which can reduce the dimension of the encoded feature vector output by the front-end feature extraction network, so as to extract the speech category discrimination information therein.
[0034] Among them, the speech category discrimination information includes information that can distinguish speech categories, which is included in the representation vector, so as to facilitate subsequent determination of the speech category of the speech to be detected according to the representation vector.
[0035] Step S13: Determine the speech category of the speech to be detected according to the similarity between the representation vector and the target vector.
[0036] In some embodiments, the speech categories may include the first type of speech and the second type of speech. The types of speech categories in this embodiment are not limited. For example, the speech categories may also include the third type of speech, the fourth type of speech, etc.
[0037] The target vector is a vector set in advance. Among them, the more similar the representation vector is to the target vector, the closer the speech to be detected is to the target class corresponding to the target vector. On the contrary, it means that the speech to be detected is closer to the non-target class. In other embodiments, the speech categories can also be divided into multiple categories according to the similarity between the representation vector and the target vector. For example, the similarity degree is 80% - 100% for the first type of speech, 60% - 80% for the second type of speech, 40% - 60% for the third type of speech, and 0% - 40% for the fourth type of speech.
[0038] In this embodiment, first use the feature extraction network to extract features from the speech to be detected to obtain the encoded feature vector of the speech to be detected, then use the deep residual convolutional network to perform dimensionality reduction on the encoded feature vector of the speech to be detected to obtain the representation vector of the speech to be detected, where the representation vector contains the speech category discrimination information of the speech to be detected, and then determine the speech category of the speech to be detected according to the similarity between the representation vector and the target vector. Through the above, the feature extraction network can retain more time-frequency information in the short-time speech, and the deep residual convolutional network can capture the speech category discrimination information in the time-frequency and frequency domain of the short-time speech, so as to improve the problem of low accuracy of short-time speech classification detection.
[0039] Please refer to Figure 2 , Figure 2It is a schematic flowchart of another embodiment of the voice detection method of the present application.
[0040] The method may include the following steps:
[0041] Step S21: Use a feature extraction network to extract features from the voice to be detected, so as to obtain an encoded feature vector of the voice to be detected.
[0042] In this embodiment, the feature extraction network includes an encoder module. The encoder module is used to perform encoding processing on the voice to be detected, so as to obtain an encoded feature vector of the voice to be detected.
[0043] In some embodiments, the feature extraction network includes a cascaded encoder module and decoder module. Among them, the decoding module is used to perform decoding processing on the encoded feature vector output by the encoding module to obtain a decoded feature vector, so as to use the decoded feature vector to train the feature extraction network. It can be understood that after the training is completed, only the encoder module can be used to obtain the encoded feature vector.
[0044] In a specific embodiment, the network structure of the feature extraction network is as follows:
[0045] Table 1 An example network structure of the feature extraction network
[0046]
[0047] Among them, the encoder module includes 5 different convolutional layers, namely Conv1, Conv2, Conv3, Conv4, and Conv5, and the corresponding kernel sizes are (10; 8; 4; 4; 4), and the corresponding strides are (5; 4; 2; 2; 2). The decoder module includes 9 identical convolutional layers (Conv6), and the kernel size of each is 3 and the stride is 1. Of course, the above network structure is only an example of the feature extraction network. In other embodiments, the network structure of the feature extraction network can also be adjusted according to actual situations, and this is not limited. In some embodiments, the decoder module further includes a fully connected layer.
[0048] Step S22: Use a deep residual convolutional network to perform dimensionality reduction processing on the encoded feature vector of the voice to be detected, so as to obtain a representation vector of the voice to be detected, where the representation vector contains voice category discrimination information of the voice to be detected.
[0049] In a specific embodiment, the network structure of the deep residual convolutional network is as follows:
[0050] Table 2 An example network structure of the deep residual convolutional network
[0051]
[0052] It can be seen that the deep residual convolutional network of the present application adopts a structure of multi-layer convolution + residual, that is, it includes the above network layers Conv1, Conv2_x, Conv3_x, Conv4_x, Conv5_x and Conv6, and the specific parameters are shown in Table 2. In addition, the commonly used statistical pooling method (Statistic Pooling) in the field of speaker recognition is used for dimensionality reduction. After linear layer transformation, a 256-dimensional feature vector - the output of FC is finally obtained. Of course, the above network structure is only an example of the deep residual convolutional network. In other embodiments, the network structure of the deep residual convolutional network can also be adjusted according to actual situations, and this is not limited herein.
[0053] In this embodiment, steps S23 to S26 are another implementation manner of step S13 in the above embodiment.
[0054] Step S23: Calculate the distance between the feature vector and the target vector.
[0055] In some embodiments, the similarity between two vectors can be reflected by the distance between the vectors. For example, the smaller the distance between two vectors, the more similar the two vectors are, and the larger the distance between two vectors, the less similar the two vectors are.
[0056] In an example, the target class corresponding to the target vector is the first type of speech. The closer the distance between the feature vector and the target vector is, the greater the possibility that the speech to be detected is the first type of speech. On the contrary, the smaller the possibility that the speech to be detected is the first type of speech, that is, the greater the possibility that it is other types of speech.
[0057] Among them, the methods for calculating the distance between two vectors include: Euclidean Distance, Manhattan Distance, Chebyshev Distance, Minkowski Distance, Mahalanobis Distance, cosine distance or cosine of the included angle (Cosine), etc. In practical applications, a suitable calculation method can be selected according to needs to calculate the distance between the feature vector and the target vector, and this is not limited in this embodiment.
[0058] As an example, the cosine of the included angle can be used to calculate the distance between the feature vector and the target vector. Specifically, the following formula can be used:
[0059]
[0060] As described above, x is the feature vector, y is the target vector, and cos(θ) is the cosine of the angle. The value range of the cosine of the angle is [-1, 1]. The larger the cosine of the angle, the smaller the angle between the two vectors, and the greater the similarity; the smaller the cosine of the angle, the larger the angle between the two vectors, and the smaller the similarity. When the directions of the two vectors coincide, the cosine of the angle takes the maximum value of 1; when the directions of the two vectors are completely opposite, the cosine of the angle takes the minimum value of -1.
[0061] Step S24: Determine whether the distance between the feature vector and the target vector is less than a preset distance threshold.
[0062] If so, execute step S25; otherwise, execute step S26.
[0063] The preset distance threshold can be set according to the actual situation and is not limited here.
[0064] In some embodiments, determine whether the cosine of the angle between the feature vector and the target vector is greater than a preset cosine threshold; if so, execute step S133; otherwise, execute step S134. It can be understood that the preset cosine threshold corresponds to the preset distance threshold. When the cosine of the angle is greater than the preset cosine threshold, the distance between the corresponding two vectors is less than the preset distance threshold.
[0065] The preset cosine threshold is, for example, 0.7, 0.8, 0.9, etc. For example, when the cosine of the angle between the feature vector and the target vector is greater than 0.8, it is determined that the speech to be detected is the first type of speech; otherwise, it is determined that the detected speech is the second type of speech.
[0066] Step S25: Determine that the speech to be detected is the first type of speech.
[0067] Step S26: Determine that the speech to be detected is the second type of speech.
[0068] In some embodiments, the first type of speech is real speech, and the second type of speech is synthetic speech (also known as fake speech). Among them, real speech refers to the speech emitted by humans through their vocal organs, while synthetic speech refers to artificial speech generated by mechanical or electronic methods. Through computer speech synthesis, any text can be converted into speech with high naturalness at any time, thus truly enabling the machine to "speak like a human".
[0069] At present, computer-synthesized speech is getting closer and closer to human pronunciation and is applied to all aspects of our lives, such as intelligent customer service, mobile phone assistants, etc., which facilitates our lives. On the other hand, however, some lawbreakers have begun to use this technology for illegal and criminal activities such as telecommunications fraud. Moreover, the voiceprint recognition system will show great vulnerability when under attack by synthetic voices. Such attacks have greatly affected the security of the voiceprint recognition system itself, and in turn have also brought security risks to the system that uses voiceprint recognition technology for access control. Therefore, it is very necessary to find corresponding countermeasures. However, the current methods for detecting synthetic speech are mainly based on the all-variable system, and the all-variable system uses less time-frequency information retained by the spectral features (MFCC). Thus, in scenarios where the available effective speech duration is short, the accuracy of speech classification will be affected.
[0070] Differently, this embodiment provides a speech detection method. The features extracted by the feature extraction network can retain more time-frequency characteristics, which is beneficial for the backend to distinguish between real speech and synthetic speech, and a high classification accuracy can also be obtained in the short-time speech scenario. Secondly, the deep residual convolution network adopts a network structure of deep residual convolution plus statistical pooling, which can reduce the dimension of the features output by the front-end feature extraction network and extract and capture the true and false voice information in the time domain and frequency domain of short-time speech.
[0071] Please refer to Figures 3 to 4 , Figure 3 which is a schematic flowchart of an embodiment of the training method of the feature extraction network of this application, Figure 4 and Figure 3 is a schematic flowchart of another implementation manner of step S33 in
[0072] The training method of the feature extraction network may include the following steps:
[0073] Step S31: Use the encoder module to perform convolution processing on the unsupervised training speech to obtain the encoded feature vector of the unsupervised training speech.
[0074] Among them, the unsupervised training speech is unlabeled training speech. It can be understood that during the training process of the network, a large amount of training data is required to support it. In this regard, the number of unsupervised training speeches in this embodiment can be selected according to actual needs and is not limited here.
[0075] Step S32: Use the decoder module to perform convolution processing on the encoded feature vector of the unsupervised training speech to obtain the decoded feature vector of the unsupervised training speech.
[0076] Each unsupervised training speech corresponds to an encoded feature vector and a decoded feature vector.
[0077] Step S33: Calculate the loss using the encoded feature vectors and decoded feature vectors of the unsupervised training speech to obtain the first loss result of the unsupervised training speech.
[0078] In some embodiments, step S33 may include sub-steps S331 to S333:
[0079] Step S331: Calculate the predicted feature vectors for the first to the preset number of steps in the future based on the decoded feature vectors.
[0080] The preset number of steps can be set in advance or changed later. In this embodiment, the preset number is denoted as K, and K = 12.
[0081] Specifically, the encoded feature vectors of the output of the encoding module for the k-th step in the future can be predicted through the linear transformation of the fully connected layer. In this embodiment, the above feature extraction network may further include a fully connected layer.
[0082] In some embodiments, h k (c i ) can be used to predict the encoded feature vectors of the output of the encoding module for the k-th step in the future, where the value range of k is [0, K]. h k (c i ) = W k C i + b k , where W k represents the weight, and b k represents the bias.
[0083] Step S332: Calculate the loss between the encoded feature vectors and the predicted feature vectors corresponding to each step to obtain the second loss result corresponding to each step.
[0084] Specifically, the following formula can be used for calculation:
[0085]
[0086] where, L k is the second loss result for the k-th step, Z i+k is the encoded feature vector for the k-th step in the future, h k (c i ) is the predicted feature vector for the k-th step in the future, is the preset number of encoded feature vectors that are not the k-th step in the future and are uniformly sampled from multiple encoded feature vectors, and T is the transpose. Among them, σ represents the Sigmoid function, that is, σ(x) = 1 / (1 + exp(-x)).
[0087] When k = 1, it means h 1 (c 1)The prediction feature vector for the first step of predicting the future, and the calculated loss is Z 2 and h 1 (c 1 )、Z 3 and h 1 (c 2 ), and so on. When k = 1, represents a negative example, which can be 10 coding feature vectors that are not the first step of the future uniformly sampled from multiple coding feature vectors. Here, the preset quantity is, for example, 10, or it can be selected according to the actual situation.
[0088] Step S333: Statistically sum all the second loss results to obtain the first loss result of the unsupervised training speech.
[0089] In some embodiments, the first loss result is calculated using the following formula:
[0090]
[0091] where L is the first loss result, K is the preset quantity, and K can be 12. That is, calculate the corresponding second loss results when predicting from the first step to the twelfth step of the future respectively, so as to obtain the first loss result of the unsupervised training speech, and thus the training of the feature extraction network can be realized.
[0092] Step S34: Adjust the parameters of the feature extraction network using the first loss result.
[0093] The sign that the network training is completed is that the first loss result L is less than the prediction loss threshold, or the network iteration reaches the prediction training times. Otherwise, continue to adjust the parameters of the feature extraction network using the first loss result. The prediction loss threshold and the prediction training times can be set as needed and are not limited here.
[0094] Above, when training the feature extraction network in this embodiment, an unsupervised training method is adopted, without data with true and false voice labels, thus saving the training cost and training time of the model. Secondly, by calculating the comparison loss function using the outputs of the encoding module and the decoding module, unsupervised training can be realized. By pre-training a feature extraction network composed of an encoding module and a decoding module, and replacing the traditional spectral features with the feature vectors vector extracted by the network, more time-frequency characteristics can be retained, which is beneficial for the true and false voice discrimination performed by the backend.
[0095] Please refer to Figure 5 , Figure 5 is a schematic flowchart of an embodiment of the training method of the deep residual convolutional network of the present application, Figure 6 is a schematic diagram of the distribution of the supervised training speech set and the weight vector of the present application.
[0096] The training method of the deep residual convolutional network may include the following steps:
[0097] Step S41: Process the supervised training speech set by using the deep residual convolutional network to obtain the representation vector of each supervised training speech, where the representation vector contains the speech classification information of the supervised training speech.
[0098] Among them, the supervised training speech is the labeled training speech. The supervised training speech set includes multiple supervised training speeches.
[0099] In this embodiment, the deep residual convolutional network outputs a 256-dimensional representation vector.
[0100] Step S42: Calculate the third loss result of each supervised training speech by using the single-class loss function.
[0101] The single-class loss function calculates the loss based on a single class.
[0102] Specifically, the first threshold, the second threshold, the weight vector, and the representation vector of the supervised training speech can be used to calculate the third loss result of each supervised training speech. Among them, the first threshold is used to limit the angle between the weight vector and the representation vector of the supervised training speech of the first class to be less than the first preset angle, so that the class of the supervised training speech of the first class is as compact as possible. The second threshold is used to limit the angle between the weight vector and the representation vector of the supervised training speech of the second class to be greater than the second preset angle, and the second preset angle is greater than the first preset angle, so that the class of the supervised training speech of the second class remains a certain degree of looseness, so that the network can ensure the generalization ability when facing various second-class speeches with different distributions.
[0103] Among them, the weight vector changes during each training process, and the final weight vector can be used as the target vector in subsequent speech classification after training is completed.
[0104] In some embodiments, the third loss result is calculated by the following formula:
[0105]
[0106] where L i is the third loss result of the i-th supervised training speech. y i is the class label of the i-th supervised training speech, and the value of y i is 0 or 1. Among them, when the value of y i is 1, it means that the label of this supervised training speech is the first class (such as real speech), and when the value of y i is 0, it means that the label of this supervised training speech is the second class (such as synthetic speech). m0 is the first threshold, m 1 is the second threshold, m 0 and m 1 can take values in the range [-1, 1], and m 0 > m 1 is used to limit the angle (θ) between w 0 and x. w 0 is the weight vector. x i is the representation vector of the i-th supervised training speech, and α is the scale parameter used to control the change of the loss function in amplitude.
[0107] Among them, when y i = 0, m 0 is used to limit θ to be less than arccos m 0 , and when y i = 1, m 1 is used to limit θ to be greater than arccos m 1 . In this embodiment, the real speech is used as the target class and the synthetic speech is used as the non-target class. Therefore, by using a smaller value of arccos m 0 , the class of the real speech can be made as compact as possible; at the same time, by using a larger value of arccos m 1 , the class of the synthetic speech can be kept somewhat loose, so that the network can ensure generalization ability in the face of various different distributions of fake voices. However, in the face of a wide variety of synthetic speech synthesis methods and synthesizer parameters, when the model faces unknown speech synthesized by a synthesizer different from the training set, due to different synthesizer parameters and synthesis quality, the calculation of the relevant statistics of the full-variable factor analysis model is interfered, resulting in an unstable established model and ultimately a serious decline in the detection accuracy. Different from this, the single-classification loss function provided in this embodiment can make the decision surface of the real speech more strict and the decision surface of the synthetic speech more loose by setting the first threshold m 0 and the second threshold m 1 , so that the system has better generalization when facing unknown synthetic speech types, thereby improving the accuracy of speech detection.
[0108] As Figure 6 shown, the squares represent the target class samples (such as real speech samples), and the dots represent the non-target class samples (such as synthetic speech samples). It can be seen that a smaller arccos m 0 can make the target class concentrated around the weight vector w 0 , and at the same time, by using a larger arccos m 1 , the non-target class samples can be pushed away from w 0 .
[0109] Step S43: Calculate the mean of the third loss results of a preset number of supervised training voices to obtain the fourth loss result corresponding to the preset number of supervised training voices.
[0110] In some embodiments, the fourth loss result is calculated using the following formula:
[0111]
[0112] where L OCS is the fourth loss result. N is the number of supervised training voices in the supervised training voice set, which can specifically be the number of samples in a minibatch.
[0113] Step S44: Determine whether the fourth loss result meets the preset requirements.
[0114] If so, execute Step S45; otherwise, execute Step S46.
[0115] Specifically, it can be determined whether the fourth loss result is less than a preset loss threshold, or whether the number of training times corresponding to the fourth loss result reaches a preset number of training times. If so, it is determined that the preset requirements are met; otherwise, it is determined that the preset requirements are not met.
[0116] Step S45: Determine that the training is completed, and use the weight vector as the target vector.
[0117] The target vector is the representative vector of the first type of voice (such as a real voice). The more similar the feature vector is to the target vector, the closer the corresponding voice type is to the first type of voice; otherwise, the farther it is from the first type of voice.
[0118] Step S46: Adjust the parameters of the deep residual convolutional network using the fourth loss result.
[0119] When the fourth loss result does not meet the preset requirements, it indicates that the training is not yet completed at this time. Then continue to adjust the parameters of the deep residual convolutional network using the fourth loss result, and then continue to train the deep residual network using the supervised training voices until the fourth loss result meets the preset requirements, at which point the training is completed.
[0120] As above, because it is highly likely that the synthetic speech distributions are completely dissimilar during training and testing, if binary classification objective function training is performed, overfitting is likely to occur in the synthetic speech. Therefore, it is more suitable to regard the speech classification detection as a single classification task. In this embodiment, the single classification loss function OC loss makes the decision surface of the real voice more strict and the decision surface of the synthetic voice more relaxed by introducing different angular intervals between the real voice and the synthetic voice, enabling the system to have better generalization when facing unknown synthetic voice types.
[0121] Please refer to Figure 7 , Figure 7 which is a structural schematic block diagram of an embodiment of the voice detection device of the present application.
[0122] The voice detection device 100 includes a feature extraction module 110, a voice detection module 120, and a determination module 130. Among them, the feature extraction module 110 is used to extract features from the voice to be detected by using a feature extraction network to obtain a coded feature vector of the voice to be detected; the voice detection module 120 is used to perform dimensionality reduction processing on the coded feature vector of the voice to be detected by using a deep residual convolutional network to obtain a characterization vector of the voice to be detected, where the characterization vector contains voice category discrimination information of the voice to be detected; the determination module 130 is used to determine the voice category of the voice to be detected according to the similarity between the characterization vector and the target vector.
[0123] In some embodiments, the feature extraction network includes a cascaded encoder module and a decoder module. The feature extraction module 110 is further used to perform convolutional processing on the unsupervised training voice by using the encoder module to obtain a coded feature vector of the unsupervised training voice; perform convolutional processing on the coded feature vector of the unsupervised training voice by using the decoder module to obtain a decoded feature vector of the unsupervised training voice; calculate a loss by using the coded feature vector and the decoded feature vector of the unsupervised training voice to obtain a first loss result of the unsupervised training voice; and adjust the parameters of the feature extraction network by using the first loss result to implement the training of the feature extraction network.
[0124] In some embodiments, the feature extraction module 110 is further used to calculate predicted feature vectors for the first to the preset number of steps in the future respectively based on the decoded feature vector; calculate a loss for each step by using the corresponding coded feature vector and predicted feature vector for each step to obtain a second loss result corresponding to each step; and sum up all the second loss results to obtain the first loss result of the unsupervised training voice.
[0125] In some embodiments, the first loss result is calculated by using the following formula:
[0126]
[0127]
[0128] where L is the first loss result, K is the preset number, L k is the second loss result of the k-th step, Z i+k is the coded feature vector of the k-th step in the future, h k (c i ) is the predicted feature vector of the k-th step in the future, is a preset number of encoded feature vectors sampled uniformly from multiple encoded feature vectors and not the k-th future step, and T is the transpose.
[0129] In some embodiments, the speech detection module 120 is further configured to process the supervised training speech set by using a deep residual convolutional network to obtain a representation vector for each supervised training speech, where the representation vector includes speech classification information of the supervised training speech; calculate a third loss result for each supervised training speech by using a single-class loss function; calculate the mean of the third loss results of the preset number of supervised training speeches to obtain a fourth loss result corresponding to the preset number of supervised training speeches; determine whether the fourth loss result meets a preset requirement; if so, determine that the training is completed and use the weight vector as the target vector; otherwise, adjust the parameters of the deep residual convolutional network by using the fourth loss result.
[0130] In some embodiments, the speech detection module 120 is further configured to calculate a third loss result for each supervised training speech by using a first threshold, a second threshold, a weight vector, and a representation vector of the supervised training speech; wherein, the first threshold is used to limit the angle between the weight vector and the representation vector of the supervised training speech of the first category to be less than a first preset angle, and the second threshold is used to limit the angle between the weight vector and the representation vector of the supervised training speech of the second category to be greater than a second preset angle, and the second preset angle is greater than the first preset angle.
[0131] In some embodiments, the fourth loss result is calculated by using the following formula:
[0132]
[0133] where, L OCS is the fourth loss result, N is the number of supervised training speeches in the supervised training speech set, y i is the class label of the i-th supervised training speech, taking values of 0 or 1, m 0 is the first threshold, m 1 is the second threshold, w 0 is the weight vector, x i is the representation vector of the i-th supervised training speech, and α is a scale parameter.
[0134] In some embodiments, the deep residual convolutional network reduces the dimension of the encoded feature vectors by using statistical pooling.
[0135] In some embodiments, the determination module 130 is further configured to calculate the distance between the representation vector and the target vector; determine whether the distance is greater than a preset distance threshold; if so, determine that the speech to be detected is a first-class speech; otherwise, determine that the detected speech is a second-class speech.
[0136] In some embodiments, the first type of voice is a real voice, and the second type of voice is a synthesized voice.
[0137] Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of an embodiment of the electronic device of the present application.
[0138] The electronic device 200 may include a memory 210 and a processor 220 that are coupled to each other. The memory 210 is used to store program data, and the processor 220 is used to execute the program data to implement the steps in any of the above method embodiments.
[0139] The electronic device 200 may include, but is not limited to: a television, a desktop computer, a laptop computer, a handheld computer, a wearable device, a head-mounted display, a reader device, a portable music player, a portable game console, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, as well as a cellular phone, a personal digital assistant (PDA), an augmented reality (AR), and a virtual reality (VR) device.
[0140] Specifically, the processor 220 is used to control itself and the memory 210 to implement the steps in any of the above method embodiments. The processor 220 may also be referred to as a CPU (Central Processing Unit). The processor 220 may be an integrated circuit chip with signal processing capabilities. The processor 220 may also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 220 may be implemented by multiple integrated circuit chips together.
[0141] Please refer to Figure 9 , Figure 9 which is a schematic structural diagram of an embodiment of the computer-readable storage medium of the present application.
[0142] The computer-readable storage medium 300 stores program data 310, which, when executed by a processor, is used to implement the steps in any of the above method embodiments.
[0143] The computer-readable storage medium 300 can be a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc., which can store computer programs, or it can be a server storing the computer program. The server can send the stored computer program to other devices for running, or it can also run the stored computer program itself.
[0144] In several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling, direct coupling, or communication connection can be through some interfaces. The indirect coupling or communication connection of devices or units can be in an electrical, mechanical, or other form.
[0145] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0146] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0147] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods according to various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0148] The above are only the embodiments of this application, and do not limit the patent scope of this application. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of this application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of this application.
Claims
1. A voice detection method, characterized in that, it includes: using a feature extraction network to extract features from the voice to be detected to obtain an encoded feature vector of the voice to be detected; wherein, the feature extraction network is obtained by adjusting the parameters of the feature extraction network using the first loss result of unsupervised training of the voice; the calculation steps of the first loss result include: calculating predicted feature vectors for the first to the preset number of steps in the future based on the decoded feature vectors of the unsupervised training voice; calculating the loss between the encoded feature vector and the predicted feature vector corresponding to each step to obtain a second loss result corresponding to each step; statistically summing all the second loss results to obtain the first loss result; wherein, the encoded feature vector is obtained by performing convolutional processing on the unsupervised training voice using the encoder module of the feature extraction network, and the decoded feature vector is obtained by performing convolutional processing on the encoded feature vector using the decoder module of the feature extraction network; using a deep residual convolutional network to perform dimensionality reduction processing on the encoded feature vector of the voice to be detected to obtain a representation vector of the voice to be detected, wherein the representation vector contains voice category discrimination information of the voice to be detected; determining the voice category of the voice to be detected according to the similarity between the representation vector and the target vector.
2. The method according to claim 1, characterized in that, the first loss result is calculated using the following formula: Wherein, L is the first loss result, K is a preset quantity, and L k is the second loss result at the k-th step, and Z i+k is the encoded feature vector at the k-th future step, and h k (c i ) is the predicted feature vector at the k-th future step, is a preset quantity of the encoded feature vectors that are not at the k-th future step and are uniformly sampled from multiple said encoded feature vectors, T is the transpose, and σ represents the Sigmoid function.
3. The method according to claim 1, characterized in that, the training method of the deep residual convolutional network includes: using the deep residual convolutional network to process a supervised training voice set to obtain a representation vector of each supervised training voice, wherein the representation vector contains voice classification information of the supervised training voice; using a single-class loss function to calculate a third loss result for each supervised training voice; calculating the mean of the third loss results of a preset number of supervised training voices to obtain a fourth loss result corresponding to the preset number of supervised training voices; judging whether the fourth loss result meets a preset requirement; if not, then adjusting the parameters of the deep residual convolutional network using the fourth loss result.
4. The method according to claim 3, characterized in that, the calculating the third loss result for each supervised training voice using the single-class loss function includes: using a first threshold, a second threshold, a weight vector, and the representation vector of the supervised training voice to calculate the third loss result for each supervised training voice; wherein, the first threshold is used to limit the angle between the weight vector and the representation vector of the supervised training voice of the first category to be less than a first preset angle, the second threshold is used to limit the angle between the weight vector and the representation vector of the supervised training voice of the second category to be greater than a second preset angle, and the second preset angle is greater than the first preset angle; after judging whether the fourth loss result meets a preset requirement, it further includes: If so, determine that the weight determination training is completed, and use the weight vector as the target vector.
5. The method according to claim 4, wherein, the third loss result is calculated using the following formula: where, L OCS is the third loss result, N is the number of supervised training voices in the supervised training voice set, y i is the class label of the i-th supervised training voice, taking values of 0 or 1, m 0 is the first threshold, m 1 is the second threshold, w 0 is the weight vector, x i is the feature vector of the i-th supervised training voice, and α is the scale parameter.
6. The method according to claim 1, wherein, the deep residual convolutional network reduces the dimension of the encoded feature vector by means of statistical pooling.
7. The method according to claim 1, wherein, determining the speech category of the speech to be detected according to the similarity between the characterization vector and the target vector includes: calculating the distance between the characterization vector and the target vector; judging whether the distance is greater than a preset distance threshold; if so, determining that the speech to be detected is the first type of speech; otherwise, determining that the speech to be detected is the second type of speech.
8. The method according to claim 7, wherein, the first type of speech is real speech, and the second type of speech is synthetic speech.
9. A speech detection device, wherein, comprising: a feature extraction module, configured to use a feature extraction network to extract features from a speech to be detected to obtain an encoded feature vector of the speech to be detected; wherein, the feature extraction network is obtained by adjusting parameters of the feature extraction network using a first loss result of unsupervised training speech; the calculation steps of the first loss result include: calculating predicted feature vectors for the first step to the preset number of steps in the future based on the decoded feature vector of the unsupervised training speech; performing loss calculation on the encoded feature vector and the predicted feature vector corresponding to each step to obtain a second loss result corresponding to each step; statistically summing all the second loss results to obtain the first loss result; wherein, the encoded feature vector is obtained by performing convolutional processing on the unsupervised training speech using an encoder module of the feature extraction network, and the decoded feature vector is obtained by performing convolutional processing on the encoded feature vector using a decoder module of the feature extraction network; a speech detection module, configured to use a deep residual convolutional network to perform dimensionality reduction processing on the encoded feature vector of the speech to be detected to obtain a characterization vector of the speech to be detected, wherein the characterization vector contains speech category discrimination information of the speech to be detected; a determination module, configured to determine the speech category of the speech to be detected according to the similarity between the characterization vector and the target vector.
10. An electronic device, wherein, the electronic device includes a memory and a processor coupled to each other, the memory is used to store program data, and the processor is used to execute the program data to implement the method according to any one of claims 1-8.
11. A computer-readable storage medium, wherein, the computer-readable storage medium stores program data, and when the program data is executed by a processor, it is used to implement the method according to any one of claims 1-8.
Citation Information
Patent Citations
Acoustic event detecting method under hospital noise environment
CN108648748A
Voice generation method and device
CN110930976A
Voice classification method and device based on voiceprint recognition and related equipment
CN113436634A
Speech classification network training method and device, computing equipment and storage medium
CN113593611A