Method, device, electronic equipment and storage medium for assessing voice security
By performing feature extraction and attention weighting operations on speech signals, the problem of low accuracy in evaluating low-quality speech signals in existing technologies is solved, and more accurate speech signal security evaluation is achieved.
Patent Information
- Application Number
- CN202210966288.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-12
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2042-08-12
AI Technical Summary
In the prior art, for speech signals with low quality, the average quality of the speech signal calculated over time is used for prediction. However, the accuracy of the assessment using the speech quality calculated over time is low, resulting in low accuracy in the security assessment of low-quality speech signals.
By extracting features from the speech signal to be evaluated, a feature vector of a preset length is obtained. The attention weights corresponding to the feature vectors are then obtained, and the attention weights are weighted to obtain a weighted average feature vector. The security evaluation score of the speech signal is then obtained based on the weighted average feature vector.
It enables more accurate security assessment of low-quality voice signals, improving the accuracy and security of the assessment, and enhancing the accuracy and reliability of the assessment.
Smart Images

Figure CN115359806B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of voice security, for example, to a method and device for evaluating voice security, an electronic device and a storage medium. BACKGROUND
[0002] In the process of network transmission, the security of voice signals is one of the main indicators for measuring the system and service of a telecommunications network provider. In the related art, the convolutional neural network is usually used to extract the perceptual features of the voice signal, and then the average quality of the voice signal calculated over time and the extracted perceptual features are used to evaluate the security of the voice signal.
[0003] In the process of implementing the embodiments of the present disclosure, it is found that at least the following problems exist in the related art:
[0004] Since the quality of some voice signals is low, that is, there may be noise in the voice signal, or there may be distortion in the voice signal due to encryption. If the average quality of the voice signal calculated over time is still used to predict the overall quality of the low-quality voice signal, the predicted overall quality deviates greatly from the true quality of the voice signal. Thus, the accuracy of the security evaluation of the low-quality voice signal using the average quality is low. SUMMARY
[0005] To provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. The summary is not an overall description of the application, nor is it intended to determine key / important elements or delineate the scope of these embodiments, but as a prelude to the detailed description below.
[0006] The embodiments of the present disclosure provide a method for evaluating voice security, a model and a method for training a model for evaluating voice security, so as to more accurately evaluate the security of low-quality voice signals.
[0007] In some embodiments, the method for evaluating voice security comprises: performing feature extraction on a voice signal to be evaluated to obtain a feature vector of a preset length; obtaining an attention weight corresponding to the feature vector; performing a weighting operation on the attention weight to obtain a weighted average feature vector; and obtaining a security evaluation score of the voice signal to be evaluated according to the weighted average feature vector.
[0008] In some embodiments, the device for evaluating speech security comprises: a feature extraction module configured to perform feature extraction on a speech signal to be evaluated to obtain a feature vector of a preset length; a first acquisition module configured to acquire an attention weight corresponding to the feature vector; a weighting module configured to perform a weighting operation on the attention weight to obtain a weighted average feature vector; and a second acquisition module configured to acquire a security evaluation score of the speech signal to be evaluated according to the weighted average feature vector.
[0009] In some embodiments, the electronic device comprises a processor and a memory storing program instructions, and the processor is configured to execute the method for evaluating speech security when the program instructions are run.
[0010] In some embodiments, the storage medium stores program instructions, and the program instructions are executed to perform the method for evaluating speech security.
[0011] The method, device, electronic device and storage medium for evaluating speech security provided by the embodiments of the present disclosure can achieve the following technical effects: the feature vector is obtained by performing feature extraction on the speech signal to be evaluated, and then the attention weight corresponding to the feature vector is acquired, so that speech signals of different qualities have different attention weights. The weighted average feature vector is obtained by performing a weighting operation on the attention weight, so that the overall quality of the speech signal to be evaluated can be more accurately obtained. Then, the security evaluation score of the speech signal to be evaluated is acquired according to the weighted average feature vector, so that the security evaluation score of the speech signal to be evaluated can be derived using the more accurate overall quality of the speech signal to be evaluated, that is, the low-quality speech signal can be more accurately evaluated for security.
[0012] The foregoing general description and the following description are merely exemplary and explanatory, and are not intended to limit the present application. BRIEF DESCRIPTION OF DRAWINGS
[0013] One or more embodiments are exemplarily illustrated by the corresponding drawings, which do not constitute a limitation on the embodiments, and elements with the same reference numerals in the drawings are shown as similar elements, the drawings do not constitute a proportional limitation, and wherein:
[0014] Figure 1 is a schematic diagram of a method for evaluating speech security provided by the embodiments of the present disclosure;
[0015] Figure 2 is a schematic diagram of segment processing on a speech signal to be evaluated provided by the embodiments of the present disclosure;
[0016] Figure 3is a schematic diagram of feature extraction on a voice signal to be evaluated provided by an embodiment of the present disclosure.
[0017] Figure 4 is a schematic diagram of attention pooling on a voice signal to be evaluated provided by an embodiment of the present disclosure.
[0018] Figure 5 is a schematic diagram of a device for evaluating voice security provided by an embodiment of the present disclosure.
[0019] Figure 6 is a schematic diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0020] In order to enable a more detailed understanding of the features and technical content of the embodiments of the present disclosure, the implementation of the embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings, which are only used for reference and do not limit the embodiments of the present disclosure. In the following technical description, in order to facilitate explanation, a plurality of details are provided to provide a sufficient understanding of the disclosed embodiments. However, one or more embodiments can still be implemented without these details. In other cases, well-known structures and devices can be simplified to facilitate the drawings.
[0021] The terms "first", "second", and the like in the specification and claims of the embodiments of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances to implement the embodiments of the present disclosure described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion.
[0022] Unless otherwise specified, the term "a plurality of" means two or more.
[0023] In the embodiments of the present disclosure, the character " / " represents an "or" relationship between the objects before and after it. For example, A / B represents: A or B.
[0024] The term "and / or" is a description of the association between objects, which means that there can be three relationships. For example, A and / or B, which means: A or B, or, A and B, three relationships.
[0025] The term "corresponding" can refer to an association or binding relationship. A and B correspond to each other means that there is an association or binding relationship between A and B.
[0026] In combination Figure 1 As shown, the embodiments of the present disclosure provide a method for evaluating voice security, comprising:
[0027] In step S101, the electronic device performs feature extraction on the speech signal to be evaluated to obtain a feature vector of a preset length.
[0028] In step S102, the electronic device obtains an attention weight corresponding to the feature vector.
[0029] In step S103, the electronic device performs weighting operation on the attention weight to obtain a weighted average feature vector.
[0030] In step S104, the electronic device obtains a security evaluation score of the speech signal to be evaluated according to the weighted average feature vector.
[0031] By using the method provided in the embodiments of the present disclosure, the feature vector is obtained by performing feature extraction on the speech signal to be evaluated, and then the attention weight corresponding to the feature vector is obtained. In this way, speech signals of different qualities have different attention weights. Then, the weighted average feature vector is obtained by performing weighting operation on the attention weight, so that the overall quality of the speech signal to be evaluated can be more accurately obtained. Then, the security evaluation score of the speech signal to be evaluated is obtained according to the weighted average feature vector, that is, the security evaluation score of the speech signal to be evaluated is obtained by using the more accurate overall quality, so that the low-quality speech signal can be more accurately evaluated.
[0032] Optionally, the feature extraction on the speech signal to be evaluated to obtain the feature vector of the preset length comprises: zero padding the speech signal to be evaluated, performing Fourier transform on the speech signal to be evaluated after zero padding to obtain a first mel spectrum segment, segmenting the first mel spectrum segment according to a preset width and a preset height to obtain at least one second mel spectrum segment, and inputting the second mel spectrum segment into a preset deep feedforward neural network to obtain the feature vector of the preset length. Since the mel scale in the mel spectrum is designed for human ears, and the mel frequency is in a linear relationship with the normal frequency in the low frequency stage. At the same time, due to the characteristics of weak perception ability of human ears in the high frequency stage, the mel frequency is in a logarithmic relationship with the normal frequency. Therefore, converting the speech signal into the mel spectrum can retain most of the information required by human acoustics to understand external sound signals to a great extent. Thus, the mel spectrum segment can extract rich and effective perceptual features.
[0033] In some embodiments, the window length of the Fourier transform is set to 20 milliseconds, the jump size between windows is 10 milliseconds, and the maximum frequency is selected to be 20 kHz.
[0034] Optionally, the first mel-frequency spectrum segment is segmented according to a preset width and a preset height to obtain at least one second mel-frequency spectrum segment, including: segmenting the first mel-frequency spectrum segment with a size of 150 milliseconds in width and 48 milliseconds in height, and setting a hop size between the mel-frequency spectrum segments to 40 milliseconds to obtain the at least one second mel-frequency spectrum segment.
[0035] In some embodiments, the feature of the input second mel-frequency spectrum segment is extracted by a deep feedforward neural network with six convolutional layers in a frame-by-frame manner. That is, the second mel-frequency spectrum segment with a dimension of 150 milliseconds in width and 48 milliseconds in height is input into the deep feedforward neural network, and the second mel-frequency spectrum segment with a dimension of 48x15 is reduced to a mel-frequency spectrum segment with a dimension of 6x3 through downsampling operation, and then finally reduced to a mel-frequency spectrum segment with a dimension of 6x1, wherein the number of sampling kernels is 64. That is, the second mel-frequency spectrum segment with a dimension of 48x15 is input into the deep feedforward neural network to obtain a feature vector with a length of 384.
[0036] In some embodiments, the attention weight corresponding to the feature vector is obtained, including: refining the feature vector again by a Self-Attention self-attention network structure based on a Transformer encoder to obtain the attention weight corresponding to the feature vector. In this way, by utilizing the interaction information of the time steps of the feature vector, the attention weight corresponding to the feature vector is obtained, which can express the feature vector more accurately. That is, by utilizing the self-attention mechanism to perform time pooling operation on the feature vector, different quality speech signals have different attention weights. That is, the speech signals with poor quality in the speech signal to be evaluated have different weights from others. Compared with the prior art which simply uses the average quality of the speech signal calculated over time to perform safety evaluation on the speech signal, the overall quality of the speech signal to be evaluated can be more accurately predicted, so that the speech signal to be evaluated can be more accurately evaluated by using the overall quality. Further, the disclosed embodiment adopts a single-head attention mechanism, the depth of the deep feedforward neural network is set to 2, the model dimension of the deep feedforward neural network is 64, and the deep feedforward neural network has 64 hidden units.
[0037] Optionally, the attention weight is weighted to obtain a weighted average feature vector, including: normalizing the attention weight corresponding to the feature vector except for zero padding to obtain a normalized attention weight; and multiplying the normalized attention weight and the feature vector in matrix to obtain the weighted average feature vector.
[0038] In some embodiments, the softmax() function is used to normalize the attention weight. Further, before normalizing the attention weight by using the softmax() function, it further includes masking the time steps padded with zero.
[0039] Optionally, the security evaluation score of the speech signal to be evaluated is obtained according to the weighted average feature vector, including: inputting the weighted average feature vector into a preset full connection layer to obtain the security evaluation score of the speech signal to be evaluated. That is, the mapping from the weighted average feature vector to the overall score is realized through the full connection layer, and finally the security evaluation score of the speech signal to be evaluated is obtained.
[0040] In some embodiments, the electronic device is a computer, a mobile phone or a tablet computer, etc. The model for evaluating speech security provided in the electronic device is used to perform security evaluation on the speech signal to be evaluated. The model for evaluating speech security first performs zero padding on the input speech signal to be evaluated, then performs Fourier transform on the speech signal to be evaluated after zero padding to obtain a first mel spectrum segment. Then the first mel spectrum segment is input into a deep feedforward neural network with a 6 convolution layer, and the features of the first mel spectrum segment are extracted in a frame-by-frame manner to obtain a feature vector with a length of 384. Then the feature vector is input into a Self-Attention network structure based on a Transformer encoder to further refine the feature vector and obtain the attention weight corresponding to the feature vector. Then the time steps padded with zeros are shielded, and the attention weight is normalized by using a softmax() function to obtain the normalized attention weight; then the normalized attention weight is multiplied by the feature vector to obtain a weighted average feature vector. Finally, a full connection layer is used to realize the mapping from the weighted average feature vector to the overall score, and the security evaluation score of the speech signal to be evaluated is obtained. In this way, the deep feedforward neural network is first used to extract rich and effective perceptual features from the first mel spectrum segment, and then the time pooling operation based on the attention mechanism is used to predict the quality of the input speech signal, so that the overall quality of the speech signal to be evaluated can be more accurately obtained. Finally, the speech signal to be evaluated is evaluated according to the more accurate overall quality of the speech signal to be evaluated and the rich and effective perceptual features, so that the security evaluation score obtained is more accurate.
[0041] Further, the model for evaluating speech security is obtained by: obtaining training samples in a preset training set, the training samples being first speech signals with security evaluation label scores. The training samples are input into a convolutional neural network for training to obtain a candidate model for evaluating speech security. The candidate model for evaluating speech security is used to perform security evaluation on a second speech signal with a security evaluation label score to obtain a security evaluation score of the second speech signal; in converging to a minimum value, the candidate model for evaluating speech security is determined as the model for evaluating speech security; wherein n represents the number of mel spectrum segments of the second speech signal, The safety evaluation score of the second speech signal is y, the safety evaluation label score of the second speech signal is y, and the mean square error (MSE) is used to measure the safety evaluation score The mean square error of the difference between the safety evaluation label score y and the safety evaluation score MSE is the loss value. Due to the characteristics of MSE, as the error decreases, the gradient also decreases, which is conducive to accelerating the convergence of the network model. That is, even if a fixed learning rate is used, using MSE can quickly converge and reach a minimum value. Therefore, by selecting MSE to define the objective function, the model for evaluating speech safety can be obtained quickly. In the process of constructing the model for evaluating speech safety, the training samples in the training set are used only once. The training set is divided into the same sub-training set for batch optimization, which is called minibatches. Further, the training set is optimized in batches using the ADAM (Adaptive Moment Estimation) algorithm instead of the conventional SGD (Stochastic Gradient Descent) algorithm. Since the adaptive ability of the ADAM algorithm is better than that of the SGD algorithm, using the ADAM algorithm to optimize the training set in batches can accelerate the convergence speed of the objective function.
[0042] Further, the optimal network parameter is obtained by θ * = argmin θ Loss MSE . Wherein, θ * is the optimal network parameter, θ is the network parameter, and the network parameter has several. argmin θ Loss MSE indicates converges to a minimum value. That is, in the case of converges to a minimum value, the optimal network parameter can be obtained.
[0043] In some embodiments, in combination with Figures 2 to 4 as shown, Figure 2 is a schematic diagram provided by the embodiments of the present disclosure for segmenting the speech signal to be evaluated. The speech signal to be evaluated in Figure 2 is subjected to Fourier transform to obtain the first mel spectrum segment in Figure 2 . The first mel spectrum segment is segmented with a size of 150 milliseconds in width and 48 milliseconds in height, and the jump size between the mel spectrum segments is set to 40 milliseconds to obtain a plurality of second mel spectrum segments in Figure 2 . Figure 3is a schematic diagram provided by an embodiment of the present disclosure for feature extraction on a voice signal to be evaluated. The plurality of second mel-spectral segments in Figure 2 are input into a deep feedforward neural network, i.e., the feature frame extraction model shown in Figure 3 is used to obtain the feature vector of the second mel-spectral segment, i.e., the frame feature shown in Figure 3 is obtained. Figure 4 is a schematic diagram provided by an embodiment of the present disclosure for attention pooling on a voice signal to be evaluated. The frame feature shown in Figure 3 is input into the attention pooling module in Figure 4 to obtain a weighted average feature vector, and then the weighted average feature vector is passed through a fully connected layer to obtain a security evaluation score Score of the voice signal to be evaluated.
[0044] In some embodiments, the attention pooling module includes a Self-Attention network architecture based on a Transformer encoder, a masking operation unit, a normalization unit, and a matrix multiplication unit. The Self-Attention network architecture based on the Transformer encoder is used to receive the frame feature output by the deep feedforward neural network and obtain the attention weight of the frame feature. The masking operation unit is used to mask the time steps filled with zeros. The normalization unit is used to perform a normalization operation on the attention weights corresponding to the feature vectors other than the zero padding to obtain normalized attention weights, and send the normalized attention weights to the matrix multiplication unit. The matrix multiplication unit is used to perform a weighting operation on the normalized attention weights under the condition that the normalized attention weights are received, i.e., the frame feature output by the deep feedforward neural network and the normalized attention weights are multiplied by a matrix to obtain a weighted average feature vector.
[0045] In some embodiments, the method for evaluating speech safety provided by the embodiments of the present disclosure is verified. The monotonicity of the method for evaluating speech safety provided by the embodiments of the present disclosure is evaluated by using SRCC (Spearman Rank Correlation Coefficient) and KRCC (Kendall Rank Correlation Coefficient). The monotonicity is used to measure the correlation between the prediction result of the method for evaluating speech safety and the subjective result of human hearing. The accuracy of the method for evaluating speech safety provided by the embodiments of the present disclosure is evaluated by using PLCC (Pearson Linear Correlation Coefficients) and RMSE (Root Mean Squared Error). The accuracy is used to measure the degree of conformity between the test result of the method for evaluating speech safety and the subjective result of human hearing. Among them, SRCC is a non-parametric index for measuring the dependence of two variables, and the closer the value of SRCC is to 1, the stronger the rank correlation of the two variables. KRCC is a statistical value for measuring the correlation characteristics between two random variables, and the closer the absolute value of KRCC is to 1, the stronger the rank correlation between the two variables. PLCC is used to determine whether there is a linear correlation between two variables, and its value is between -1 and 1.
[0046] In the first optional embodiment, in the case that the prediction result of the method for evaluating speech safety is a safety evaluation score, the subjective result of human hearing is a safety evaluation label score.
[0047] In the case that SRCC is used to measure the correlation between the safety evaluation score and the safety evaluation label score, Among them, S = {s1, s2, s3,..., sn} is a set of safety evaluation label scores Score' of distorted audios in the audio data test set, n represents the total number of distorted audios, E = {e1, e2, e3,..., en} is a set of safety evaluation scores Score obtained by using the method for safety evaluation provided by the present application to evaluate the speech signals in the audio data test set, S n = {s1', s2', s3',..., sn'} is a set of safety evaluation label scores Score' of distorted audios in the audio data test set, n represents the total number of distorted audios, E n = {e1, e2, e3,..., en} is a set of safety evaluation scores Score obtained by using the method for safety evaluation provided by the present application to evaluate the speech signals in the audio data test set, S R = {s1', s2', s3',..., sn'} is a set of safety evaluation label scores Score' of distorted audios in the audio data test set, n represents the total number of distorted audios, E r1 = {e1, e2, e3,..., en} is a set of safety evaluation scores Score obtained by using the method for safety evaluation provided by the present application to evaluate the speech signals in the audio data test set, S r2 = {s1', s2', s3',..., sn'} is a set of safety evaluation label scores Score' of distorted audios in the audio data test set, n represents the total number of distorted audios, E r3 = {e1, e2, e3,..., en} is a set of safety evaluation scores Score obtained by using the method for safety evaluation provided by the present application to evaluate the speech signals in the audio data test set, S rn = {s1', s2', s3',..., sn'} is a set of safety evaluation label scores Score' of distorted audios in the audio data test set, n represents the total number of distorted audios, E R = {e1, e2, e3,..., en} is a set of safety evaluation scores Score obtained by using the method for safety evaluation provided by the present application to evaluate the speech signals in the audio data test set, S r1 = {s1', s2', s3',..., sn'} is a set of safety evaluation label scores Score' of distorted audios in the audio data test set, n represents the total number of distorted audios, E r2 = {e1, e2, e3,..., en} is a set of safety evaluation scores Score obtained by using the method for safety evaluation provided by the present application to evaluate the speech signals in the audio data test set, Sr3 ,..., e rn} is a set of grade divisions of the set of Score scores, that is, the Score' scores in the set S are graded from 1 to n, and after the grading is completed, the grades of the same Score' scores are averaged to obtain the set of grade divisions S R and E R . s ri is the i-th grade in the set of grade divisions S R , and e ri is the i-th grade in the set of grade divisions E R . is the average value of the elements in the set S R , and is the average value of the elements in the set E R . In the case where the KRCC is used to measure the correlation between the safety evaluation scores and the safety evaluation label scores, where n c is the logarithmic number of harmonies, n d is the logarithmic number of disharmonies, and n represents the total number of distorted audios. In some embodiments, let S = {s1, s2, s3,..., s n} be the set of safety evaluation label scores Score' scores of distorted audios in the audio data test set. E = {e1, e2, e3,..., e n} be the set of safety evaluation scores Score obtained by performing safety evaluation on the speech signals in the audio data test set using the method for safety evaluation provided in the present application. Take [(s i , e i ), (s j , e j )] and i≠j. In the case where s i >e i and s j >e j , or s i <e i and s j <e j , determine that [(s i , e i ), (s j , e j )] is harmonious. In the case where s i >e i and s j <e j , or s i <e i and s j >e j , determine that [(s i , ei ), (s j e j )] is discordant. Among them, s i For the i-th security assessment label score Score' in the set S of security assessment label scores, s j Let S be the j-th security assessment label score in the set S of security assessment label scores. i Let e be the i-th security assessment score in the set of security assessment scores E. j Let S be the j-th security assessment score in the set of security assessment scores E. Using PLCC and RMSE to measure the degree of conformity between the security assessment score and the security assessment label score, we utilize... For the set E = {e1, e2, e3, ..., e...} n The scores in} are processed to obtain the fitted predicted score set E′={e′1,e′2,e′3,...,e′ n This allows for more accurate Pearson linear correlation coefficients and root mean square errors. Where e i Let Score be the i-th security assessment score in the set E of security assessment scores. λ1 represents the predicted score after fitting, λ2 represents the first parameter to be fitted, λ3 represents the third parameter to be fitted, λ4 represents the fourth parameter to be fitted, and λ5 represents the fifth parameter to be fitted. Where S = {s1, s2, s3, ..., s} n} represents the set of security assessment label scores (Scores) for distorted audio data in the audio data test set. σ S Let σ be the standard deviation of set S. E′ Let S be the standard deviation of set E′, and COV(S, E′) characterize the covariance of sets S and E′. Represents the true value e′ of the sample i Compared with the predicted value s of the sample i The square root of the ratio of the square of the bias to the total sample size n. Where, e′ i Let be the i-th fitted prediction score in set E′.
[0048] In some embodiments, the test dataset includes a P501 dataset, a LiveTalk dataset, and a For dataset. The P501 dataset contains simulated distortions under different codecs, background noises, and clipping conditions, and contains a total of 240 audio files, which are derived from audio recordings of 4 volunteers in 60 different environments. The LiveTalk dataset is derived from real-time conversational audio recordings of 8 volunteers in 58 different environments, such as shopping malls, elevators, subway stations, and inside a moving car, etc. It contains a total of 232 audio segments with a length of 6 to 12 seconds. The For dataset also contains distortions under different codecs, background noises, and clipping conditions, and contains a total of 240 audio files, but it is different from the P501 dataset in that the P501 dataset is mainly recorded in real-time conditions of network phone Skype, multi-person mobile cloud video conference software Zoom, communication tool WhatsApp, and mobile network recording. The For dataset is mainly recorded in real-time conditions of communication tool WhatsApp, multi-person mobile cloud video conference software Zoom, and mobile game social network platform Discord. In addition, the For dataset is derived from audio recording segments of 80 volunteers in 60 environments. Meanwhile, each voice data in the P501 dataset, the LiveTalk dataset, and the For dataset corresponds to a security evaluation label score Score’, such as a security evaluation label score.
[0049] In some embodiments, as shown in Table 1, the Spearman rank correlation coefficient, the Kendall rank correlation coefficient, the Pearson linear correlation coefficient and the root mean square error obtained when the method Proposed for security evaluation provided by the embodiments of the present disclosure is used to perform security evaluation on the For data set are 0.7938, 0.5926, 0.8207 and 1.2611 respectively. The Spearman rank correlation coefficient, the Kendall rank correlation coefficient, the Pearson linear correlation coefficient and the root mean square error obtained when the existing method P563 for evaluating speech security is used to perform security evaluation on the For data set are 0.0925, 0.0753, 0.1347 and 2.1910 respectively. The Spearman rank correlation coefficient, the Kendall rank correlation coefficient, the Pearson linear correlation coefficient and the root mean square error obtained when the existing method WEnets for evaluating speech security is used to perform security evaluation on the For data set are 0.7055, 0.5067, 0.7055 and 1.5644 respectively. The Spearman rank correlation coefficient, the Kendall rank correlation coefficient, the Pearson linear correlation coefficient and the root mean square error obtained when the method Proposed for security evaluation provided by the embodiments of the present disclosure is used to perform security evaluation on the LiveTalk data set are 0.7553, 0.5723, 0.7849 and 1.3964 respectively. The Spearman rank correlation coefficient, the Kendall rank correlation coefficient, the Pearson linear correlation coefficient and the root mean square error obtained when the existing method P563 for evaluating speech security is used to perform security evaluation on the LiveTalk data set are 0.1802, 0.1450, 0.2271 and 2.1624 respectively. The Spearman rank correlation coefficient, the Kendall rank correlation coefficient, the Pearson linear correlation coefficient and the root mean square error obtained when the existing method WEnets for evaluating speech security is used to perform security evaluation on the LiveTalk data set are 0.5405, 0.3833, 0.5687 and 1.8533 respectively. The Spearman rank correlation coefficient, the Kendall rank correlation coefficient, the Pearson linear correlation coefficient and the root mean square error obtained when the method Proposed for security evaluation provided by the embodiments of the present disclosure is used to perform security evaluation on the P501 data set are 0.9064, 0.7316, 0.9047 and 1.0037 respectively. The Spearman rank correlation coefficient, the Kendall rank correlation coefficient, the Pearson linear correlation coefficient and the root mean square error obtained when the existing method P563 for evaluating speech security is used to perform security evaluation on the P501 data set are 0.1499, 0.1210, 0.1010 and 2.3441 respectively. The Spearman rank correlation coefficient, the Kendall rank correlation coefficient, the Pearson linear correlation coefficient and the root mean square error obtained when the existing method WEnets for evaluating speech security is used to perform security evaluation on the P501 data set are 0.7251, 0.5221, 0.7571 and 1.5393 respectively.It can be seen that when the method for evaluating speech security provided in the embodiments of the present disclosure is used to perform security evaluation on the For dataset, the LiveTalk dataset and the P501 dataset, the obtained SRCC value, KRCC value and PLCC value are all greater than the SRCC value, KRCC value and PLCC value obtained when the existing method for evaluating speech security, i.e., the method P563 and WEnets, is used to perform security evaluation on the For dataset, the LiveTalk dataset and the P501 dataset. At the same time, when the method for evaluating speech security provided in the embodiments of the present disclosure is used to perform security evaluation on the For dataset, the LiveTalk dataset and the P501 dataset, the obtained RMSE value is less than the RMSE value obtained when the method P563 and WEnets are used to perform security evaluation on the For dataset, the LiveTalk dataset and the P501 dataset. Since a method for evaluating speech with excellent performance should have relatively higher SRCC value, KRCC value and PLCC value, and relatively lower RMSE value. Therefore, the method for evaluating speech security provided in the embodiments of the present disclosure can more accurately perform security evaluation on low-quality speech signals compared with the method P563 and WEnets. At the same time, in the process of verifying the performance of the method for evaluating speech provided in the embodiments of the present disclosure by using the test dataset, for the For dataset and the P501 dataset, the performance indicators SRCC values of the two methods reach 0.7938 and 0.9064 respectively. For the LiveTalk dataset, the performance indicator SRCC value reaches 0.7553. It can be seen that the method for evaluating speech security provided in the embodiments of the present disclosure can perform security evaluation on simulated distortion and real distortion speech signals. Moreover, it can more accurately perform security evaluation on simulated distortion speech signals.
[0050]
[0051] Table 1
[0052] Further, the first speech signal also has a discontinuity label score, a loudness label score, a noise label score and a voice color label score. The training sample is input into the convolutional neural network for training to obtain a candidate model for evaluating speech security, and the candidate model for evaluating speech security is used to perform security evaluation on the second speech signal with the security evaluation label score to obtain a security evaluation score of the second speech signal; in After convergence to the minimum value, and confirming the candidate model for evaluating speech security, the process further includes: using the speech security evaluation model to perform discontinuity, loudness, noise, and timbre evaluations on the speech signal under evaluation, obtaining discontinuity evaluation scores, loudness evaluation scores, noise evaluation scores, and timbre evaluation scores. This allows for not only security evaluation of the speech signal under evaluation but also discontinuity, loudness, noise, and timbre evaluations, thus providing a more comprehensive reflection of the applicability of the model for evaluating speech security. By prioritizing different aspects, such as loudness or timbre, the degree of content leakage in the speech signal can be determined, broadening the application scenarios of the model for evaluating speech security.
[0053] In a second alternative embodiment, where the prediction result of the method used to assess speech security is a discontinuity assessment score, the subjective human auditory result is a discontinuity label score. When SRCC is used to measure the correlation between the discontinuity assessment score and the discontinuity label score, Where A = {a1, a2, a3, ..., a...} n Let} be the set of discontinuous label scores for distorted audio in the audio data test set, where n represents the total number of distorted audio, and H = {h1, h2, h3, ..., h...} n This refers to the set of discontinuity assessment scores obtained by using the security assessment method provided in this application to assess the discontinuity of speech signals in an audio data test set. R ={a r1 a r2 a r3 , ..., a rn Let H be the set of gradations for the set of discontinuous label scores. R ={h r1 h r2 h r3 , ..., h rn Let A be the grade set for a set of discontinuous assessment scores. Specifically, discontinuous label scores within set A are graded from 1 to n. After grading, the grades of scores with the same discontinuous label are averaged to obtain the final grade set A. R and H R a ri Let A be the set of levels. R The i-th level, h ri For the set of levels H R The i-th level in the series. For A R The average value of the elements in the set. For H RThe average of the elements within a set. In the case where the KRCC is used to measure the correlation between discontinuity evaluation scores and discontinuity label scores, where n k is the logarithmic number of harmonious, n g is the logarithmic number of disharmonious, and n represents the total number of distorted audios. In some embodiments, let A = {a1, a2, a3,..., a n} be the set of discontinuity label scores of distorted audios in the test set of audio data. H = {h1, h2, h3,..., h n} be the set of discontinuity evaluation scores obtained by performing discontinuity evaluation on the speech signals in the test set of audio data using the method for security evaluation provided in the present application. Take [(a i , h i ), (a j , h j )] and i≠j. In the case where a i >h i and a j >h j , or in the case where a i <h i and a j <h j , determine [(a i , h i ), (a j , h j )] to be harmonious. In the case where a i >h i and a j <h j , or in the case where a i <h i and a j >h j , determine [(a i , h i ), (a j , h j )] to be disharmonious. Wherein a i is the i-th discontinuity label score in the set of discontinuity label scores A, a j is the j-th discontinuity label score in the set of discontinuity label scores A. h i is the i-th discontinuity evaluation score in the set of discontinuity evaluation scores H, h j is the j-th discontinuity evaluation score in the set of discontinuity evaluation scores H. In the case where the PLCC and RMSE are used to measure the correlation between discontinuity evaluation scores and discontinuity label scores, use to measure the correlation between the set H = {h1, h2, h3,..., h nthe fitting of the prediction scores in the set H' = {h'1, h'2, h'3,..., h'N} is obtained. In this way, a more accurate Pearson linear correlation coefficient and root mean square error can be obtained. Wherein, h'1, h'2, h'3,..., h'N are the fitting prediction scores in the set H'. n} are obtained. In this way, a more accurate Pearson linear correlation coefficient and root mean square error can be obtained. Wherein, h'1, h'2, h'3,..., h'N are the fitting prediction scores in the set H'. i is the i-th discontinuity evaluation score in the set H of discontinuity evaluation scores, is the fitting prediction score, the parameter a1 is the sixth parameter to be fitted, the parameter a2 is the seventh parameter to be fitted, the parameter a3 is the eighth parameter to be fitted, the parameter a4 is the ninth parameter to be fitted, and the parameter a5 is the tenth parameter to be fitted. Wherein, A = {a1, a2, a3,..., aN} is the set of audio data test set discontinuity label scores. σ n} is the set of discontinuity label scores of distorted audio in the audio data test set. σ A is the standard deviation of the set A, σ H′ is the standard deviation of the set H', and COV(A, H') represents the covariance of the set A and the set H'. represents the true value h'N of the sample. i is the predicted value aN of the sample. i is the square root of the ratio of the square of the deviation to the total number of samples n. Wherein, h'N is the true value of the sample. i is the i-th fitting prediction score in the set H'.
[0054] In some embodiments, each voice data in the P501 data set, the LiveTalk data set and the For data set corresponds to a discontinuity label score. As shown in Table 2, when the For data set is evaluated for discontinuity Discontinuity using the method for security evaluation provided by the embodiments of the present disclosure, the Spearman rank correlation coefficient, the Kendall rank correlation coefficient, the Pearson linear correlation coefficient and the root mean square error obtained are 0.8437, 0.6470, 0.8566, 1.4611, respectively. When the LiveTalk data set is evaluated for discontinuity using the method for security evaluation provided by the embodiments of the present disclosure, the Spearman rank correlation coefficient, the Kendall rank correlation coefficient, the Pearson linear correlation coefficient and the root mean square error obtained are 0.6506, 0.4725, 0.6937, 2.8873, respectively. When the P501 data set is evaluated for discontinuity using the method for security evaluation provided by the embodiments of the present disclosure, the Spearman rank correlation coefficient, the Kendall rank correlation coefficient, the Pearson linear correlation coefficient and the root mean square error obtained are 0.8333, 0.6417, 0.8348, 1.3925, respectively.
[0055]
[0056]
[0057] Table 2
[0058] In a third optional embodiment, in the case that the prediction result of the method for evaluating speech security is the loudness evaluation score, the human auditory subjective result is the loudness label score. In the case that SRCC is used to measure the correlation between the loudness evaluation score and the loudness label score, wherein B = {b1, b2, b3,..., b n} is the set of loudness label scores of the distorted audios in the audio data test set, n represents the total number of distorted audios, M = {m1, m2, m3,..., m n} is the set of loudness evaluation scores obtained by using the method for security evaluation provided in the present application to evaluate the loudness of the speech signals in the audio data test set. B R = {b r1 , b r2 , b r3 ,..., b rn} is the set of grade divisions of the set of loudness label scores, M R = {m r1 , m r2 , m r3 ,..., m rn} is the set of grade divisions of the set of loudness evaluation scores, i.e. the loudness label scores in the set B are divided into grades from 1 to n, and after the division, the grades of the same loudness label scores are averaged to obtain the set of grade divisions B R and M R . b ri is the i-th grade in the set of grade divisions B R , and m ri is the i-th grade in the set of grade divisions M R . is the average value of the elements in the set B R , is the average value of the elements in the set M R . In the case that KRCC is used to measure the correlation between the loudness evaluation score and the loudness label score, wherein no is the logarithmic number of harmony, n p is the logarithmic number of disharmony, and n represents the total number of distorted audios. In some embodiments, let B = {b1, b2, b3,..., b n} be the set of loudness label scores of the distorted audios in the audio data test set. M = {m1, m2, m3,..., m n} be the set of loudness evaluation scores obtained by using the method for security evaluation provided in the present application to evaluate the loudness of the speech signals in the audio data test set. Any of [(b i , mi ), (b j m j And i≠j. In b i >m i And b j >m j In the case of, or, b i <m i And b j <m j In the case of [(b)], determine i m i ), (b j m j [] is harmonious. In b i >m i And b j <m j In the case of b i <m i And b j >m j In the case of [(b)], determine i m i ), (b j m j )] is discordant. Among them, b i Let b be the i-th loudness label score in the set of loudness label scores B. j Let be the j-th loudness label score in the set of loudness label scores B. i Let m be the i-th loudness assessment score in the set of loudness assessment scores M. j Let be the j-th loudness assessment score in the set M of loudness assessment scores. Using PLCC and RMSE to measure loudness assessment scores and loudness label scores, For set M = {m1, m2, m3, ..., m} n The scores in} are processed to obtain the fitted predicted score set M′={m′1,m′2,m′3,...,m′ n This allows for more accurate Pearson linear correlation coefficients and root mean square errors. Where m... i Let i be the loudness assessment score in the set M of loudness assessment scores. The parameter β1 represents the eleventh parameter to be fitted, β2 represents the twelfth parameter to be fitted, β3 represents the thirteenth parameter to be fitted, β4 represents the fourteenth parameter to be fitted, and β5 represents the fifteenth parameter to be fitted. Where B = {b1, b2, b3, ..., b} n} represents the set of loudness label scores for distorted audio data in the audio data test set. σ Bis the standard deviation of the set B M′ is the standard deviation of the set M', COV(B, M') represents the covariance of the set B and the set M'. represents the true value of the sample m' i is the predicted value of the sample b i is the square root of the ratio of the square of the deviation to the total number of samples n. Wherein, m' i is the i-th fitted prediction score in the set M'.
[0059] In some embodiments, each voice data in the P501 dataset, the LiveTalk dataset and the For dataset corresponds to a loudness label score. As shown in Table 3, when the For dataset is evaluated for loudness using the method for safety evaluation provided by the embodiments of the present disclosure, the Spearman rank correlation coefficient, the Kendall rank correlation coefficient, the Pearson linear correlation coefficient and the root mean square error obtained are 0.8014, 0.6305, 0.8793, 1.0630 respectively. When the LiveTalk dataset is evaluated for loudness using the method for safety evaluation provided by the embodiments of the present disclosure, the Spearman rank correlation coefficient, the Kendall rank correlation coefficient, the Pearson linear correlation coefficient and the root mean square error obtained are 0.6016, 0.4314, 0.6653, 1.3327 respectively. When the P501 dataset is evaluated for loudness using the method for safety evaluation provided by the embodiments of the present disclosure, the Spearman rank correlation coefficient, the Kendall rank correlation coefficient, the Pearson linear correlation coefficient and the root mean square error obtained are 0.8725, 0.7003, 0.9167, 0.9531 respectively.
[0060]
[0061] Table 3
[0062] In a fourth optional embodiment, in the case where the prediction result of the method for evaluating the safety of voice is a noise evaluation score, the subjective result of human hearing is a noise label score. In the case where SRCC is used to measure the correlation between the noise evaluation score and the noise label score, wherein F = {f1, f2, f3,..., f n} is the set of noise label scores of distorted audio in the audio data test set, n represents the total number of distorted audio, U = {u1, u2, u3,..., u n} is the set of noise evaluation scores obtained by evaluating the voice signal in the audio data test set using the method for safety evaluation provided by the present application. F R = {f r1 , f r2 , f r3 ,..., f rnLet U be the set of levels that categorizes the set of noise label scores. R ={u r1 u r2 u r3 , ..., u rn Let F be the set of levels for the noise assessment score set. Specifically, the noise label scores within set F are divided into levels from 1 to n. After the level division, the levels of the same noise label scores are averaged to obtain the final level set F. R and U R f ri For the set of levels F R The i-th level, u ri For the set of levels U R The i-th level in the series. For F R The average value of the elements in the set. For U R The average value of the elements within the set. In the context of KRCC being used to measure the correlation between noise assessment scores and noise label scores, Where nx is the number of harmonious logarithms, n y Let F be the logarithm of the number of dissonances, and n represent the total number of distorted audio frequencies. In some embodiments, let F = {f1, f2, f3, ..., f...} n Let U be the set of noise label scores for distorted audio data in the audio data test set. U = {u1, u2, u3, ..., u...} n} represents the set of noise assessment scores obtained by performing noise assessment on speech signals in an audio data test set using the security assessment method provided in this application. (f) i u i ), (f j u j And i≠j. In f i >u i And f j >u j In the case of, or, f i i And f j j In the case of [(f)] i u i ), (f j u j [] is harmonious. In f i >u i And f j j In the case of, or f i i And f j >uj In the case of [(f)] i u i ), (f j u j )] is discordant. Among them, f i Let f be the i-th noise label score in the set of noise label scores F. j Let be the j-th noise label score in the set of noise label scores F. i Let u be the i-th noise evaluation score in the noise evaluation score set U. j Let be the j-th noise assessment score in the noise assessment score set U. Using PLCC and RMSE to measure the noise assessment score and noise label score, [the following is used]. For the set U = {u1, u2, u3, ..., u...} n The scores in} are processed to obtain the fitted predicted score set U′={u′1,u′2,u′3,...,u′ n This allows for more accurate Pearson linear correlation coefficients and root mean square errors. Where, u i Let i be the i-th noise evaluation score in the noise evaluation score set U. The parameter δ1 represents the sixteenth parameter to be fitted, δ2 represents the seventeenth parameter to be fitted, δ3 represents the eighteenth parameter to be fitted, δ4 represents the nineteenth parameter to be fitted, and δ5 represents the twentieth parameter to be fitted. Where F = {f1, f2, f3, ..., f n} represents the set of noise label scores for distorted audio data in the audio data test set. σ F Let σ be the standard deviation of set F. U′ Let F be the standard deviation of set U′, and COV(F, U′) characterize the covariance of set F and set U′. Represents the true value u′ of the sample i Compared with the predicted value f of the sample i The square root of the ratio of the square of the bias to the total number of samples n. Where, u′ i Let be the i-th fitted prediction score in set U′.
[0063] In some embodiments, each voice data in the P501 dataset, the LiveTalk dataset and the For dataset corresponds to a noisiness label score. In combination with Table 4, the Spearman rank correlation coefficient, the Kendall rank correlation coefficient, the Pearson linear correlation coefficient and the root mean square error obtained when the For dataset is evaluated for noisiness using the method for safety evaluation provided in the embodiments of the present disclosure are 0.7604, 0.5765, 0.8032 and 1.2374, respectively. The Spearman rank correlation coefficient, the Kendall rank correlation coefficient, the Pearson linear correlation coefficient and the root mean square error obtained when the LiveTalk dataset is evaluated for loudness using the method for safety evaluation provided in the embodiments of the present disclosure are 0.7296, 0.5352, 0.8061 and 1.3238, respectively. The Spearman rank correlation coefficient, the Kendall rank correlation coefficient, the Pearson linear correlation coefficient and the root mean square error obtained when the P501 dataset is evaluated for loudness using the method for safety evaluation provided in the embodiments of the present disclosure are 0.8891, 0.7123, 0.9012 and 1.0895, respectively.
[0064]
[0065] Table 4
[0066] In a fifth optional embodiment, in the case where the prediction result of the method for evaluating voice safety is a color evaluation score, the human auditory subjective result is a color label score. In the case where SRCC is used to measure the correlation between the color evaluation score and the color label score, wherein Q = {q1, q2, q3,..., qn} is a set of color label scores of distorted audios in the audio data test set, n represents the total number of distorted audios, T = {t1, t2, t3,..., tn} is a set of color evaluation scores obtained by evaluating the voice signals in the audio data test set using the method for safety evaluation provided in the present application. Q n n is a set of color label scores of distorted audios in the audio data test set, n represents the total number of distorted audios, T = {t1, t2, t3,..., tn} is a set of color evaluation scores obtained by evaluating the voice signals in the audio data test set using the method for safety evaluation provided in the present application. Q R r1 r2 r3 rn is a set of color label scores of distorted audios in the audio data test set, n represents the total number of distorted audios, T = {t1, t2, t3,..., tn} is a set of color evaluation scores obtained by evaluating the voice signals in the audio data test set using the method for safety evaluation provided in the present application. Q R r1 r2 r3 rn is a set of color label scores of distorted audios in the audio data test set, n represents the total number of distorted audios, T = {t1, t2, t3,..., tn} is a set of color evaluation scores obtained by evaluating the voice signals in the audio data test set using the method for safety evaluation provided in the present application. Q R and TR q ri is the i-th level in the level partition set Q R t ri is the i-th level in the level partition set T R . is the average value of elements within the set Q R , and is the average value of elements within the set T R . In the case where KRCC is used to measure the correlation between the color evaluation scores and the color label scores, where nzis the logarithmic number of harmonies, n v is the logarithmic number of disharmonies, and n represents the total number of distorted audios. In some embodiments, let Q = {q1, q2, q3,..., q n} be the set of color label scores of distorted audios in the test set of audio data. T = {t1, t2, t3,..., t n} be the set of color evaluation scores obtained by performing color evaluation on the speech signals in the test set of audio data using the method for safety evaluation provided in the present application. For any [(q i , t i ), (q j , t j )] and i≠j. If q i > t i and q j > t j , or, q i < t i and q j < t j , then it is determined that [(q i , t i ), (q j , t j )] is harmonious. If q i > t i and q j < t j , or q i < t i and q j > t j , then it is determined that [(q i , t i ), (q j , t j )] is disharmonious. Where q i is the i-th color label score in the set of color label scores Q, q j is the j-th color label score in the set of color label scores Q. t iLet t be the i-th voice assessment score in the set of voice assessment scores T. j Let be the j-th timbre assessment score in the set T of timbre assessment scores. Using PLCC and RMSE to measure timbre assessment scores and timbre label scores, [the following is used]... For the set T = {t1, t2, t3, ..., t...} n The scores in} are processed to obtain the fitted predicted score set T′={t′1,t′2,t′3,...,t′ n This allows for more accurate Pearson linear correlation coefficients and root mean square errors. Where t i Let be the i-th sound and color assessment score in the set T of sound and color assessment scores. ε1 represents the predicted score after fitting, ε2 represents the twenty-first parameter to be fitted, ε3 represents the twenty-third parameter to be fitted, ε4 represents the twenty-fourth parameter to be fitted, and ε5 represents the twenty-fifth parameter to be fitted. Where Q = {q1, q2, q3, ..., q} n} represents the set of timbre label scores for distorted audio data in the audio data test set. σ Q Let σ be the standard deviation of set Q. T′ Let Q be the standard deviation of set T′, and COV(Q, T′) represent the covariance of sets Q and T′. Represents the true value t′ of the sample i With the predicted value q of the sample i The square root of the ratio of the square of the bias to the total sample size n. Where t′ i Let be the i-th fitted predicted score in set T′.
[0067] In some embodiments, each voice data in the P501 dataset, the LiveTalk dataset and the For dataset corresponds to a voice coloration label score. As shown in Table 5, when the For dataset is evaluated for voice coloration using the method for security evaluation provided by the embodiments of the present disclosure, the Spearman rank correlation coefficient, the Kendall rank correlation coefficient, the Pearson linear correlation coefficient and the root mean square error obtained are 0.8018, 0.5986, 0.8305 and 1.3047, respectively. When the LiveTalk dataset is evaluated for voice coloration using the method for security evaluation provided by the embodiments of the present disclosure, the Spearman rank correlation coefficient, the Kendall rank correlation coefficient, the Pearson linear correlation coefficient and the root mean square error obtained are 0.7726, 0.5751, 0.7915 and 1.2252, respectively. When the P501 dataset is evaluated for voice coloration using the method for security evaluation provided by the embodiments of the present disclosure, the Spearman rank correlation coefficient, the Kendall rank correlation coefficient, the Pearson linear correlation coefficient and the root mean square error obtained are 0.8248, 0.6250, 0.8457 and 1.2138, respectively.
[0068]
[0069] Table 5
[0070] In combination with Figure 5 As shown in Table 5, the embodiments of the present disclosure provide a device for evaluating the security of voice, which comprises a feature extraction module 501, a first acquisition module 502, a weighting module 503 and a second acquisition module 504. The feature extraction module 501 is configured to perform feature extraction on a voice signal to be evaluated, and obtain a feature vector of a preset length. The first acquisition module 502 is configured to acquire an attention weight corresponding to the feature vector. The weighting module 503 is configured to perform a weighting operation on the attention weight, and obtain a weighted average feature vector. The second acquisition module 504 is configured to acquire a security evaluation score of the voice signal to be evaluated according to the weighted average feature vector.
[0071] By using the device provided by the embodiments of the present disclosure, the feature vector is obtained by performing feature extraction on the voice signal to be evaluated, and then the attention weight corresponding to the feature vector is acquired, so that voice signals of different qualities have different attention weights. The weighted average feature vector is obtained by performing a weighting operation on the attention weight, that is, the overall quality of the voice signal to be evaluated can be more accurately obtained. Then, the security evaluation score of the voice signal to be evaluated is acquired according to the weighted average feature vector, that is, the security evaluation score of the voice signal to be evaluated is derived using the more accurate overall quality of the voice signal to be evaluated, so that the security evaluation of the voice signal of low quality can be more accurately performed.
[0072] Optionally, the feature extraction module is configured to perform feature extraction on the speech signal to be evaluated to obtain a feature vector of a preset length by performing zero padding on the speech signal to be evaluated, performing Fourier transform on the speech signal to be evaluated after zero padding to obtain a first mel-frequency spectrum segment, segmenting the first mel-frequency spectrum segment according to a preset width and a preset height to obtain at least one second mel-frequency spectrum segment, and inputting the second mel-frequency spectrum segment into a preset deep feedforward neural network to obtain the feature vector of the preset length.
[0073] Optionally, the weighting module is configured to perform weighting operation on the attention weight to obtain a weighted average feature vector by performing normalization operation on the attention weight corresponding to the feature vector other than the zero padding to obtain a normalized attention weight, and performing matrix multiplication on the normalized attention weight and the feature vector to obtain the weighted average feature vector.
[0074] Optionally, the second acquisition module is configured to acquire the security evaluation score of the speech signal to be evaluated according to the weighted average feature vector by inputting the weighted average feature vector into a preset fully connected layer to obtain the security evaluation score of the speech signal to be evaluated.
[0075] In combination Figure 6 As shown in the figure, the electronic device provided by the embodiment of the present disclosure includes a processor 600 and a memory 601. Optionally, the electronic device can further include a communication interface 602 and a bus 603. The processor 600, the communication interface 602, and the memory 601 can complete mutual communication through the bus 603. The communication interface 602 can be used for information transmission. The processor 600 can invoke the logical instructions in the memory 601 to execute the method for evaluating speech security of the above-mentioned embodiments.
[0076] By using the electronic device provided by the embodiment of the present disclosure, the feature vector is obtained by performing feature extraction on the speech signal to be evaluated, and then the attention weight corresponding to the feature vector is obtained, so that different quality speech signals have different attention weights. The weighted average feature vector is obtained by performing weighting operation on the attention weight, that is, the overall quality of the speech signal to be evaluated can be more accurately obtained. Then, the security evaluation score of the speech signal to be evaluated is acquired according to the weighted average feature vector, that is, the security evaluation score of the speech signal to be evaluated is derived by using the more accurate overall quality of the speech signal to be evaluated, so that the low-quality speech signal can be more accurately evaluated for security.
[0077] In addition, the logic instructions in the memory 601 described above can be implemented in the form of a software function unit and sold or used as an independent product, which can be stored in a computer readable storage medium.
[0078] The memory 601 as a computer readable storage medium can be used to store software programs, computer executable programs, such as program instructions / modules corresponding to the method in the embodiments of the present disclosure. The processor 600 executes the function application and data processing by running the program instructions / modules stored in the memory 601, that is, implements the method for evaluating voice security in the above embodiments.
[0079] The memory 601 can include a program storage area and a data storage area, wherein the program storage area can store an operating system and application programs required by at least one function; the data storage area can store data created according to the use of the terminal device, etc. In addition, the memory 601 can include a high-speed random access memory, and can also include a non-volatile memory.
[0080] The embodiments of the present disclosure provide a storage medium, which stores program instructions, and the program instructions execute the method for evaluating voice security when running.
[0081] The embodiments of the present disclosure provide a computer program product, which includes a computer program stored on a computer readable storage medium, and the computer program includes program instructions, and when the program instructions are executed by a computer, the computer executes the method for evaluating voice security.
[0082] The computer readable storage medium described above can be a transitory computer readable storage medium or a non-transitory computer readable storage medium.
[0083] The technical solutions of the embodiments of the present disclosure can be embodied in the form of a software product, and the computer software product is stored in a storage medium, and includes one or more instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the method described in the embodiments of the present disclosure. The aforementioned storage medium can be a non-transitory storage medium, including: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, etc. Various media that can store program codes, or a transitory storage medium.
[0084] The foregoing description and accompanying drawings fully illustrate embodiments of this disclosure to enable those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, procedural, and other changes. The embodiments represent only possible variations. Individual components and functions are optional unless explicitly required, and the order of operation may vary. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. Moreover, the terminology used in this application is for describing embodiments only and is not intended to limit the claims. As used in the description of embodiments and claims, the singular forms “a,” “an,” and “the” are intended to equally include the plural forms unless the context clearly indicates otherwise. Similarly, the term “and / or” as used in this application means including one or more of the associated listed items and all possible combinations thereof. Additionally, when used in this application, the term "comprise" and its variations "comprises" and / or "comprising" refer to the presence of stated features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof. Without further limitations, an element defined by the phrase "comprises a..." does not exclude the presence of other identical elements in the process, method, or apparatus that includes said element. In this document, each embodiment may focus on the differences from other embodiments, and similar or identical parts between embodiments can be referred to mutually. For methods, products, etc., disclosed in the embodiments, if they correspond to the method section disclosed in the embodiments, the relevant parts can be referred to the description of the method section.
[0085] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of this disclosure. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0086] The methods and products (including but not limited to devices and equipment) disclosed in the embodiments herein can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of units may be merely a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the shown or discussed units may be through some interfaces, and the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected to implement this embodiment according to actual needs. Furthermore, the functional units in the embodiments of this disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0087] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than that shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order than disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. Each block in a block diagram and / or flowchart, and combinations of blocks in a block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
Claims
1. A method for evaluating voice security, characterized in that, include: Feature extraction is performed on the speech signal to be evaluated to obtain a feature vector of a preset length; The feature vector is passed through a Self-Attention network structure based on a Transformer encoder to obtain the attention weights corresponding to the feature vector; The attention weights are weighted to obtain a weighted average feature vector; The security assessment score of the speech signal to be evaluated is obtained based on the weighted average feature vector. The process of weighting the attention weights to obtain a weighted average feature vector includes: normalizing the attention weights corresponding to the feature vectors other than those with zero padding to obtain normalized attention weights; and performing matrix multiplication between the normalized attention weights and the feature vectors to obtain the weighted average feature vector. The method further includes: acquiring training samples in a preset training set; inputting the training samples into a convolutional neural network for training to obtain a candidate model for evaluating speech security; using the candidate model for evaluating speech security to perform a security evaluation on a second speech signal with a security evaluation label score to obtain a security evaluation score for the second speech signal; When convergence reaches the minimum value, the candidate model for evaluating speech security is determined as the model for evaluating speech security; the model for evaluating speech security is used to perform discontinuity evaluation, loudness evaluation, noise evaluation and timbre evaluation on the speech signal to be evaluated, and obtain discontinuity evaluation score, loudness evaluation score, noise evaluation score and timbre evaluation score; The training samples consist of a first speech signal with security assessment label scores, including discontinuity label scores, loudness label scores, noise label scores, and timbre label scores; n represents the number of Mel-spectrum segments in the second speech signal. y represents the security assessment score of the second speech signal, and y represents the security assessment label score of the second speech signal. MSE is used to measure the security assessment score. The mean of the sum of squared differences between the safety assessment label score y and the safety assessment label score y, LOSS MSE This is the loss value.
2. The method according to claim 1, characterized in that, Feature extraction is performed on the speech signal to be evaluated to obtain a feature vector of a preset length, including: The speech signal to be evaluated is zero-filled; Perform a Fourier transform on the speech signal to be evaluated after zero padding to obtain the first Mel spectrum band; The first Mel spectrum segment is segmented according to a preset width and a preset height to obtain at least one second Mel spectrum segment. The second Mel spectrum band is input into a preset deep feedforward neural network to obtain a feature vector of a preset length.
3. The method according to claim 1, characterized in that, The security assessment score of the speech signal to be evaluated is obtained based on the weighted average feature vector, including: The weighted average feature vector is input into a preset fully connected layer to obtain the security assessment score of the speech signal to be evaluated.
4. An apparatus for evaluating voice security, characterized in that, include: The feature extraction module is configured to extract features from the speech signal to be evaluated and obtain a feature vector of a preset length. The first acquisition module is configured to pass the feature vector through a Self-Attention network structure based on a Transformer encoder to obtain the attention weights corresponding to the feature vector; The weighting module is configured to perform a weighting operation on the attention weights to obtain a weighted average feature vector. The second acquisition module is configured to acquire the security assessment score of the speech signal to be evaluated based on the weighted average feature vector. The weighting module is configured to perform a weighted operation on the attention weights in the following manner to obtain a weighted average feature vector: normalize the attention weights corresponding to the feature vectors other than zero-padding to obtain normalized attention weights; and perform matrix multiplication of the normalized attention weights with the feature vectors to obtain the weighted average feature vector. The device further includes: acquiring training samples in a preset training set; inputting the training samples into a convolutional neural network for training to obtain a candidate model for evaluating speech security; using the candidate model for evaluating speech security to perform a security evaluation on a second speech signal with a security evaluation label score to obtain a security evaluation score for the second speech signal; in When convergence reaches the minimum value, the candidate model for evaluating speech security is determined as the model for evaluating speech security; the model for evaluating speech security is used to perform discontinuity evaluation, loudness evaluation, noise evaluation and timbre evaluation on the speech signal to be evaluated, and obtain discontinuity evaluation score, loudness evaluation score, noise evaluation score and timbre evaluation score; The training samples consist of a first speech signal with security assessment label scores, including discontinuity label scores, loudness label scores, noise label scores, and timbre label scores; n represents the number of Mel-spectrum segments in the second speech signal. y represents the security assessment score of the second speech signal, and y represents the security assessment label score of the second speech signal. MSE is used to measure the security assessment score. The mean of the sum of squared differences between the safety assessment label score y and the safety assessment label score y, LOSS MSE This is the loss value.
5. The apparatus according to claim 4, characterized in that, The feature extraction module is configured to extract features from the speech signal to be evaluated in the following manner to obtain a feature vector of a preset length: The speech signal to be evaluated is zero-filled; Perform a Fourier transform on the speech signal to be evaluated after zero padding to obtain the first Mel spectrum band; The first Mel spectrum segment is segmented according to a preset width and a preset height to obtain at least one second Mel spectrum segment. The second Mel spectrum band is input into a preset deep feedforward neural network to obtain a feature vector of a preset length.
6. The apparatus according to claim 5, characterized in that, The second acquisition module is configured to acquire the security assessment score of the speech signal to be evaluated based on the weighted average feature vector in the following manner: The weighted average feature vector is input into a preset fully connected layer to obtain the security assessment score of the speech signal to be evaluated.
7. An electronic device comprising a processor and a memory storing program instructions, characterized in that, The processor is configured to, when running the program instructions, perform the method for evaluating voice security as described in any one of claims 1 to 3.
8. A storage medium storing program instructions, characterized in that, When the program instructions are executed, they perform the method for assessing voice security as described in any one of claims 1 to 3.
Citation Information
Patent Citations
Specific sound recognition method and device and storage medium
CN109074822A
Class cross entropy weighted intelligent construction noise assessment method based on pre-classification result
CN112837696A
Speech recognition method and device, equipment and storage medium
CN113838466A