A synthetic speech recognition method based on acoustic characteristics
By combining the DNN model of FFV, RMSA and SNS features, using the Spec-Attention module and the MDCN model, the problem of lack of a unified framework for synthetic speech recognition models is solved, and the acoustic features are efficiently integrated, which improves the recognition accuracy and computing efficiency.
Patent Information
- Application Number
- CN202211510517.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-29
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-11-29
AI Technical Summary
The existing synthetic speech recognition models lack a unified framework, and the complex and independent model construction lacks effective combinations between models, making it difficult to efficiently integrate high-value data in acoustic features and reduce noise interference.
FFV, RMSA and SNS features are combined with DNN models, and in-depth representation is performed through the Spec-Attention module and the MDCN model. The maximum dense convolutional neural network and attention mechanism are used to focus on high-value data in the acoustic features, reduce noise interference, and realize comprehensive identification of cross-modal data.
The accuracy of synthetic speech recognition is improved, and by analyzing the differences in acoustic characteristics, it provides a theoretical basis, enhances the model's perception of speech acoustic characteristics, reduces computing resource consumption, and improves recognition capabilities.
Smart Images

Figure CN115954016B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech recognition, and in particular, to a synthetic speech recognition method based on acoustic characteristics. Background Art
[0002] Currently, the mainstream recognition models are mainly DNN, CNN, and RNN types of models or models improved on the basis of the former. By using them independently or in combination, the purpose of recognizing synthetic speech can be achieved.
[0003] Some researchers, in addition to the above models, also utilize other optimized model technologies. By introducing the attention mechanism, the performance of existing time series models, convolutional network models, etc. can be improved to a certain extent, greatly improving the deficiencies of existing models, saving machine computing power, and reducing operation costs. For example, Li X et al. proposed the combination of the SE channel attention module and the Res2Net network model, and input CQT spectrogram features to carry out synthetic speech recognition; Ma X et al. proposed a lightweight convolutional network model combined with a four-layer CBAM attention module, and used spectrogram features for learning and training to further develop the utilization rate of spectrogram features. The research on attention algorithms has become one of the mainstream methods and research hotspots in the field of deep learning, and is a great leap forward in the development trend of neural network models gradually reaching a bottleneck.
[0004] The above are the research results of existing recognition models in the field of synthetic speech recognition, with a wide variety and advanced technologies. However, there is less combination and more independence among various models, lacking a unified framework that allows different models to be combined and used, and the model construction process is complex and the experimental difficulty is large.
[0005] Therefore, it is necessary to provide a new synthetic speech recognition method based on acoustic characteristics to solve the above technical problems. Summary of the Invention
[0006] The technical problem solved by the present invention is to provide a synthetic speech recognition method based on acoustic characteristics that can help a machine focus on high-value data in acoustic features, reduce interference from noise data irrelevant to acoustic characteristics, effectively integrate image-type data and time-series data in acoustic features, and utilize cross-modal data to comprehensively complete the target task to a certain extent.
[0007] To solve the above technical problems, the synthetic speech recognition method based on acoustic characteristics provided by the present invention includes the following steps:
[0008] S1: Connect the FFV feature and the RMSA feature, and then input them into the DNN model together;
[0009] S2: Use the DNN model to perform deep representation on it;
[0010] S3: Deeply represent the SNS features through the Spec-Attention module and the MDCN model in sequence;
[0011] S4: Connect the two and input them into the fully connected layer for binary classification output, and finally output real or synthetic.
[0012] Preferably, the Spec-Attention module first cuts the input SNS feature images in the frequency and segment directions respectively, that is, the segmentation images containing the harmonic morphological characteristics are obtained by frequency segmentation, and the segmentation images containing the phoneme spectrum distribution characteristics are obtained by segment segmentation. Subsequently, spatial and channel attention are calculated for each segmentation image and summed to obtain a single attention feature map, and after two connections, finally, an attention weight distribution feature map of the features is obtained.
[0013] Preferably, in the Spec-Attention module, the original image is first segmented, then input into the maximum and minimum pooling layers, and then input into the convolutional layer with a 7*7 convolutional kernel. After sigmoid function normalization, the dot product is performed with the original image data to output the attention feature map.
[0014] Preferably, the channel attention part in the Spec-Attention module uses the ECA module. After one-dimensional convolution and normalization of the segmentation image channels, selective attention is completed.
[0015] Preferably, the MDCN model improves the dense blocks in the dense neural network model, and incorporates the maximum feature mapping MFM1 / 2 operation after each layer of convolution in the dense block and the last layer of the transition layer to obtain the maximum dense block and the maximum transition layer.
[0016] Preferably, the inputs of the maximum dense block and the maximum transition layer first pass through a batch normalization layer, then through a Relu activation function, and are input into the 3*3 convolutional layer. Then, the MFM1 / 2 layer activates the key information of the features after convolution.
[0017] Preferably, first pass through a convolutional layer with 64 3*3 convolutional kernels and a stride of 1, then through a batch normalization layer and the maximum pooling layer and input into the densely connected maximum dense block; the MDCN model contains three maximum dense blocks and maximum transition layers, and each maximum dense block contains 3 convolutions and MFM1 / 2 calculations; after three dense connection calculations, then through a batch normalization layer and the global average pooling layer, and finally input into the fully connected layer to achieve binary classification output.
[0018] Preferably, the data characteristics required for the target task are obtained based on the experimental results of the previous chapter; according to the data characteristics required by the target, a specific algorithm is designed; the speech acoustic data is extracted; the acoustic data is processed using the specific algorithm; the data is transformed to highlight the key characteristics in the data, highlight the high-value part of the data, and weaken the redundant data, and finally the characteristics targeted for the synthetic speech recognition task are characterized.
[0019] Preferably, the acoustic features reflecting the intensity dispersion degree, fundamental frequency dispersion degree, and speech spectrum characteristics are selected, namely the root mean square energy angle feature RMSA, FFV, and SNS. Among them, the RMSA and FFV features are time-domain features containing timing information, and the SNS feature is a frequency-domain feature containing spectrum information.
[0020] Compared with the related technologies, the synthetic speech recognition method based on acoustic characteristics provided by the present invention has the following beneficial effects:
[0021] The present invention provides a synthetic speech recognition method based on acoustic characteristics, through the acoustic characteristics difference between the synthetic speech and the real speech. By comparing and analyzing the performance of the synthetic speech and the real speech in acoustic characteristics such as fundamental frequency, intensity, and spectrogram, analyzing the differences, and drawing regular conclusions, it explains the acoustic principle by which the synthetic speech can be recognized, and provides a theoretical basis for further automatic recognition.
[0022] Through the acoustic feature RMSA characterizing the intensity dispersion degree. This feature quantifies and characterizes the difference in the intensity change rate between the synthetic speech and the real speech, and after being fused with the FFV feature and the SNS feature, it is used as a high-dimensional feature input to the recognition model, providing a new feature design idea for synthetic speech recognition.
[0023] Through the maximum dense convolutional neural network model MDCN. When constructing the dense block of the dense convolutional neural network, the maximum feature mapping function is used. While retaining the dense connection of the model and reducing information forgetting, it also strengthens the effective information in the content learned by the convolutional neurons, providing a good model for improving the classification and recognition ability.
[0024] Through the attention module named Spec-Attention Block. According to the speech harmonic morphology and the distribution of the spectrum of a single phoneme, the narrowband spectrogram is segmented, and the refined segmentation results are selectively focused on from two dimensions of space and channel, making the model more focused on the harmonic position and spectrum breadth that can distinguish between synthetic and real speech, enhancing the model's perception of the speech acoustic characteristics, and further improving the recognition ability. Description of the Drawings
[0025] Figure 1 It is a principle block diagram of the synthetic speech recognition method based on acoustic characteristics provided by the present invention;
[0026] Figure 2 The process diagram of the characteristic acoustic features for the synthetic speech recognition method based on acoustic features provided by the present invention;
[0027] Figure 3 The filtering diagram of the filter bank for the synthetic speech recognition method based on acoustic features provided by the present invention;
[0028] Figure 4 The structure diagram of the dense block and transition layer in the MDCN model for the synthetic speech recognition method based on acoustic features provided by the present invention;
[0029] Figure 5 The schematic diagram of the Spec-Attention Block for the synthetic speech recognition method based on acoustic features provided by the present invention. Detailed implementation manners
[0030] The present invention will be further described below in conjunction with the accompanying drawings and implementation manners.
[0031] Please refer to Figure 1 、 Figure 2 、 Figure 3 、 Figure 4 and Figure 5 , where Figure 1 is the principle block diagram of the synthetic speech recognition method based on acoustic features provided by the present invention; Figure 2 is the process diagram of the characteristic acoustic features for the synthetic speech recognition method based on acoustic features provided by the present invention; Figure 3 is the filtering diagram of the filter bank for the synthetic speech recognition method based on acoustic features provided by the present invention; Figure 4 is the structure diagram of the dense block and transition layer in the MDCN model for the synthetic speech recognition method based on acoustic features provided by the present invention; Figure 5 is the schematic diagram of the Spec-Attention Block for the synthetic speech recognition method based on acoustic features provided by the present invention. The synthetic speech recognition method based on acoustic features: 1) Obtain the data characteristics required for the target task according to the experimental results of the previous chapter; 2) Design a specific algorithm according to the data characteristics required by the target; 3) Extract the speech acoustic data; 4) Process the acoustic data using the specific algorithm; 5) Transform the data to highlight the key characteristics in the data, highlight the high-value part of the data, and weaken the redundant data, and finally characterize the features targeted at the synthetic speech recognition task, such as Figure 2Therefore, according to the requirements of this experiment, the present invention designs and selects acoustic features that can reflect the discrete degree of sound intensity, the discrete degree of fundamental frequency, and the speech spectrum characteristics, namely the root mean square energy angle feature (Root Mean Square Angle, hereinafter referred to as RMSA for short), the fundamental frequency variation rate feature (Fundamental Frequency Variation, hereinafter referred to as FFV for short), and the speech narrowband spectrogram feature (Speech Narrowband Spectrogram, hereinafter referred to as SNS for short). Among them, the RMSA and FFV features are time-domain features containing timing information, and the SNS feature is a frequency-domain feature containing spectrum information. Further combining the three features will be more suitable for the synthetic speech recognition task. The specific feature extraction process is as follows.
[0032] The present invention proposes the RMSA feature to characterize the discreteness difference characteristics shown in the sound intensity between the synthetic speech and the real speech. The specific feature extraction process is as follows:
[0033] 1) Speech data acquisition. Input the speech, and extract the speech digital signal through sampling at 16000 Hz and 8-bit quantization.
[0034] 2) Calculate the root mean square energy (RMS) of the speech as shown in Equation 1. First, frame the speech signal. Each frame contains 2048 sampling points, and the overlapping part between frames contains 512 sampling points. Then calculate the root mean square energy of the speech signal for each frame. The specific process is to sum the squares of the signals within a period, divide by the time, and then take the square root as a whole.
[0035] 3) Vectorize the input data. To transform the one-dimensional time-series data into two-dimensional data, add the time point data to the original data as dimension one, and the value at that point as dimension two.
[0036] 4) Calculate the cosine distance between adjacent vectors;
[0037] 5) Finally, according to the calculated cosine distance d, obtain the cosine value of the angle, and use the inverse cosine function to calculate the corresponding angle degree to obtain the RMSA feature.
[0038] The reason why the present invention uses the calculation method of root mean square energy as a specific processing algorithm to characterize the sound intensity is as follows:
[0039] Sound intensity is an important indicator that describes the magnitude of sound energy in decibels from the perspective of human hearing. At the same time, the energy of speech signals includes both sound intensity and frequency intensity. For a periodically changing signal, due to its fluctuations, it is difficult to estimate the actual magnitude of the entire signal through statistical quantities such as the mean and amplitude. By calculating the root mean square (RMS) value of the periodically changing signal to represent the overall energy of the signal, it is equivalent to the energy of a steady signal with a constant value. For example, the alternating current in daily life has periodic changes, and its effective value is calculated to reflect the amount of current passing through, which is equivalent to the amount of direct current passing through at a certain fixed ampere. The root mean square (RMS) is a commonly used method for calculating the effective value. By calculating the RMS energy of speech signals, the true intensity of speech signals can be more accurately characterized. The energy of each frame of the periodically changing speech signal is extracted, and the RMS value of the speech energy of each frame can better characterize the energy of the signal in a short period of time. It can be seen that by calculating the RMS energy value, the degree of fluctuation of the speech signal is effectively characterized in the form of numerical changes, providing a calculation condition for the next step of extracting the sound intensity dispersion.
[0040] Considering that the difference in the acoustic characteristics between synthetic speech and real speech is reflected in the different degrees of dispersion of sound intensity, the present invention proposes to further calculate the cosine angle between adjacent data on the basis of the RMS energy. This can improve the data fineness, amplify the part with large differences between adjacent data, reduce the part with small differences, and reduce the smoothness of the data, thereby enhancing the characteristics of the data. This is because in the process of natural speech production, it often has large fluctuations and a strong sense of rhythm. The large fluctuations often have a greater impact on sound intensity, and the degree of expansion of the included angle between adjacent vectors is relatively large. While the sound intensity of synthetic speech tends to be more stable during pronunciation, which makes the change in the included angle between two adjacent vectors smaller.
[0041] Therefore, by calculating the included angle between two adjacent vectors, it can be used to measure the difference between data points and quantify the degree of fluctuation reflected in the acoustic characteristics of sound intensity in speech.
[0042] The intensity of real speech will change accordingly with the pronunciation of human beings. The sound intensity will also increase or decrease significantly, resulting in a larger overall degree of sound intensity dispersion. Its essence is caused by the change of speech energy. By extracting the RMS value of speech signal energy and calculating the difference between adjacent two frames, to a certain extent, it can reflect the change of acoustic characteristics of speech and the state of the speaker during pronunciation. Therefore, extracting the RMSA feature of speech to characterize the sound intensity dispersion can extract the instantaneous change of sound intensity from the perspective of speech acoustic characteristics, which is conducive to distinguishing between synthetic speech and real speech and helps to improve the accuracy of automatic recognition of synthetic speech.
[0043] The discreteness of the fundamental frequency refers to the fluctuation range of the fundamental frequency values, which will be characterized using the FFV feature in prosodic features. The FFV feature represents the instantaneous change of the pitch frequency between consecutive frames in the form of a vector, and can better reflect the change of the fundamental frequency data. In addition, different from the conventional method that only uses the fundamental frequency to calculate the change amount by taking the difference of data at time points, this feature also proposes the concept of vanishing point inner product, calculates the Fourier transform on the original signal data, so as to achieve the purpose of simultaneously using the fundamental frequency difference between harmonics between frames. This feature can not only reflect the pitch fluctuation degree in acoustics, but also be well applicable to the synthetic speech recognition task, and is an ideal acoustic feature. The design concept of this feature is to find the best frequency amplitude expansion factor required for each of the adjacent two frames to align the harmonic spacing, calculate the similarity of the expansion factors within the adjacent two frames, so as to characterize the increase and decrease degree of the fundamental frequency. The overall steps are frame division and windowing, calculation of the fast Fourier transform, calculation of the vanishing point inner product, and FFV filter processing.
[0044] The specific extraction steps are as follows:
[0045] 1) First, perform frame division and windowing on the speech signal. That is, the speech signal is segmented into short segments with a duration of 32 milliseconds for each frame on the time axis, and a new frame is obtained by overlapping a part of the adjacent two frames. The repetition rate of frame overlapping is 75%, that is, 24 milliseconds are repeated, so as to improve the continuity and relevance between data.
[0046] 2) Subsequently, each frame of the speech signal is processed using a window function. The window functions used are the Hann window and the Hamming window. In the present invention, the speech signal is successively processed through the Hann window and the Hamming window functions. This approach is essentially to multiply each specific value in each frame of the speech signal by a weight w[n], so as to avoid the redundant high-frequency components caused by the discontinuous points generated by frame division processing. Windowing is beneficial to improving the continuity of the signal after frame division. The window function is as shown in Formula 4, where when α = 0.5, it is the Hann window. When α = 0.46, it is the Hamming window. N is the length of the discrete data after sampling of the speech signal, and n is each data point.
[0047] 3) Perform a fast Fourier transform (Fast Fourier transform, FFT) on the data within the window. The data of each frame is divided into two windows FL and FR on the left and right. Each window is 20 milliseconds wide, the data sorting is mirror images of each other, and the peak interval time is 12 milliseconds. A fast Fourier transform with N = 512 will be used for each window. Among them, the fast Fourier transform is a faster way to calculate the one-dimensional discrete Fourier transform, which can more efficiently separate the discrete one-dimensional speech signal data according to different frequency bands, and finally the spectral representation form of the speech signal can be obtained.
[0048] 4) Perspective projection, vanishing - point inner product. Perform vanishing - point projection on the left and right windows after fast Fourier transform and calculate the scaling factor. Specifically, it means performing perspective projection on the spectra at the - T0 and T0 instants within the same frame through the position of a point τ (which can be positive or negative) on the coordinate axis. These two instants are taken from the FL and FR windows respectively. When the point τ approaches positive or negative infinity, the projections emitted from this point will tend to be parallel infinitely, becoming parallel projection, which is equivalent to performing orthogonal projection on vectors in the vector space, equivalent to performing standard dot - product on the parallel data on two vectors.
[0049] Different from the standard dot - product corresponding to orthogonal projection, the vanishing - point inner product refers to the process of performing projection on the - T0 and T0 instants using a finite point τ, and then calculating the sum of dot - products of the scaled points. It is to perform dot - product summation after scaling the data in the two windows proportionally, obtaining the compressed windows FL and FR or compressed FR and FL. The calculation of the spectral scaling is as shown in Equation 5. This formula uses the method of linear interpolation to achieve the estimation of discrete data, and k takes all integer values in the range from - 255 to 256.
[0050] The concept of the vanishing point comes from the principle of geometric perspective. Because in two - dimensional plane vision, observing two sets of line segments in the same direction and parallel to each other in the same space will extend and converge to a point, which is called the vanishing point. Since the vanishing point can fix the direction of all line segments in the same direction in the picture, and under the limitation of this point, all line segments are proportional to each other. Therefore, this property of the vanishing point can be used to obtain the scaling ratio of data at different times in the same frame, so as to realize obtaining the expansion factor in the data space by using the way of spatial deformation. The expansion factor can be used to characterize the change rate of the current frame.
[0051] 5) Obtain the expansion factor and get the FFV spectrum. Obtain the expansion factors that align with the harmonic spacing at all times to characterize the change rate of a frame of speech. Accumulate the results of the vanishing - point products of FL and the scaled FR (vanishing point τ < t0) or FR and the enlarged FL (vanishing point τ > t0) at the current time, and normalize the accumulated result, that is, divide it by the square root of the sum of their squares, to obtain the expansion factor of the signal at the t0 moment. Take the expansion factor with the largest value among all times to represent the fundamental - frequency change rate of the current frame. This is because in the current frame, the expansion factor with the largest value can reflect the two times with the largest scaling change among all times. The specific calculation process is as shown in Equation 6, F* means the conjugate - complex form of the vector. Where ρ is the expansion coefficient, determined by the position of the point τ, N = 512, and r takes values from - 255 to 255. When the value of τ is less than t0, ρ < 0, and when the value of τ is greater than t0, ρ > 0.
[0052] After the above operations, substituting the data of all frames into the operations can obtain the maximum dilation factors of all frames, and then calculating the cosine distance of the dilation factors between adjacent two frames, the change rate of the fundamental frequency between frames can be characterized by the similarity degree of the dilation factors between frames.
[0053] 6) Process using a filter bank (FFVfilterbank) to obtain the final fundamental frequency change rate feature. The filter bank is as Figure 3 shown, which contains 7 types of filters, multiplying the fundamental frequency change spectrum with different weight values respectively, and finally retaining the data of the first three segments, the middle segment and the last three segments, reducing the fundamental frequency change spectrum from 512 dimensions to 7 dimensions.
[0054] The fundamental frequency dispersion of real speech is greater than that of synthetic speech. Therefore, by extracting the FFV feature (FFV) of the speech signal to characterize the fundamental frequency dispersion degree, the differences in the fundamental frequency acoustic characteristics between synthetic speech and real speech can be reflected. This feature extracts the instantaneous change situation of the fundamental frequency from the perspective of speech acoustic characteristics, which helps to distinguish synthetic speech and real speech and improve the recognition accuracy for further carrying out deep learning experiments.
[0055] The present invention uses the method of extracting SNS features
[60] to characterize the differences in spectral characteristics between synthetic speech and real speech. The specific steps for extracting SNS features are as follows:
[0056] 1) First, perform frame windowing processing on the speech signal. That is, each frame is 32 milliseconds, and the frame repetition rate is 20%, and it is processed by Hanning and Hamming window functions in sequence.
[0057] 2) Then perform discrete Fourier transform (Discrete Fourier Transform, DFT) on each frame of the signal. That is, the original discrete sampling signal is transformed through discrete Fourier transform to obtain the frequency domain characteristics of the speech signal.
[0058] Finally, speech spectrum data is obtained. Subsequently, by plotting the data on each pixel and using heat maps with different color shades to characterize the frequency intensity of different regions, a characterization image containing the frequency in terms of time, position, and intensity can be obtained.
[0059] Directly extracting SNS features and using the machine to directly learn and recognize the input narrowband spectrum image can more intuitively learn the differences in spectral characteristics between synthetic speech and real speech. Compared with other methods, it has the advantages of being more efficient and intuitive, and also facilitates improving the interpretability of the automatic recognition process.
[0060] Based on the dense convolutional neural network model (DenseNet), the present invention incorporates the MFM operation to obtain the maximum dense convolutional neural network model (MDCN). The reasons for such a design are as follows:
[0061] First, by using the dense connection method of the convolutional layer, the forgetting of SNS feature information by the model as the convolutional layer deepens is reduced, and the loss of key data in the SNS features is avoided, thereby improving the utilization efficiency of the SNS features.
[0062] Second, by using the MFM operation after each convolutional layer, the less valuable and irrelevant redundant data information in the SNS features is further suppressed, the number of model parameters is reduced, the model is lightweighted, and the training speed of complex multi-dimensional acoustic features is improved.
[0063] The present invention designs a MaxDense Convolutional Neural Network Model (MaxDenseConvolutionNet, hereinafter referred to as MDCN). Its main design is to improve the dense block in the Dense Neural Network Model (DenseNet), and integrate the maximum feature mapping MFM1 / 2 operation after each layer of convolution in the dense block and the last layer of the transition layer to obtain the DenseMaxblock and the TransitionMaxBlock, thereby strengthening the model's ability to suppress non-critical information and further lightweighting the model structure.
[0064] The structures of the DenseMaxblock and the TransitionMaxBlock designed by the present invention are as Figure 4 shown. The input features first pass through a batch normalization layer, then through a Relu activation function, and are input into a 3*3 convolutional layer. Then, the MFM1 / 2 layer is used to activate the key information of the features after convolution.
[0065] The structures of the dense block and the transition layer in DenseNet can be seen in Figure 4 (a) of. Comparing with (b), it can be seen that:
[0066] 1) The present invention replaces the Relu activation function after convolution in the dense block of DenseNet by using the MFM1 / 2 operation. This can further suppress the neurons that learn non-critical information in the SNS features and activate the neurons containing key information.
[0067] This is because the process of using the Relu activation function to transform the neuron output is essentially to map to the function space of Formula 9 according to the numerical size, and still retains the non-critical data in the original data.
[0068] The MFM1 / 2 operation, on the other hand, directly selects the maximum value at the same position in the feature map as the output, thus only retaining the most valuable data in the neuron output.
[0069] 2) The present invention replaces the max pooling layer in the transition layer of DenseNet by using the MFM1 / 2 operation. This can reduce the loss of key information of features. This is because the essence of the max pooling layer is to utilize the locality and translational invariance of image data. The purpose is to gradually reduce the pixel dimension and aggregate important information through the operation of finding the maximum value of the corresponding positions of local pixels in the image. However, this process only retains the maximum value of local pixels at the corresponding positions, which will cause some key data of SNS features to be omitted.
[0070] The MFM1 / 2 operation selects the maximum value in the channel dimension as the output, without changing the position information of the image data, and retains the key data at each position in the neuron output. At the same time, it is more suitable for SNS features because the SNS features contain key information in the spatial position, such as harmonic morphology and spectral distribution position.
[0071] The input features of the MDCN network model first pass through a convolutional layer with 64 3*3 convolutional kernels and a stride of 1, and then pass through a batch normalization layer and a max pooling layer and are input into the maximally dense block with dense connections; there are three maximally dense blocks and a max transition layer in the MDCN model, and each maximally dense block contains 3 convolutions and MFM1 / 2 calculations; after three dense connection calculations, it passes through a batch normalization layer and a global average pooling layer, and finally is input into the fully connected layer to achieve binary classification output.
[0072] The present invention further explores the performance of SNS features and focuses on the high-value data in acoustic features. By designing a Spec-Attention module targeted at acoustic features on the basis of the existing visual attention, selective attention is respectively performed on data in different frequency bands and different segments of SNS features. The reasons for such design are as follows:
[0073] First, by introducing the visual attention mechanism into the convolutional neural network, the learning object of the model can be made more targeted, enabling the model to focus on the key information in the image and pay more attention to the more representative data in the SNS of acoustic features, such as harmonic morphology and spectral distribution in the SNS features, and improving the deficiencies of the existing model. It can greatly reduce the consumption of GPU resources, more efficiently complete the operation tasks of huge amounts of image data, and to a certain extent, it can reveal the black box of the neural network model and help people understand the principle of machine learning.
[0074] Second, by introducing an attention mechanism in speech recognition, the connections between different SNS feature data can be discovered. It can improve the machine's perception of the arrangement characteristics in the speech time series in the form of coherent pronunciation, and can perceive the internal connections of speech and the pronunciation status of speech at a deeper and more multi-dimensional level, such as the spectral broadness of phonemes in SNS features, etc. Thus, it can strengthen the machine's learning of human pronunciation habits and speech coherence characteristics, discover the habitual pronunciation dynamic stereotype of the speaker, and find the internal characteristics of sounds, which helps to improve the machine recognition performance and accuracy.
[0075] Based on the above attention mechanism, the present invention combines the acoustic characteristics in SNS features, and designs a Spec-Attention module for the harmonic and segment spectral distribution characteristics in the narrowband spectrogram. Further improve the performance of the neural network model and the utilization rate of the acoustic feature SNS to achieve a further breakthrough in the accuracy of synthetic speech recognition.
[0076] 1) Main structure of the Spec-Attention module
[0077] In the design of the Spec-Attention module of the present invention, the input SNS feature images are first cut in the frequency and segment directions respectively, that is, the cut images containing the harmonic morphological characteristics are obtained by frequency segmentation, and the cut images containing the phoneme spectral distribution characteristics are obtained by segment segmentation. Subsequently, spatial and channel attention is calculated for each cut image and summed to obtain a single attention feature map, and then after two connections, a feature attention weight distribution feature map is finally obtained. As Figure 5 shown.
[0078] The reasons for such a design are as follows:
[0079] First, from the previous experiments, it can be seen that the SNS features contain the acoustic characteristic differences between synthetic speech and real speech, mainly the morphology of harmonics and the spectral data broadness. By dividing the SNS feature map into intervals of 1000 Hz in frequency, the model can focus on the important data in this area, reduce global calculations, thereby improving the attention focusing accuracy and making it easier to focus on the harmonic position.
[0080] Second, since the spectral broadness is essentially the frequency distribution of phonemes in a single segment, by dividing the SNS feature map into intervals of a single segment, the model can perform selective attention in this area, thus focusing on the spectral broadness of a single segment and grasping the important acoustic characteristics for identifying synthetic speech.
[0081] 2) Spatial attention module
[0082] In the Spec-Attention module designed by the present invention, spatial attention is calculated for the segmented image. Different from the design of using max pooling and average pooling in the CBAM module, in the present invention, the original image is input into the max pooling layer and the min pooling layer to extract the original image data, and then input into the convolutional layer with a 7×7 convolutional kernel. After normalization by the sigmoid function, the attention feature map is output. Such a design can expand the difference degree of image information data, so that the model can more easily focus on the data with higher value.
[0083] 3) Channel attention module
[0084] The channel attention module in the Spec-Attention module designed by the present invention uses the ECA module. After one-dimensional convolution and normalization processing on the channels of the segmented image, selective attention is completed.
[0085] 4) Parallel spatial and channel attention module
[0086] Different from the way of cascading spatial and channel attention in the CBAM module, in the present invention, the channel attention module and the spatial attention module are operated in parallel, and the results are then connected, so as to improve the utilization rate of SNS features and avoid ignoring key information in cascading connection.
[0087] The Spec-Attention module designed by the present invention can improve the model's learning ability for SNS feature data, strengthen the model's focus on more valuable data, further enhance the machine's perception of speech acoustic characteristics, and thus improve the performance of acoustic feature recognition and synthesized speech.
[0088] Due to the introduction of the attention mechanism, the model's learning ability for speech feature data can be greatly improved, and the machine's learning ability can be better focused on the data worthy of attention. Therefore, the present invention combines the above MDCN model with the Spec-Attention module and fuses SNS features, RMSA features, and FFV features to achieve the purpose of recognizing and synthesizing speech. The scheme for recognizing and synthesizing speech in the present invention is shown in Figure 1 .
[0089] The present invention first connects the FFV feature and the RMSA feature, and then inputs them into the DNN model together. The DNN model is used to perform deep representation on them, and then the SNS feature is successively subjected to deep representation through the Spec-Attention module and the MDCN model. Subsequently, the two are connected and input into the fully connected layer for binary classification output.
[0090] The specific parameter settings of the model are as follows: The DNN model has 5 hidden layers each with 2048 neurons; The Spec-Attention module cuts by frequency 10 times and by segment 5 times; The MDCN model has a total of 3 maximum dense blocks, and each dense block contains 3 convolutions and 3 MFM1 / 2 operations, and each convolution contains 32 3*3 convolution kernels.
[0091] Compared with the related technologies, the synthetic speech recognition method based on acoustic characteristics provided by the present invention has the following beneficial effects:
[0092] Through the acoustic characteristics differences between synthetic speech and real speech. By comparing and analyzing the performance of synthetic speech and real speech in terms of acoustic characteristics such as fundamental frequency, intensity, spectrogram, etc., analyzing the differences, drawing regular conclusions, explaining the acoustic principle by which synthetic speech can be recognized, and providing a theoretical basis for further automatic recognition.
[0093] Through the acoustic feature RMSA that characterizes the discrete degree of intensity. This feature quantifies and characterizes the difference in intensity change rate between synthetic speech and real speech, and after fusing with the FFV feature and SNS feature, it is used as a high-dimensional feature for inputting into the recognition model, providing a new feature design idea for synthetic speech recognition.
[0094] Through the maximum dense convolutional neural network model MDCN. When constructing the dense blocks of the dense convolutional neural network, the maximum feature mapping function is used, which not only retains the dense connection of the model and reduces information forgetting, but also strengthens the effective information in the content learned by the convolutional neurons, providing a good model for improving the classification and recognition ability.
[0095] Through the attention module named Spec-AttentionBlock. The narrowband spectrogram is segmented according to the harmonic morphology of the speech and the distribution of the spectrum of a single phoneme, and the selectively focused attention is carried out on the refined segmentation results from two dimensions of space and channel, making the model more focused on the harmonic position and spectral broadness that can distinguish synthetic and real speech, enhancing the model's perception of the acoustic characteristics of speech, and further improving the recognition ability.
[0096] The above are only the embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be included in the patent protection scope of the present invention by the same token.
Claims
1. A synthetic speech recognition method based on acoustic characteristics, characterized in that It includes the following steps: S1: Connect the FFV feature and the RMSA feature, and then input them into the DNN model together; S2: Use the DNN model to perform deep representation on it; S3: Represent the SNS feature deeply through the Spec-Attention module and the MDCN model in sequence; S4: Connect the deep representation in step S2 and the deep representation in step S3, input it into the fully connected layer for binary classification output, and finally output real or synthetic; The Spec-Attention module first cuts the input SNS feature image in the frequency and segment directions respectively, that is, it is cut by frequency to obtain a segmented image containing the harmonic morphological characteristics, and cut by segment to obtain a segmented image containing the phoneme spectral distribution characteristics. Then, spatial and channel attention is calculated for each segmented image and summed to obtain a single attention feature map. After two connections, a feature attention weight distribution feature map is finally obtained; The MDCN model improves the dense block in the dense neural network model, and incorporates the maximum feature mapping MFM1 / 2 operation after each layer of convolution in the dense block and the last layer of the transition layer to obtain the maximum dense block and the maximum transition layer.
2. The method for synthesizing speech recognition based on acoustic characteristics according to claim 1, wherein In the Spec-Attention module, the original image is first segmented, then input into the maximum and minimum pooling layers, and then input into the convolutional layer with a 7*7 convolutional kernel. After sigmoid function normalization, it is dot-producted with the original image data to output the attention feature map.
3. The method for synthetic speech recognition based on acoustic characteristics according to claim 1, wherein The ECA module is used in the Spec-Attention module. After one-dimensional convolution and normalization of the segmented image channels, selective attention is completed.
4. The method for synthetic speech recognition based on acoustic characteristics according to claim 1, wherein The input of the maximum dense block and the input of the maximum transition layer first pass through a batch normalization layer, then through a Relu activation function, and are input into the 3*3 convolutional layer. Then, the MFM1 / 2 layer activates the key information of the features after convolution.
5. The method for synthetic speech recognition based on acoustic characteristics according to claim 1, characterized in that, First, it passes through a convolutional layer with 64 3*3 convolutional kernels and a stride of 1, then through a batch normalization layer and the maximum pooling layer and is input into the densely connected maximum dense block; The MDCN model contains three maximum dense blocks and maximum transition layers. Each maximum dense block contains 3 convolutions and MFM1 / 2 calculations; After three dense connection calculations, it passes through a batch normalization layer and the global average pooling layer, and finally is input into the fully connected layer to achieve binary classification output.
Citation Information
Patent Citations
Parallel feature extraction system and method for general specific voice in voice signal
CN110992987A
Method and device for signal synthesis
JP1996123484A