Novel dynamic time delay neural network for dialect recognition

By introducing a weighted average mechanism of parallel convolution kernel and attention weight parameters in the delay neural network, the convolution kernel weights are dynamically adjusted, and the problem of low recognition accuracy in dialect recognition in dialect recognition is solved, achieving higher dialect type recognition accuracy and speech content recognition accuracy.

CN120108419AActive Publication Date: 2025-06-06BEIJING FANGWEI ZHILIAN TECHNOLOGY CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510255371.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-06
Estimated Expiration
2045-03-05

AI Technical Summary

Technical Problem

In the prior art, it is difficult to accurately identify dialects with rich tones and many regional variants, such as Minnan and Cantonese, in dialect recognition, resulting in a decrease in recognition accuracy.

Method used

A new dynamic delay neural network is adopted, and the respective attention weight parameters are calculated by introducing k parallel convolution kernels, and the weighted average fusion is used to obtain a convolution kernel with the optimal weight parameters. Replace the conventional convolution kernel of the existing delay neural network, and dynamically adjust the convolution kernel weight parameters to extract the deep-level features of dialect audio.

Benefits of technology

It improves the accuracy of dialect type recognition, can more accurately identify the voice content of different dialects, and enhances the performance of the acoustic model of delay neural networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108419A_ABST
    Figure CN120108419A_ABST
Patent Text Reader

Abstract

The invention discloses a novel dynamic time delay neural network for dialect recognition, and belongs to the field of acoustic model modeling. Introducing an existing time delay neural network into k parallel convolution kernels, calculating respective attention weight parameters, and carrying out weighted averaging and fusion to obtain a convolution kernel with an optimal weight parameter; the convolution kernel of the optimal weight parameter replaces a conventional convolution kernel of an existing time delay neural network to obtain an improved dynamic time delay neural network, according to different dialect inputs, each weight parameter in the convolution kernel of the optimal weight parameter is dynamically adjusted, deeper-level feature information of the audio is extracted, and the dynamic time delay neural network is obtained. The extracted depth features and a dialect acoustic template are compared and analyzed, and finally the dialect category is judged in combination with a classifier; according to the invention, the accuracy of determining the dialect type to which the audio belongs is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of acoustic model building, and in particular is a novel dynamic time-delay neural network for dialect recognition. Background Art

[0002] With the development of emerging technologies such as machine learning and artificial intelligence, Mandarin speech recognition technology is sufficient to meet the requirements of most application scenarios. However, in the context of cultural heritage and regional communication, dialects still dominate, and their complex speech features and lack of standardized data pose severe challenges to existing recognition systems. Especially when faced with dialects with rich tones and regional variants such as Minnan and Cantonese, the subtle differences in pronunciation often cause the accuracy of general speech recognition models to drop drastically.

[0003] In the current research on dialect recognition, there are two main methods:

[0004] One is phased dialect recognition: first identify which dialect the speaker is using, and then use a model specific to that dialect to determine the content of the speech; the advantage of this method is that it can take advantage of the specific features of each dialect, thereby improving the accuracy of recognition. In this way, the model can be optimized for the speech characteristics of different dialects.

[0005] The other is unified model training: directly use the data of all dialects to train a unified model, and then use this model to identify the content of the dialect; this method only needs to maintain one model instead of multiple dialect-specific models; it simplifies the management and deployment of the model; however, this method may not perform as well as specially trained models on the characteristics of some dialects.

[0006] Time-delay neural networks are mainly used to process one-dimensional time series tasks. Compared with traditional HMM time series models, they have better performance in acoustic modeling because of their advantage of fewer parameters. However, whether the dialect type can be accurately judged is particularly important in dialect recognition. The acoustic model modeled by ordinary time-delay neural networks cannot achieve a relatively high recognition rate for all dialect types at the same time due to their limited representation ability, which will lead to errors in the selection of dialect acoustic models and thus convert them into incorrect speech content. Based on this, the accuracy and precision of dialect content recognition can be improved by first identifying the type of dialect through the time-delay neural network model, and then further identifying the speech content of this specific dialect. Summary of the invention

[0007] Based on the above problems, the present invention constructs a new type of dynamic time-delay neural network, which improves the accuracy of dialect recognition and identifies the dialect type used by the speaker. Subsequently, for staged dialect recognition, an acoustic model specific to the dialect is used to determine the content of the speech.

[0008] The novel dynamic time-delay neural network for dialect recognition has the following specific structure: the convolution kernel of the existing time-delay neural network is updated: first, k parallel convolution kernels are introduced, the respective attention weight parameters are calculated, and then weighted average fusion is performed to obtain a convolution kernel with an optimal weight parameter; finally, the convolution kernel with the optimal weight parameter replaces the conventional convolution kernel of the existing time-delay neural network to obtain an improved dynamic time-delay neural network; when different dialects are input, the weight parameters in the convolution kernel with the optimal weight parameter are dynamically adjusted to extract deeper feature information of the audio, and the extracted deep features are compared and analyzed with the dialect acoustic template, and finally the dialect category is determined in combination with the classifier;

[0009] The specific update process is as follows:

[0010] Step 1: Perform global average pooling on the input audio features to obtain global spatial features of length C;

[0011] The input audio feature is a H×W×C tensor;

[0012] The global average pooling operation calculates the mean of each channel in its spatial dimension. The output of the global average pooling of the cth channel is GlobalFeature c The formula is as follows:

[0013]

[0014] Where H is the height of the feature map, W is the width of the feature map, and C is the total number of channels. i,j,c Represents the eigenvalue of the cth channel at position (i, j), c∈C.

[0015] Step 2: Map the global spatial features to k dimensions through two fully connected layers in sequence, perform softmax normalization, and obtain the weight coefficients of the k convolution kernels;

[0016] The formula is as follows:

[0017]

[0018] Among them, Z k is the kth output of the second fully connected layer, Softmax(Z k ) is the weight coefficient of the kth convolution kernel.

[0019] Step 3: Take the weighted average of all parallel convolution kernel weights according to their respective weight coefficients and fuse them to obtain a new convolution kernel.

[0020] The weight of each convolution kernel is W conv,k ∈R N×N×C , where N is the size of the convolution kernel;

[0021] Weighted average to get the new convolution kernel W new , the formula is:

[0022]

[0023] π k is the weight coefficient of the kth convolution kernel, W conv,k is the original weight of the kth convolution kernel.

[0024] Step 4: Input the new convolution kernel into the time-delay neural network model so that it can adaptively capture the deep features of the input dialect.

[0025] The output formula of the improved dynamic time-delay neural network model is:

[0026]

[0027] C t ={X t-n , X t+n}

[0028]

[0029] Y t ' is the output feature, f is the activation function, W T and b are fixed weight matrices and bias vectors, respectively, C t is the input feature, X t-n The number of upper information frames input to the lower network at time t is n; X t+n The number of context information frames input to the lower network at time t is n, where n is greater than or equal to 1;

[0030] The improved time-delay neural network model can adaptively select the optimal feature extraction mode according to the audio content of the input dialect: it strengthens the short-time impulse response in the transient stage of dialect pronunciation and enhances the resonance peak tracking ability in the steady-state stage.

[0031] Step 5: Compare and analyze the extracted deep features with the dialect acoustic template, and finally determine the dialect category by combining the classifier.

[0032] The advantages of the present invention are:

[0033] The present invention discloses a novel dynamic time-delay neural network for dialect recognition. The weight coefficient of each convolution kernel can be dynamically adjusted according to different dialect audio inputs, and an optimal convolution kernel parameter is obtained by weighted average, so as to extract deeper features of the dialect audio changes and improve the accuracy of dialect recognition. This operation not only does not bring additional depth or width to the time-delay neural network, but also greatly enhances the performance of the time-delay neural network acoustic model. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 It is a schematic diagram of a novel dynamic time-delay neural network for dialect recognition according to the present invention;

[0035] Figure 2 Schematic diagram of a typical time-delay neural network structure.

[0036] Figure 3 A schematic diagram of a specific implementation method for dynamically adjusting the structure of the present invention is provided.

[0037] Figure 4 A schematic diagram of the structure of the dynamic time-delay neural network of the present invention is inserted into the existing time-delay neural network structure. DETAILED DESCRIPTION

[0038] In order to facilitate those skilled in the art to understand and implement the present invention, the present invention is described in further detail and in depth with reference to the accompanying drawings.

[0039] In the time-delay neural network acoustic modeling model, the acoustic modeling effect is good. However, for the input of audio in different dialects, the time-delay neural network with fixed parameters may not have high recognition accuracy for some dialects. The acoustic model lacks certain generalization and flexibility, resulting in low recognition rate for some dialects.

[0040] Compared with the traditional time-delay neural network acoustic model, once the training is completed, all weight parameters are fixed. For any input, all parameters are treated equally, and the ability of the acoustic model to extract effective audio features is very limited. In order to improve the expressiveness of the acoustic model of the time-delay neural network and increase the flexibility of the acoustic model in representing features, the present invention improves the convolution method of the time-delay neural network in the architecture of the time-delay neural network, replaces a single convolution kernel operation with multiple parallel convolution kernels, and then obtains a group of convolution kernel coefficients, each convolution kernel has different convolution weight parameters, generates different convolution kernel weight parameters according to different inputs, and fuses multiple convolution kernels with fixed parameters into a dynamic kernel, which can capture deeper time-varying information in dialect audio and enhance the expressiveness of the acoustic model of dialect recognition.

[0041] The present invention proposes a novel dynamic time-delay neural network for dialect recognition, the specific structure of which is as follows:

[0042] The novel dynamic time-delay neural network for dialect recognition has the following specific structure: the convolution kernel of the existing time-delay neural network is updated: first, k parallel convolution kernels are introduced, the respective attention weight parameters are calculated, and then weighted average fusion is performed to obtain a convolution kernel with an optimal weight parameter; finally, the convolution kernel with the optimal weight parameter replaces the conventional convolution kernel of the existing time-delay neural network to obtain an improved dynamic time-delay neural network; when different dialects are input, the weight parameters in the convolution kernel with the optimal weight parameter are dynamically adjusted to extract deeper feature information of the audio, and the extracted deep features are compared and analyzed with the dialect acoustic template, and finally the dialect category is determined in combination with the classifier;

[0043] like Figure 1 As shown, the specific update process is as follows:

[0044] Step 1: Perform global average pooling on the input audio feature x to obtain a global spatial feature with a length of C;

[0045] The input audio feature is a H×W×C tensor;

[0046] The global average pooling operation calculates the mean of each channel in its spatial dimension. The output of the global average pooling of the cth channel is GlobalFeature c The formula is as follows:

[0047]

[0048] Where H is the height of the feature map, W is the width of the feature map, and C is the total number of channels. i,j,c Represents the eigenvalue of the cth channel at position (i, j), c∈C.

[0049] Step 2: Map the global spatial features to k dimensions through two fully connected layers in sequence, perform softmax normalization, and obtain the weight coefficients of the k convolution kernels;

[0050] The formula is as follows:

[0051]

[0052] Among them, Z k is the kth output of the second fully connected layer, Softmax(Z k ) is the weight coefficient of the kth convolution kernel.

[0053] Step 3: Take the weighted average of all parallel convolution kernel weights according to their respective weight coefficients and fuse them to obtain a new convolution kernel.

[0054] k parallel convolution kernels, each with a weight of W conv,k ∈R N×N×C, where N is the size of the convolution kernel; the weight coefficients of each convolution kernel are weighted averaged to obtain a new convolution kernel W new , the formula is:

[0055]

[0056] π k is the weight coefficient of the kth convolution kernel, and the set of weight coefficients corresponding to k parallel convolution kernels is {π 1 ,π 2 ,...,π k}; W conv,k is the original weight of the kth convolution kernel.

[0057] Step 4: Add the new convolution kernel W new The input is improved into the time-delay neural network model by dynamically changing the analysis period so that it can adaptively capture the deep features in the input dialect.

[0058] By dynamically changing the analysis period like adjusting the focus of a microscope: when encountering short dialect plosive sounds, "increase the magnification" for detailed observation; when encountering long tone changes, "zoom out" to grasp the overall trend.

[0059] In the structure of the common time-delay neural network, the output is calculated according to the following formula:

[0060] Y t =f(W T C t +b)

[0061] C t ={X t-n , X t+n}

[0062] Where X t-n The number of upper information frames input to the lower network at time t is n; X t+n The number of context information frames input to the lower network at time t is n, where n is greater than or equal to 1; C t is the input feature, Y t is the output feature, W T and b are the fixed weight matrix and bias vector respectively, and f is the activation function.

[0063] Traditional methods have obvious deficiencies in capturing features specific to each dialect, especially in temporal features with context. The weight matrix W with fixed parameters T It cannot adapt to the unique phoneme variation characteristics of each dialect. The output formula of the improved dynamic time delay neural network is:

[0064]

[0065]

[0066] Use k convolution kernels of the same size and use weight coefficient π k Combine them. It is worth noting that the weight coefficient π here k It is not fixed; instead, it changes according to the input data 'x', and aggregates these kernel coefficients into an optimal convolution kernel by averaging weighted aggregation. Since the convolution kernel parameters are not fixed according to the input, they tend to be dynamically adjusted according to the current dialect type. Finally, the features output by the acoustic model are more conducive to determining which dialect it is.

[0067] The convolution kernel coefficient parameters of the improved time-delay neural network model are not fixed, and the optimal feature extraction mode can be adaptively selected according to the audio content of the input dialect: the short-time impulse response is strengthened in the transient stage of dialect pronunciation, and the resonance peak tracking ability is enhanced in the steady-state stage.

[0068] Step 5: Compare and analyze the extracted deep features with the dialect acoustic template, and finally determine the dialect category by combining the classifier.

[0069] Traditional fixed convolution kernels use homogenized weight processing in time domain feature extraction, which is difficult to adapt to the time-varying characteristics of dialects. The convolution method proposed in the present invention generates a content-adaptive convolution kernel coefficient group by sensing the time-frequency feature differences of the input audio in real time, and adopts a weighted fusion strategy to achieve dynamic parameter adjustment. This design enables the convolution kernel of each time window to have feature sensitivity, and the model can adaptively select the optimal feature extraction mode according to the audio content: strengthen the short-time pulse response in the transient stage of dialect pronunciation (such as the ending of the glottal stop in Minnan dialect), and enhance the resonance peak tracking ability in the steady-state stage (such as the continuous voiced consonant in Wu dialect). This dynamic weight adjustment mechanism effectively captures the deep time-varying information such as tone glide and transition phonemes in dialect audio, so that the accuracy of dialect recognition tasks can be effectively improved.

[0070] Compared with the existing time-delay neural network model, the convolution method is improved: k parallel convolution kernels are used to replace the conventional single convolution kernel, each of which has different convolution weight parameters, and each convolution kernel coefficient is also different. The weight parameters of these parallel convolution kernels are weighted averaged according to the convolution kernel coefficients, and a new convolution kernel is obtained by fusion. Because the characteristics of the input audio are different, the obtained set of convolution kernel coefficients are not exactly the same, and the weight parameters of the convolution kernel weights after the final weighted average fusion are not exactly the same. It is realized that the convolution weight parameters can be dynamically adjusted according to the different inputs. By improving in this way, the model can more flexibly select the most relevant feature extraction method according to the input content, capture deeper time-varying information in the dialect audio, and thus improve the accuracy of the recognition task.

[0071] like Figure 2 As shown in the figure, it is a schematic diagram of a typical time-delay neural network structure. The input of each layer of the time-delay neural network is obtained through the context window of the lower layer, so it can describe the temporal relationship between the nodes of the upper and lower layers. After each layer is affine transformed, like DNN, there will be an activation function as the total output of the layer. This output is used as the extracted audio feature to determine the type of dialect to which the audio belongs.

[0072] For the acoustic model of the time-delay neural network after training, the weight parameters of each node in each layer are fixed. The audio judgment ability for different dialects may be different. The time-delay neural network essentially uses one-dimensional convolution operations to process time series data. Each layer can contain multiple convolution kernels to extract local features.

[0073] The present invention realizes dynamic adjustment of convolution parameters by improving the way of extracting features by convolution kernel. The present invention mainly improves the convolution kernel by replacing the current convolution kernel with multiple parallel convolution kernels. During the model training process, the parameters of multiple parallel convolution kernels are not the same. Finally, they are weighted averaged by different input weights to aggregate into an optimal convolution kernel.

[0074] like Figure 3 As shown in the figure, it is a schematic diagram of a specific implementation method for dynamic adjustment. k parameters are adjustable. When k is 1, it is no different from an ordinary time-delay neural network, and the output is calculated through a convolution. The dynamically adjusted time-delay neural network adjusts the current weight parameters continuously by adjusting the attention weights according to the different input audios, so as to extract deeper information from the audio features.

[0075] During the training process, assuming k is 2, there are two parallel convolution weight parameters, and each convolution weight parameter has an attention weight coefficient π k , the weight coefficient of the first convolution is π 1 , the weight coefficient of the second convolution is π 2 , π 1 Add π 2 =1, when the training audio is input, the neural network will continue to forward propagate and back propagate to update the convolution weight parameters. Because there are two parallel convolution kernels, the parameters in the two convolution kernels are also different during the training process. During the training process, it is assumed that the first convolution weight parameter is more suitable for extracting the audio features of dialect A, so π will be increased. 1 The weight coefficient weakens π 2 The weight coefficient of the convolution kernel is then weighted averaged. The extracted features after this operation are more conducive to judging dialect A. Assuming that the second convolution weight parameter is more suitable for extracting the audio features of dialect B, it will increase π 2 The weight coefficient weakens π1 The weight coefficient of , the extracted features are more conducive to judging that it is dialect B.

[0076] Here we encounter a problem: during the training process, the model cannot actively judge the tendency to increase π 1 Coefficient or π 2 The coefficient of π needs to be calculated using the attention mechanism k process, such as Figure 2 As shown in the dashed box, when audio features are input, the input audio features need to be globally averaged pooled (GlobalAvgPooling) to obtain global spatial features, and then mapped to k dimensions through two FC layers, and softmax normalized. The obtained k attention weights are assigned to the k convolution kernels of this layer.

[0077] During the model training process, the attention mechanism is also trained so that the model can dynamically adjust the convolution kernel coefficients according to the input. Figure 2 The attention mechanism in the model runs in parallel with the convolution and shares input with the convolution, which enables the model to dynamically adjust different convolution kernel parameters according to different input audio. The neural network will continuously optimize the parameters in convolution kernel 1 and convolution kernel 2 through the back-propagation mechanism, and will also optimize the parameters of FC in the attention mechanism, and the model training process is completed. However, this method is not limited to dynamic changes based on the input of different dialect audio, but in a piece of audio, different frames of audio can be used as input of the time-delay neural network for continuous dynamic adjustment and optimization of convolution weight parameters.

[0078] In the model inference stage, similar to the training process stage, there are also two parallel convolution kernels with fixed parameters. However, due to the existence of the attention mechanism, the weight coefficient will dynamically change in the range of 0 to 1 according to the different inputs, and then multiply it by the convolution kernel parameters to take a weighted average of the two convolution kernels. In this way, for different inputs, even if there are two convolution kernels with fixed parameters, the convolution kernel parameters are not the same during each inference process, which improves the flexibility of the model and enhances the representation ability of the model.

[0079] In the final implementation process, Figure 4 As shown in the figure, each convolution of the time-delay neural network is replaced with this dynamically adjusted convolution kernel. Some nodes can also be selectively replaced according to actual conditions. In addition, k is also an adjustable parameter. Here, k does not mean that the more k, the stronger the model representation ability. The increase in k may also lead to an increase in model parameters, and the model is prone to overfitting. The specific setting for the optimal effect needs to be further verified according to actual conditions.

[0080] The above is only a specific implementation of the present invention, which enables those skilled in the art to understand or implement the present invention. The present invention improves the original time-delay neural network so that the network can dynamically adjust the acoustic model parameters according to different inputs, thereby improving the accuracy of dialect recognition.

Claims

1. A new dynamic time-delay neural network for dialect recognition, characterized in that: The specific structure is as follows: the convolution kernel of the existing time-delay neural network is updated: first, k parallel convolution kernels are introduced, and their respective attention weight parameters are calculated, and then weighted average fusion is performed to obtain a convolution kernel with an optimal weight parameter; finally, the convolution kernel with the optimal weight parameter replaces the conventional convolution kernel of the existing time-delay neural network to obtain an improved dynamic time-delay neural network; when different dialects are input, the weight parameters in the convolution kernel with the optimal weight parameter are dynamically adjusted to extract deeper feature information of the audio, and the extracted deep features are compared and analyzed with the dialect acoustic template, and finally the dialect category is determined in combination with the classifier.

2. A novel dynamic time-delay neural network for dialect recognition as claimed in claim 1, characterized in that: The specific process of updating the convolution kernel is as follows: Step 1: Perform global average pooling on the input audio features to obtain global spatial features of length C; Step 2: Map the global spatial features to k dimensions through two fully connected layers in sequence, perform softmax normalization, and obtain the weight coefficients of the k convolution kernels; The formula is as follows: Among them, Z k is the kth output of the second fully connected layer, Softmax(Z k ) is the weight coefficient of the kth convolution kernel; Step 3: Take the weighted average of all parallel convolution kernel weights according to their respective weight coefficients and fuse them to obtain a new convolution kernel. The weight of each convolution kernel is W conv,k ∈R N×N×C , where N is the size of the convolution kernel; Fusion gets the new convolution kernel W new , the formula is: π k is the weight coefficient of the kth convolution kernel, W conv,k is the original weight of the kth convolution kernel; Step 4: Input the new convolution kernel into the time-delay neural network model so that it can adaptively capture the deep features of the input dialect; The output formula of the improved dynamic time-delay neural network model is: C t ={X t-n ,X t+n } 0≤π k ≤1, Y t ' is the output feature, f is the activation function, W T and b are fixed weight matrices and bias vectors, respectively, C t is the input feature, X t-n The number of upper information frames input to the lower network at time t is n; X t+n The number of context information frames input to the lower network at time t is n, where n is greater than or equal to 1; Step 5: Compare and analyze the extracted deep features with the dialect acoustic template, and finally determine the dialect category by combining the classifier.

3. A novel dynamic time-delay neural network for dialect recognition as claimed in claim 2, characterized in that: In the step 1, the input audio feature is a tensor of H×W×C; The global average pooling operation calculates the mean of each channel in its spatial dimension. The output of the global average pooling of the cth channel is GlobalFeature c The formula is as follows: Where H is the height of the feature map, W is the width of the feature map, C is the total number of channels, and X i,j,c Represents the eigenvalue of the cth channel at position (i, j), c∈C.

4. A novel dynamic time-delay neural network for dialect recognition as claimed in claim 2, characterized in that: In the step 4, the improved time-delay neural network model can adaptively select the optimal feature extraction mode according to the audio content of the input dialect: strengthen the short-time impulse response in the transient stage of the dialect pronunciation, and enhance the resonance peak tracking ability in the steady-state stage.

5. The novel dynamic time-delay neural network for dialect recognition as claimed in claim 1, characterized in that: k is an adjustable parameter. Each computing node of the time-delay neural network is replaced with a dynamically adjusted convolution weight, or some convolution weights are selectively replaced according to actual conditions.

Citation Information

Patent Citations

  • Dialect species recognition method based on extended convolutional neural network

    CN111243575A

  • Dialect recognition system and training method thereof

    CN116030793A

  • Voice task processing method and device, electronic equipment and storage medium

    CN116129881A

  • Multi-language speech recognition method and system based on generative learning model

    CN118841000A

  • Speech dialect classification for automatic speech recognition

    US20120109649A1