Dialect language recognition method and system based on self-supervised speaker representation decoupling
By introducing gradient inversion layers and classifiers into the self-supervised model, decoupling speaker information is solved, and the problem of difficulty in decoupling of supervised learning in the existing technology is solved, and a more efficient dialect language recognition effect is achieved.
Patent Information
- Application Number
- CN202510102857.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-01-22
AI Technical Summary
The prior art relies on supervised learning of a large number of labeled data in dialect language recognition, and it is not clear whether the information contained in the self-supervised representation is beneficial to the dialect language recognition task, which makes it difficult to achieve information decoupling.
Through a self-supervised speaker representation decoupling method, a self-supervised model composed of convolutional neural network and Transformer is used to extract speech features and convert them into self-supervised speech representation through an average pooling layer. Combined with a gradient inversion layer and a classifier, the speaker information is decoupled to improve the effect of dialect language recognition.
The unnecessary speaker information in the self-supervised model was successfully decoupled, improving the effectiveness of dialectical language tasks, especially showing excellent performance under low resource conditions.
Smart Images

Figure CN119964554A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a dialect variety recognition method and system based on self-supervised speaker representation decoupling, and the technical field of natural language processing. Background Art
[0002] Speech recognition technology has a wide range of applications, including voice assistants, smart homes, car voice interaction, and many other fields. Dialects are a popular way for Chinese people to communicate in daily life. Dialect Identification (DID) refers to the analysis and processing of input speech sequences to determine the dialect to which they belong. The core challenge of dialect identification is how to extract features that can effectively distinguish different dialects.
[0003] Traditional dialect recognition methods use a large amount of labeled data for training in a supervised learning manner. First, the underlying acoustic features such as Mel-frequency cepstral coefficients, filter banks, i-vectors, and x-vectors are extracted from the speech signal. The corresponding dialect representation is then extracted using a front-end encoder. The back-end independently trained classifier is then applied to identify the dialect category. Benefiting from the rapid development of deep neural network technology, the front-end and back-end methods are integrated into a single network through an end-to-end approach. Although supervised learning has become the core method of speech processing, it requires a large amount of labeled data for each task and scenario.
[0004] In recent years, self-supervised learning (SSL) has made significant progress in the field of speech processing. Similar to traditional acoustic features, self-supervised speech representations contain a lot of information that is very beneficial to downstream tasks, such as speech emotion recognition, language identification, and speaker verification. Although self-supervised representations have made progress in dialect recognition tasks, it is not clear whether all the information contained in these representations is beneficial to dialect recognition tasks. Therefore, it is crucial to analyze the information in the self-supervised representations and decouple irrelevant or unnecessary information. Summary of the invention
[0005] The technical problem solved by the present invention is: the present invention provides a dialect variety recognition method and system based on self-supervised speaker representation decoupling, which is used to decouple unnecessary speaker information in the self-supervised model to extract the difference representations between dialects. The present invention improves the effect of dialect variety tasks.
[0006] The technical solution of the present invention is: a dialect recognition method based on self-supervised speaker representation decoupling, the method comprising:
[0007] Step 1: Input the original speech into the context encoder of the self-supervised model to extract the output context representation, where the self-supervised model consists of a convolutional neural network feature extractor and a Transformer-based context encoder;
[0008] Step 2: Perform weighted summation of the context representations and then use an average pooling layer to convert the result into a self-supervised speech representation.
[0009] Step 3: The self-supervised speech representation is input into the speaker gender classifier and the dialect classifier to obtain the decoupled speaker representation S and dialect representation D respectively;
[0010] Step 4: Use speaker representation and dialect representation to train the self-supervised model based on the gradient descent algorithm, adjust the network weight parameters by back-propagating the gradient of the loss function relative to each weight in the network, and use the trained self-supervised model to identify the dialect.
[0011] Furthermore, in Step 2, the generation process of the self-supervised speech representation U is expressed as follows:
[0012]
[0013] Among them, w i is the weight of layer i, C i represents the output of the i-th layer of the context encoder, Mean represents the mean pooling operation, and N represents the number of encoder layers.
[0014] Furthermore, in Step 3, the speaker gender classifier and the dialect classifier both use a cross-entropy loss function to calculate their respective losses; the loss functions for dialect and gender classification are defined as follows:
[0015]
[0016] Among them, d i and i Represent the real dialect and speaker gender label respectively; and Represents the predicted probability score; N d and N s is the number of all dialect labels and the number of gender labels; L d and L s They are the loss functions for dialect and gender classification respectively.
[0017] Furthermore, the Step 3 also includes adding a gradient reversal layer before the speaker gender classifier to decouple speaker information.
[0018] Furthermore, in the Step 4, during the training of the self-supervised model, the gradient reversal layer does not substantially participate in the forward propagation of the self-supervised model; during the reverse propagation of the self-supervised model training, unlike the general gradient descent algorithm, the gradient reversal layer multiplies the gradient of the subsequent layer by -λ and passes it to the previous layer, so that the weight parameters of the model will be trained to be away from the target distribution when updated; the weight parameters θ of the model m The update formula is as follows:
[0019]
[0020] Where μ is the learning rate, λ is a manually adjustable hyperparameter, and θ s represents the weight parameter of the speaker classifier, μ represents the learning rate, and θ d Represents the weight parameter of the dialect classifier.
[0021] The present invention also provides a dialect variety recognition system based on self-supervised speaker representation decoupling, the system comprising: a module for executing the above-mentioned dialect variety recognition method based on self-supervised speaker representation decoupling.
[0022] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the above-mentioned dialect recognition method based on self-supervised speaker representation decoupling is implemented.
[0023] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the dialect recognition method based on self-supervised speaker representation decoupling is implemented.
[0024] The present invention also provides a computer program product, including a computer program, characterized in that when the computer program is executed by a processor, the method for dialect recognition based on self-supervised speaker representation decoupling is implemented.
[0025] This paper explores the application of the Self-Supervised Model (SSM) in the dialect recognition task; analyzes the impact of speaker information on dialect recognition and visualizes the speaker information between different layers of the self-supervised model; and proposes a method of decoupling speaker-related information from the self-supervised representation by hierarchical use of the Gradient Reversal Layer (GRL).
[0026] The beneficial effects of the present invention are:
[0027] 1. The feature decoupling method proposed in the present invention successfully decouples speaker information from the self-supervised model, improving the effect of dialect tasks.
[0028] 2. Experimental results on the KeSpeech dialect public dataset show that the model proposed in this paper exhibits excellent performance in dialect recognition tasks.
[0029] 3. The model proposed in this invention shows excellent performance in language recognition under low resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 This is a diagram of the architecture of the dialect recognition model based on self-supervised speaker representation decoupling proposed by the present invention; DETAILED DESCRIPTION
[0031] Example 1: Figure 1 As shown, a dialect recognition method based on self-supervised speaker representation decoupling comprises:
[0032] Step 1: Input the original speech into the context encoder of the self-supervised model to extract the output context representation, where the self-supervised model consists of a convolutional neural network feature extractor and a Transformer-based context encoder;
[0033] Step 2: Perform weighted summation of the context representations and then use an average pooling layer to convert the result into a self-supervised speech representation.
[0034] Furthermore, in Step 2, the generation process of the self-supervised speech representation U is expressed as follows:
[0035]
[0036] Among them, w i is the weight of layer i, C i represents the output of the i-th layer of the context encoder, Mean represents the mean pooling operation, and N represents the number of encoder layers.
[0037] Step 3: The self-supervised speech representation is input into the speaker gender classifier and the dialect classifier to obtain the decoupled speaker representation S and dialect representation D respectively;
[0038] Furthermore, in Step 3, the speaker gender classifier and the dialect classifier both use a cross-entropy loss function to calculate their respective losses; the loss functions for dialect and gender classification are defined as follows:
[0039]
[0040] Among them, d i and i Represent the real dialect and speaker gender label respectively; and Represents the predicted probability score; N d and N s is the number of all dialect labels and the number of gender labels; L d and L s They are the loss functions for dialect and gender classification respectively.
[0041] Furthermore, the Step 3 also includes adding a gradient reversal layer before the speaker gender classifier to effectively decouple speaker information.
[0042] Step 4: Use speaker representation and dialect representation to train the self-supervised model based on the gradient descent algorithm, adjust the network weight parameters by back-propagating the gradient of the loss function relative to each weight in the network, and use the trained self-supervised model to identify the dialect.
[0043] Furthermore, in the Step 4, during the training of the self-supervised model, the gradient reversal layer does not substantially participate in the forward propagation of the self-supervised model; during the reverse propagation of the self-supervised model training, unlike the general gradient descent algorithm, the gradient reversal layer multiplies the gradient of the subsequent layer by -λ and passes it to the previous layer, so that the weight parameters of the model will be trained to be away from the target distribution when updated; the weight parameters θ of the model m The update formula is as follows:
[0044]
[0045] Wherein, μ is the learning rate, λ is a manually adjustable hyperparameter, which is set to 0.15 in the present invention, and θ s represents the weight parameter of the speaker classifier, μ represents the learning rate, and θ d Represents the weight parameter of the dialect classifier.
[0046] The present invention also provides a dialect recognition system based on self-supervised speaker representation decoupling, the system comprising:
[0047] A context representation extraction module is used to extract the output context representation from the original speech input into the context encoder of the self-supervised model, where the self-supervised model is composed of a convolutional neural network feature extractor and a Transformer-based context encoder;
[0048] A self-supervised speech representation generation module that performs a weighted sum of the context representations and then uses an average pooling layer to convert the result into a self-supervised speech representation;
[0049] A speaker representation and dialect representation generation module is used to input the self-supervised speech representation into the speaker gender classifier and the dialect classifier to obtain the decoupled speaker representation and dialect representation respectively;
[0050] The dialect recognition module is used to use speaker representation and dialect representation to train the self-supervised model based on the gradient descent algorithm, adjust the weight parameters of the network by back-propagating the gradient of the loss function relative to each weight in the network, and use the trained self-supervised model to recognize the dialect.
[0051] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the above-mentioned dialect recognition method based on self-supervised speaker representation decoupling is implemented.
[0052] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the dialect recognition method based on self-supervised speaker representation decoupling is implemented.
[0053] The present invention also provides a computer program product, including a computer program, characterized in that when the computer program is executed by a processor, the method for dialect recognition based on self-supervised speaker representation decoupling is implemented.
[0054] In order to decouple unnecessary speaker information in the self-supervised model, the present invention explores the application of the self-supervised model (SSM) in the dialect recognition task; analyzes the impact of speaker information on dialect recognition and visualizes the speaker information between different layers of the self-supervised model; and proposes a method of decoupling speaker-related information from the self-supervised representation by hierarchical use of gradient reversal layers (GRL).
[0055] For dialect identification tasks, current research usually adds an average pooling layer and a dialect classifier to the self-supervised model; the input of the self-supervised model is the original speech sequence, and then the feature extractor based on the convolutional neural network extracts the features, and the context encoder based on the Transformer extracts the context representation; the dialect classifier usually contains complex structures, such as convolutional neural networks or self-attention layers. In order to fully reflect the superiority of the proposed representation decoupling method, the classifiers used in the invention are all simple linear projection layers. Figure 1 As shown, an average pooling layer, a gradient reversal layer, a speaker gender classifier and a dialect classifier following the self-supervised model are integrated on the basis of the self-supervised model.
[0056] In order to illustrate the effect of the present invention, the present invention has done the following experiment:
[0057] (a) Experimental settings and evaluation indicators
[0058] The experiments were mainly conducted and analyzed on the public Chinese dialect dataset Kespeech. The Kespeech dataset consists of 1542 hours of speech data annotated by 27,237 speakers from 34 cities in China. The dataset includes eight Mandarin dialects and Putonghua data: Northeast, Jilu, Central Plains, Jiaoliao, Jianghuai, Southwest, Lanyin, and Beijing.
[0059] The evaluation index in the experiment mainly uses multi-classification accuracy (Acc), and the accuracy is the number of correctly classified samples L c The number of all samples is L a Ratio:
[0060]
[0061] In the experiment, all audio samples were set to 16KHz, single channel, and 16-bit bit depth. The three pre-trained self-supervised models (Wav2vec2.0, WavLM, and HuBERT) all used open source models from the Huggingface community, and all used the base model. The fine-tuning process was 10 rounds, during which the convolutional feature extractor was frozen and parameters were not updated. The learning rate was set to 0.003, and the lambda of the gradient reversal layer was set to 0.15.
[0062] (b) Comparative experiment
[0063] The experiment of this invention aims to compare the performance of different models in dialect recognition, and the experimental results are shown in Table 1. The table shows the baseline method and the existing state-of-the-art method on the KeSpeech multi-dialect dataset, where each column corresponds to the dialect classification accuracy of a specific dialect, and the last column calculates the overall classification accuracy.
[0064] Table 1 shows the comparative experiments of different models on the test set.
[0065] Model northeast Giroud Central Plains Jiaoliao JAC southwest Blue Silver Beijing overall M1 - 43.22 52.63 58.56 53.19 82.47 43.76 - 56.34 M2 - 39.37 76.65 37.84 64.44 82.32 42.67 - 61.13 M3 - 48.45 68.54 36.52 68.56 80.91 44.18 - 60.77 M4 - - - - - - - - 78.57 M5 - - - - - - - - 79.10 M6 76.53 62.41 67.91 76.45 72.33 89.34 69.51 64.13 75.86 M7 78.26 68.32 73.61 76.94 72.75 89.79 70.13 64.91 75.50 M8 80.47 70.56 71.25 77.45 75.31 90.34 71.16 66.35 76.45 M9 83.76 72.25 63.69 81.57 77.92 91.64 76.41 69.21 78.12 M10 85.18 71.72 73.81 82.43 78.65 92.03 76.13 70.09 78.62 M11 87.72 72.45 77.19 82.06 82.34 93.73 77.68 72.33 79.93
[0066] M1-M3 correspond to traditional methods and network architectures. Traditional methods are not applicable to low-resource conditions, such as Northeastern Mandarin and Beijing Mandarin in the table. M1 is a ResNet-34 model, whose structure is a stacked residual convolutional network structure. M2 uses the Kaldi open source toolkit to extract x-vector features from audio for dialect classification. M3 is ECAPA-TDNN, which is a speaker model based on time-delay neural network. Due to its superior performance, it has been widely used in the field of speaker recognition. ECAPA-TDNN has also been applied to tasks such as speaker classification and language recognition. Since the network architecture is not applicable to low-resource tasks, no evaluation indicators are given for low-resource dialects such as Northeastern dialect and Beijing dialect.
[0067] M4 and M5 correspond to DIMNet and Qifusion-Net respectively, which are two multi-task models trained on dialect recognition and speech recognition tasks simultaneously. Due to the multi-task environment, the accuracy of each dialect classification is not provided separately.
[0068] M6-M8 are the experimental results of Wav2vec2.0, HuBERT and WavLM after fine-tuning using multi-dialect data, while M9-M11 are the experimental results of the three self-supervised models after using the feature decoupling method proposed in this invention.
[0069] From the analysis of Table 1, it can be clearly observed that even without the feature decoupling method, the experimental results of the self-supervised models (M6-M8) after fine-tuning are better than those of the traditional methods and model architectures, reaching between 75.96% and 76.45% respectively. Since the multi-task architecture learns both speech representation and semantic representation at the same time, that is, the corresponding dialect recognition task and speech recognition task, it may be richer in representation information than the self-supervised representation. The performance of the self-supervised models (M9-M11) fine-tuned using the feature decoupling method has been further improved, with accuracy rates of 78.12%, 78.62% and 79.93% respectively, surpassing all current baseline models, thus proving the superiority of the method.
[0070] The specific implementation modes of the present invention are described in detail above in conjunction with the accompanying drawings, but the present invention is not limited to the above implementation modes, and various changes can be made within the knowledge scope of ordinary technicians in this field without departing from the purpose of the present invention.
Claims
1. A dialect recognition method based on self-supervised speaker representation decoupling, characterized by: The method comprises: Step 1: Input the original speech into the context encoder of the self-supervised model to extract the output context representation; Step 2: Perform weighted summation of the context representations and then use an average pooling layer to convert the result into a self-supervised speech representation. Step 3: The self-supervised speech representation is input into the speaker gender classifier and the dialect classifier to obtain the decoupled speaker representation and dialect representation respectively; Step 4: Use speaker representation and dialect representation to train the self-supervised model based on the gradient descent algorithm, adjust the network weight parameters by back-propagating the gradient of the loss function relative to each weight in the network, and use the trained self-supervised model to identify the dialect.
2. The dialect identification method based on self-supervised speaker representation decoupling according to claim 1, characterized in that: In Step 2, the generation process of the self-supervised speech representation U is expressed as follows: Among them, w i is the weight of layer i, C i represents the output of the i-th layer of the context encoder, Mean represents the mean pooling operation, and N represents the number of encoder layers.
3. The dialect identification method based on self-supervised speaker representation decoupling according to claim 1, characterized in that: In Step 3, the speaker gender classifier and the dialect classifier both use the cross entropy loss function to calculate their respective losses; the loss functions for dialect and gender classification are defined as follows: Among them, d i and i Represent the real dialect and speaker gender label respectively; and Represents the predicted probability score; N d and N s is the number of all dialect labels and the number of gender labels; L d and L s They are the loss functions for dialect and gender classification respectively.
4. The dialect identification method based on self-supervised speaker representation decoupling according to claim 1, characterized in that: The Step 3 also includes adding a gradient reversal layer before the speaker gender classifier to decouple speaker information.
5. The dialect identification method based on self-supervised speaker representation decoupling according to claim 1, characterized in that: In the step 4, during the training of the self-supervised model, the gradient reversal layer does not substantially participate in the forward propagation of the self-supervised model; during the reverse propagation of the self-supervised model training, the gradient reversal layer multiplies the gradient of the subsequent layer by -λ and passes it to the previous layer, so that the weight parameters of the model will be trained to be far away from the target distribution when updated; the weight parameters of the model θ m The update formula is as follows: Where μ is the learning rate, λ is a manually adjustable hyperparameter, and θ s represents the weight parameter of the speaker classifier, μ represents the learning rate, and θ d Represents the weight parameter of the dialect classifier.
6. A dialect recognition system based on self-supervised speaker representation decoupling, characterized by: The system comprises: a module for executing the dialect recognition method based on self-supervised speaker representation decoupling according to any one of claims 1 to 5.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the dialect recognition method based on self-supervised speaker representation decoupling as described in any one of claims 1 to 5 is implemented.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for dialect recognition based on self-supervised speaker representation decoupling as claimed in any one of claims 1 to 5 is implemented.
9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for dialect recognition based on self-supervised speaker representation decoupling as claimed in any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Dialect classification method and system based on self-supervised voice representation
CN116631375A
Decoupling type voice self-supervision pre-training method
CN118841029A
Multi-dialect speech recognition method, device, equipment and medium
CN119107936A